Eval: honest product imagery with model consent (77)
Run on our server, 27 Sep 2026, branch the pre-release branch, gateway route (Qwen3.8-27B through the model gateway,
receipted), document reader service on :8497, under heavy load from other workloads (load average up to 80). Scripts:
scripts/imagery_scenes.py (the renders), scripts/imagery_eval.py (build, run, score), scripts/imagery_opus.py
(the frontier reference judge, eval only). Raw results on our server in <internal path>:
eval/results-dev-final.jsonl, eval/results.jsonl (dev + test as frozen), eval-post/results.jsonl (11 test cases re-run
after the fixes), eval/results-merged.jsonl, eval/opus_results.jsonl, eval/opus_results_zoom.jsonl, sample-runs.json.
What was measured
- Fidelity check: does an AI image that misrepresents the product get flagged, and does a faithful one pass?
- Consent gate: is every image with a person refused unless a consent-ledger entry covers this advertising use?
- Disclosure: does each target get the right marking, and does it survive (XMP read back, label OCR'd, C2PA valid)?
- Frontier gap (The owner's request): the same comparison prompt through Claude Opus 5.5, per error type.
Data
- Six made-up products drawn with Pillow (
decosa_api/verticals/imagery/synth.py; Lato SIL OFL 1.1 and DejaVu fonts): a hand wash, a tea box, a drink can, a vitamin jar, a toothpaste carton, a candle. Each has a logo mark, a wordmark, a name, a variant, a claim, a net quantity and a warning. The seller's reference photo is the pack on a plain studio background. - Lifestyle renders: 22 scenes drawn by Wan2.2-VACE-Fun-A14B (Apache-2.0) in the shared ComfyUI on GPU0, one job at a time, only after checking the queue was empty and at least 14 GB was free (about 30-45 s each; nothing else was loaded, no service was stopped). The pack is placed on the canvas and the generator draws the scene around it.
- Split by product, fixed before any check ran: dev = tea box, candle; test = hand wash, can, vitamin jar, toothpaste.
- Planted set: every render labelled faithful by the building agent gets three versions (as rendered; cropped, resized to 1024 px and JPEG q75; warm grade, slight blur, 3 degree tilt, JPEG q85) and six planted misrepresentations, each in one of those versions: size (net quantity or count changed, e.g. 500 ml to 750 ml), colour (the pack's main colour hue-rotated in the render's own pixels), text (a claim the product does not make), logo (another mark and wordmark font), warning (left off), feature (a pump, a second unit or a red "FREE GIFT" badge). Plants are rendered with the product spec and pasted over the render where they differ, colour-matched to the render.
- Wild set: every raw render with a hand label. Six were not faithful by themselves: the generator swapped the hand wash's screw cap for a pump three times (once when the prompt said "screw cap") and for a spray top once, drew two extra toothpaste cartons with garbled brand text, and redrew a carton as a bottle. Two renders contain people the prompt did not ask for (a hand pouring tea; diners at the next tables).
Results
Held-out test (4 products; thresholds frozen before it ran)
Cells: flagged / cases (right category in brackets). "Fail or review" counts an image a person must look at; "fail" is an outright refusal. Faithful images: the 3 versions of each of the 9 faithful test renders (27); the consent gate refused 2 cafe versions and 1 run failed with a 422, leaving 24 (22 with a model comparison).
| layer | size | colour | text | logo | warning | feature | all planted | flagged faithful |
|---|---|---|---|---|---|---|---|---|
| Qwen3.8 + checks, fail or review | 8/8 | 8/8 | 8/8 | 8/8 | 8/8 (7) | 8/8 | 48/48 | 3/24 |
| Qwen3.8 + checks, fail only | 8/8 | 8/8 | 8/8 | 8/8 | 8/8 (7) | 4/8 (3) | 44/48 | 2/24 |
| deterministic checks alone | 8/8 | 8/8 | 8/8 | 8/8 | 7/8 | 6/8 (3) | 45/48 | 1/24 |
| Qwen3.8 alone | 8/8 | 7/8 | 8/8 | 8/8 | 7/7 | 8/8 | 46/47 | 0/22 |
| Claude Opus 5.5 alone (reference) | 8/8 | 8/8 | 8/8 | 8/8 | 7/7 | 8/8 | 47/47 | 9/22 |
- The 3 flagged faithful images: two versions of the vitamin jar on a marble counter, where the model boxed no product and the run said "not found" (counted as a false flag); one toothpaste carton on a stool whose box the generator drew about 12% taller (the new outline check fires at 12%).
- 2 runs failed with HTTP 422 (one planted-set version and the wild copy of the same toothpaste render): the model found no product on the reference photo, twice in a row, under load.
- 6 of the 30 planted cases of the can on a cafe table are not in the table: the consent gate refused them (the generator drew diners in the background; see Consent).
- Qwen's one colour miss (a hue shift on the hand wash, CIEDE2000 13.6) was caught by the colour measurement; the checks' misses (one warning, two features: a pump and a second, partly hidden carton) were caught by Qwen.
After the test run (not held out): four fixes were made because of what the test showed: code finds the product when the model boxes none (a template search; a plain-background packshot is boxed without the model), "not found" became a review instead of a fail, a pack text line and the model's "feature" back each other (a FREE GIFT badge carries text), and the model's unit count counts as a feature. The 11 affected test cases were re-run: 48/48 flagged (46 fail, 2 review), 2/25 faithful flagged (both the taller toothpaste carton), and the 422s are gone.
Dev (2 products; thresholds were set here in three rounds)
24/24 planted flagged (22 fail, 2 review), 0/12 faithful flagged. Qwen3.8 alone 24/24 and 0/12; Opus 5.5 24/24 and 2/12. Round 1 used a logo threshold of 0.5 (planted logos scored 0.60-0.63, faithful 0.95+: set to 0.8) and a pack-area rule that fired on a 3 degree tilt (fixed with a rotation search and blurring before the comparison, and by leaving out the studio floor shadow).
Wild renders (not planted)
- The 5 generator misrepresentations in test were all flagged (4 fail, 1 review): three pump or spray swaps, the extra
garbled cartons, the carton drawn as a bottle. The pump is the demo sample
pump-swap. - Of the 9 faithful test renders (by the first hand label), 6 passed as frozen; the cafe scene was refused for its bystanders, the jar on marble was "not found" and the carton on a stool hit the 422. Both pass after the fixes.
What the frontier reference showed, and the gap it found
Claude Opus 5.5 (claude-opus-5-5, Anthropic API, 27 Sep 2026) got the same system prompt, schema and images (JPEG q90,
≤1024 px) as the Qwen path. Qwen also got a zoomed crop of the product when it is small; a second Opus run with a crop from
the known pack position ("zoom") is reported too. 270 calls, 838,858 input and 90,327 output tokens, US$5.16 in all
(list price $4/$20 per million), median 4.5 s per call. Eval only: the hosted service stays on Qwen3.8 through the gateway.
- On the planted errors the two models are level: Opus 47/47 and 24/24, Qwen 46/47 and 24/24. Qwen's one miss was closed by the colour check, so Qwen plus the deterministic checks catches everything Opus catches on the planted set.
- Opus flagged 9 of 22 "faithful" test images (12 with the zoom crop). Looking at those images again at full size, most of its flags were right and our hand labels were wrong: the generator had garbled the small warning and claim text on five renders ("High caffaino cuntare. Nat fr children", "Far extarnal use anly. Aroid contact", "No aulded sugar", "Bum time"), and drawn two cartons with an open top. Re-labelled by eye (prompted by Opus's flags, so not blind); dev and test together, images with a model comparison:
| faithful renders, re-inspected | images | Qwen + checks | checks alone | Qwen alone | Opus | Opus + zoom |
|---|---|---|---|---|---|---|
| clean | 17 | 0 | 0 | 0 | 1 | 1 |
| small text garbled by the generator | 17 | 0 | 0 | 0 | 11 | 15 |
| carton top redrawn | 10 | 1 | 1 | 0 | 2 | 2 |
- This is the gap: legibility of small printed text. The document reader's parser (PaddleOCR-VL, a vision-language reader) "reads through" garbled type and returns the correct words, and Qwen3.8 does not notice either. A probe after the eval (not wired in, measured on the same 13 renders, so not held out): Tesseract, which reads the pixels literally, got 25-50% of the warning's words exactly on the 5 garbled renders and 58-100% on the 8 clean ones.
Consent (live runs on the pre-release server, sample-runs.json)
Every refusal came before the comparison (2 model calls: the two locate calls), with a signed consent decision:
| case | result |
|---|---|
person, no identity given (no-identity) |
refused: "the image shows 1 person and no consent-ledger identity was given" |
generator-drawn bystanders (bystanders) |
refused: 2 people, no identity |
consent for another campaign (wrong-campaign) |
refused, code project_not_in_scope |
consent withdrawn (withdrawn-consent) |
refused, code revoked |
| consent expired on 31 Mar 2026 | refused, code expired |
| an id that is not in the ledger | refused, code no_entry |
| synthetic performer used in Japan (enrolled for US, GB, DE, FR) | refused, code territory_not_covered |
synthetic performer, covered (synthetic-model) |
approved: label burned in and read back, Amazon keyword, NY s.396-b "met" |
real model, covered, but the bottle has a spray top (consented-but-wrong-cap) |
consent granted, image not approved (feature) |
| In the eval the gate also refused the tea scene where the generator added a hand, and 9 of the 10 cafe images (the one | |
| crop that cut the diners out was checked and approved). |
Disclosure
On every approved image: XMP with the IPTC DigitalSourceType was read back from the file (Google Merchant Center, Meta, EU
Art. 50(2)); the C2PA credential validated ("Valid", development CA: public validators show it as untrusted) with
c2pa.created and the same source type. For a synthetic performer, Amazon's contains-synthetic-performer keyword was
added and the "AI-generated image" label was burned in and read back by Tesseract (NY s.396-b "met"); for a real model's
consented replica, no Amazon keyword (Amazon says not to use it for real people) and the EU Art. 50(4) label. Etsy is
reported as advice only (its rule covers items made with AI, not photos of real items), Meta's ad-disclosure rule as not
verified. Unit tests cover each rule row.
Cost and time (under load)
Per checked image: 3 Qwen3.8 image calls (about 4,400 prompt tokens), 6-8 document-reader parses (model-call receipts), about US$0.0018 at gateway list price; about 10 receipts. p50 23 s, p90 38 s per check on test, with 2 checks at a time on a shared, overloaded box (12.9 s median measured by the site's recorder at a quieter moment). A consent refusal costs 2 calls (about $0.001) and 3 s.
Expected properties of the demo samples (the rehearsal bundle checks these)
size-750ml: statusnot_approved, asizeviolation from the document reader ("750 ml" against 500 ml), no image.no-identityandbystanders: statusconsent_refusedafter two model calls; no comparison, no credential.faithful-shelf: statusapproved,xmp.digital_source_typeends intrainedAlgorithmicMedia,credential.kindisc2pa, the record verifies at/record/verifyand fails onceimage_sha256is changed.synthetic-model: approved withlabel.text[0] == "AI-generated image",label_readback.ok, the Amazon keyword inxmp.keywords, and the NY rulemet.- Every Qwen3.8 call has a gateway receipt; every page parse a signed model-call receipt; every consent lookup a decision id.
Caveats
- Synthetic products and one generator. The same author made the products, the plants, the checker and the labels. Plants are clean edits; a generator's own drift (the wild set) is the more realistic case and there are only 6 of those.
- Small: 48 planted and 24 faithful test images from 8 renders of 4 products.
- The first hand labels missed garbled small text on 5 of 13 faithful renders; the re-inspection was prompted by Opus's flags. The planted-set numbers are unaffected (each plant is on top of its render), but "flagged faithful" for Opus is mostly real problems, and for our checker the garbled text is a miss we did not see.
- Thresholds were set on dev in three rounds; four fixes followed the test run (above) and their numbers are not held out.
- The consent gate refuses any human likeness without an identity, including generator-drawn bystanders and hands: strict by design (G3), and a real cost in friction for sellers.
- Not legal advice and not a compliance determination.
Verdict
Worth a paying pilot for sellers and agencies that make product images with AI: on this data every planted misrepresentation and every generator swap we found was flagged, a person with no consent on file never gets through, and the disclosure is applied and read back. It is not yet a legibility checker: it passed generator-garbled warnings that a frontier model caught. That is the next build (see Capability opportunities in the report).