Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: honest product imagery with model consent (77)

Run on our server, 27 Sep 2026, branch the pre-release branch, gateway route (Qwen3.8-27B through the model gateway, receipted), document reader service on :8497, under heavy load from other workloads (load average up to 80). Scripts: scripts/imagery_scenes.py (the renders), scripts/imagery_eval.py (build, run, score), scripts/imagery_opus.py (the frontier reference judge, eval only). Raw results on our server in <internal path>: eval/results-dev-final.jsonl, eval/results.jsonl (dev + test as frozen), eval-post/results.jsonl (11 test cases re-run after the fixes), eval/results-merged.jsonl, eval/opus_results.jsonl, eval/opus_results_zoom.jsonl, sample-runs.json.

What was measured

  1. Fidelity check: does an AI image that misrepresents the product get flagged, and does a faithful one pass?
  2. Consent gate: is every image with a person refused unless a consent-ledger entry covers this advertising use?
  3. Disclosure: does each target get the right marking, and does it survive (XMP read back, label OCR'd, C2PA valid)?
  4. Frontier gap (The owner's request): the same comparison prompt through Claude Opus 5.5, per error type.

Data

  • Six made-up products drawn with Pillow (decosa_api/verticals/imagery/synth.py; Lato SIL OFL 1.1 and DejaVu fonts): a hand wash, a tea box, a drink can, a vitamin jar, a toothpaste carton, a candle. Each has a logo mark, a wordmark, a name, a variant, a claim, a net quantity and a warning. The seller's reference photo is the pack on a plain studio background.
  • Lifestyle renders: 22 scenes drawn by Wan2.2-VACE-Fun-A14B (Apache-2.0) in the shared ComfyUI on GPU0, one job at a time, only after checking the queue was empty and at least 14 GB was free (about 30-45 s each; nothing else was loaded, no service was stopped). The pack is placed on the canvas and the generator draws the scene around it.
  • Split by product, fixed before any check ran: dev = tea box, candle; test = hand wash, can, vitamin jar, toothpaste.
  • Planted set: every render labelled faithful by the building agent gets three versions (as rendered; cropped, resized to 1024 px and JPEG q75; warm grade, slight blur, 3 degree tilt, JPEG q85) and six planted misrepresentations, each in one of those versions: size (net quantity or count changed, e.g. 500 ml to 750 ml), colour (the pack's main colour hue-rotated in the render's own pixels), text (a claim the product does not make), logo (another mark and wordmark font), warning (left off), feature (a pump, a second unit or a red "FREE GIFT" badge). Plants are rendered with the product spec and pasted over the render where they differ, colour-matched to the render.
  • Wild set: every raw render with a hand label. Six were not faithful by themselves: the generator swapped the hand wash's screw cap for a pump three times (once when the prompt said "screw cap") and for a spray top once, drew two extra toothpaste cartons with garbled brand text, and redrew a carton as a bottle. Two renders contain people the prompt did not ask for (a hand pouring tea; diners at the next tables).

Results

Held-out test (4 products; thresholds frozen before it ran)

Cells: flagged / cases (right category in brackets). "Fail or review" counts an image a person must look at; "fail" is an outright refusal. Faithful images: the 3 versions of each of the 9 faithful test renders (27); the consent gate refused 2 cafe versions and 1 run failed with a 422, leaving 24 (22 with a model comparison).

layer size colour text logo warning feature all planted flagged faithful
Qwen3.8 + checks, fail or review 8/8 8/8 8/8 8/8 8/8 (7) 8/8 48/48 3/24
Qwen3.8 + checks, fail only 8/8 8/8 8/8 8/8 8/8 (7) 4/8 (3) 44/48 2/24
deterministic checks alone 8/8 8/8 8/8 8/8 7/8 6/8 (3) 45/48 1/24
Qwen3.8 alone 8/8 7/8 8/8 8/8 7/7 8/8 46/47 0/22
Claude Opus 5.5 alone (reference) 8/8 8/8 8/8 8/8 7/7 8/8 47/47 9/22
  • The 3 flagged faithful images: two versions of the vitamin jar on a marble counter, where the model boxed no product and the run said "not found" (counted as a false flag); one toothpaste carton on a stool whose box the generator drew about 12% taller (the new outline check fires at 12%).
  • 2 runs failed with HTTP 422 (one planted-set version and the wild copy of the same toothpaste render): the model found no product on the reference photo, twice in a row, under load.
  • 6 of the 30 planted cases of the can on a cafe table are not in the table: the consent gate refused them (the generator drew diners in the background; see Consent).
  • Qwen's one colour miss (a hue shift on the hand wash, CIEDE2000 13.6) was caught by the colour measurement; the checks' misses (one warning, two features: a pump and a second, partly hidden carton) were caught by Qwen.

After the test run (not held out): four fixes were made because of what the test showed: code finds the product when the model boxes none (a template search; a plain-background packshot is boxed without the model), "not found" became a review instead of a fail, a pack text line and the model's "feature" back each other (a FREE GIFT badge carries text), and the model's unit count counts as a feature. The 11 affected test cases were re-run: 48/48 flagged (46 fail, 2 review), 2/25 faithful flagged (both the taller toothpaste carton), and the 422s are gone.

Dev (2 products; thresholds were set here in three rounds)

24/24 planted flagged (22 fail, 2 review), 0/12 faithful flagged. Qwen3.8 alone 24/24 and 0/12; Opus 5.5 24/24 and 2/12. Round 1 used a logo threshold of 0.5 (planted logos scored 0.60-0.63, faithful 0.95+: set to 0.8) and a pack-area rule that fired on a 3 degree tilt (fixed with a rotation search and blurring before the comparison, and by leaving out the studio floor shadow).

Wild renders (not planted)

  • The 5 generator misrepresentations in test were all flagged (4 fail, 1 review): three pump or spray swaps, the extra garbled cartons, the carton drawn as a bottle. The pump is the demo sample pump-swap.
  • Of the 9 faithful test renders (by the first hand label), 6 passed as frozen; the cafe scene was refused for its bystanders, the jar on marble was "not found" and the carton on a stool hit the 422. Both pass after the fixes.

What the frontier reference showed, and the gap it found

Claude Opus 5.5 (claude-opus-5-5, Anthropic API, 27 Sep 2026) got the same system prompt, schema and images (JPEG q90, ≤1024 px) as the Qwen path. Qwen also got a zoomed crop of the product when it is small; a second Opus run with a crop from the known pack position ("zoom") is reported too. 270 calls, 838,858 input and 90,327 output tokens, US$5.16 in all (list price $4/$20 per million), median 4.5 s per call. Eval only: the hosted service stays on Qwen3.8 through the gateway.

  • On the planted errors the two models are level: Opus 47/47 and 24/24, Qwen 46/47 and 24/24. Qwen's one miss was closed by the colour check, so Qwen plus the deterministic checks catches everything Opus catches on the planted set.
  • Opus flagged 9 of 22 "faithful" test images (12 with the zoom crop). Looking at those images again at full size, most of its flags were right and our hand labels were wrong: the generator had garbled the small warning and claim text on five renders ("High caffaino cuntare. Nat fr children", "Far extarnal use anly. Aroid contact", "No aulded sugar", "Bum time"), and drawn two cartons with an open top. Re-labelled by eye (prompted by Opus's flags, so not blind); dev and test together, images with a model comparison:
faithful renders, re-inspected images Qwen + checks checks alone Qwen alone Opus Opus + zoom
clean 17 0 0 0 1 1
small text garbled by the generator 17 0 0 0 11 15
carton top redrawn 10 1 1 0 2 2
  • This is the gap: legibility of small printed text. The document reader's parser (PaddleOCR-VL, a vision-language reader) "reads through" garbled type and returns the correct words, and Qwen3.8 does not notice either. A probe after the eval (not wired in, measured on the same 13 renders, so not held out): Tesseract, which reads the pixels literally, got 25-50% of the warning's words exactly on the 5 garbled renders and 58-100% on the 8 clean ones.

Consent (live runs on the pre-release server, sample-runs.json)

Every refusal came before the comparison (2 model calls: the two locate calls), with a signed consent decision:

case result
person, no identity given (no-identity) refused: "the image shows 1 person and no consent-ledger identity was given"
generator-drawn bystanders (bystanders) refused: 2 people, no identity
consent for another campaign (wrong-campaign) refused, code project_not_in_scope
consent withdrawn (withdrawn-consent) refused, code revoked
consent expired on 31 Mar 2026 refused, code expired
an id that is not in the ledger refused, code no_entry
synthetic performer used in Japan (enrolled for US, GB, DE, FR) refused, code territory_not_covered
synthetic performer, covered (synthetic-model) approved: label burned in and read back, Amazon keyword, NY s.396-b "met"
real model, covered, but the bottle has a spray top (consented-but-wrong-cap) consent granted, image not approved (feature)
In the eval the gate also refused the tea scene where the generator added a hand, and 9 of the 10 cafe images (the one
crop that cut the diners out was checked and approved).

Disclosure

On every approved image: XMP with the IPTC DigitalSourceType was read back from the file (Google Merchant Center, Meta, EU Art. 50(2)); the C2PA credential validated ("Valid", development CA: public validators show it as untrusted) with c2pa.created and the same source type. For a synthetic performer, Amazon's contains-synthetic-performer keyword was added and the "AI-generated image" label was burned in and read back by Tesseract (NY s.396-b "met"); for a real model's consented replica, no Amazon keyword (Amazon says not to use it for real people) and the EU Art. 50(4) label. Etsy is reported as advice only (its rule covers items made with AI, not photos of real items), Meta's ad-disclosure rule as not verified. Unit tests cover each rule row.

Cost and time (under load)

Per checked image: 3 Qwen3.8 image calls (about 4,400 prompt tokens), 6-8 document-reader parses (model-call receipts), about US$0.0018 at gateway list price; about 10 receipts. p50 23 s, p90 38 s per check on test, with 2 checks at a time on a shared, overloaded box (12.9 s median measured by the site's recorder at a quieter moment). A consent refusal costs 2 calls (about $0.001) and 3 s.

Expected properties of the demo samples (the rehearsal bundle checks these)

  1. size-750ml: status not_approved, a size violation from the document reader ("750 ml" against 500 ml), no image.
  2. no-identity and bystanders: status consent_refused after two model calls; no comparison, no credential.
  3. faithful-shelf: status approved, xmp.digital_source_type ends in trainedAlgorithmicMedia, credential.kind is c2pa, the record verifies at /record/verify and fails once image_sha256 is changed.
  4. synthetic-model: approved with label.text[0] == "AI-generated image", label_readback.ok, the Amazon keyword in xmp.keywords, and the NY rule met.
  5. Every Qwen3.8 call has a gateway receipt; every page parse a signed model-call receipt; every consent lookup a decision id.

Caveats

  • Synthetic products and one generator. The same author made the products, the plants, the checker and the labels. Plants are clean edits; a generator's own drift (the wild set) is the more realistic case and there are only 6 of those.
  • Small: 48 planted and 24 faithful test images from 8 renders of 4 products.
  • The first hand labels missed garbled small text on 5 of 13 faithful renders; the re-inspection was prompted by Opus's flags. The planted-set numbers are unaffected (each plant is on top of its render), but "flagged faithful" for Opus is mostly real problems, and for our checker the garbled text is a miss we did not see.
  • Thresholds were set on dev in three rounds; four fixes followed the test run (above) and their numbers are not held out.
  • The consent gate refuses any human likeness without an identity, including generator-drawn bystanders and hands: strict by design (G3), and a real cost in friction for sellers.
  • Not legal advice and not a compliance determination.

Verdict

Worth a paying pilot for sellers and agencies that make product images with AI: on this data every planted misrepresentation and every generator swap we found was flagged, a person with no consent on file never gets through, and the disclosure is applied and read back. It is not yet a legibility checker: it passed generator-garbled warnings that a frontier model caught. That is the next build (see Capability opportunities in the report).