77 · Sales and marketing · Compliance and trust · preview
Honest product imagery
Eval results
Not held outRun 27 Sep 2026Eval write-up (decosa-api, access required)
- Planted misrepresentations flagged48 / 48test splitn = 4844 refused, 4 sent to a person; size, colour, text, logo, warning and feature, 8 each. Thresholds were frozen before the test ran; fixes followed (below).
- Faithful images flagged3 / 24test splitn = 242 where the model boxed no product (now found by a template search), 1 carton the generator drew 12% taller. 2 / 25 after the fixes.
- Generator misrepresentations flagged (not planted)5 / 5test splitn = 5The image model swapped a screw cap for a pump or spray top, added cartons, redrew a carton as a bottle.
- Planted flagged by Qwen3.8 alone / Claude Opus 5.5 alone46 / 47 and 47 / 47test splitn = 47Opus 5.5 is an eval-only reference judge (same prompt and images); the colour check closed Qwen's one miss.
- Generator-garbled small text caught0 / 17 (Opus 5.5: 15 / 17)syntheticn = 17Found after the test run: re-labelled by eye, prompted by Opus's flags. The main gap.
- Consent refusals7 / 7syntheticn = 7No identity, bystanders, other campaign, withdrawn, expired, not enrolled, territory; each before the comparison, with a signed decision.
- Dev: planted flagged / faithful flagged24 / 24 and 0 / 12dev (tuned on)n = 36Thresholds set here in three rounds.
Dataset
Six made-up products (2 dev, 4 test, split by product); 22 lifestyle scenes drawn by Wan2.2-VACE-Fun-A14B around the packs; each faithful render in three versions plus six planted misrepresentations; the raw renders hand-labelled.
Caveats
- Same author made the products, the plants, the checker and the labels; one image generator; synthetic products only.
- Small: 48 planted and 24 faithful test images from 8 renders of 4 products, and 5 generator misrepresentations.
- The test split was used after the frozen run to fix four things; those re-run numbers are not held out.
- The first hand labels missed generator-garbled warnings on 5 of 13 faithful renders; the checker passed all of them.
- Any person, including generator-drawn bystanders and hands, needs a consent-ledger identity: strict by design.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 27 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 13 s
- Receipts
- 10
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.002
Self-host verification
Verified on 27 Sep 2026: Fresh clone of the branch into a clean directory, docker build of the api image (45 s), the api with named volumes on host networking against the running local Qwen3.8-27B (vision, direct route) and document reader, the C2PA dev certificate from the prompt's step; then torn down.
The rehearsal bundle passed 11 of 11 in 17 s (750 ml refused, a person without consent refused after two calls, a faithful image approved with IPTC metadata and a C2PA credential, the record verifies and a tampered copy fails); receipts attested; no product text in the logs. The first try without the C2PA step approved the image with no credential, which the bundle caught. The model and document-reader servers' own startup was not re-verified (no new GPU load).
Rehearsal bundle: honest-product-imagery.zip (404 KB, 11 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route); production gets this tool when the branch merges.
- Measured on six made-up products with scenes drawn by one image model; real product photos and other generators were not tested.
- The alignment handles upright shots and small tilts; strong perspective or a product held at an angle skips the colour and logo checks.
- A generated hand or body part counts as a person, so it needs an identity (a synthetic performer can be enrolled as a fictional identity).
- The C2PA credential uses a development certificate; public validators show it as untrusted.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Finds the product units, the logo and label, and every person or human likeness in the reference photo and the AI image; compares the two side by side with the seller's facts, per category (size, colour, text, logo, warning, feature)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
- Pack text, step 1: finds the text regions on crops of the product (the document reader block)Docling 2.130 with the Heron layout modelMIT (Docling) + Apache-2.0 (weights)
- Pack text, step 2: reads each region (name, variant, claims, quantity, warnings)PaddleOCR-VL-1.6 (0.9B)Apache-2.0
- Alignment (NCC on edges over scale and a few degrees of rotation), colour (CIEDE2000 after white balance on the pack's neutral areas), logo match, unit count, changed pack area, the quantity and text comparisons, the verdict rules, the consent gate, XMP, the burned-in label and its OCR read-back, the C2PA credential and the signed record (no model; CPU)decosa-api imagery module (decosa_api/verticals/imagery)AGPL-3.0-or-later
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · pixels and the model, no document reader (1)
- not measured as a separate tier (the standard tier was measured): not measureddecosa-api docs/evals/honest-product-imagery.md
Standard · pixels, pack text and the model (hosted demo) (1)
- held-out test (thresholds frozen): planted misrepresentations flagged / faithful images flagged: 48 of 48 (44 refused, 4 for review) / 3 of 24decosa-api docs/evals/honest-product-imagery.md, measured on our server 2026-09-27, gateway route