45 · Compliance and trust · live
Green-claims substantiation check
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Verdict accuracy, v2 on test2 (held out, first run)43/44 (98%)held outn = 44v2.2 reruns on test2: 41/44 each.
- Planted violations flagged, v2 on test2 (held out)19/19held outn = 19v2.2 reruns: 18/19 each.
- False alarms on clean claims and non-claims, v2 on test2 (held out)0/25held outn = 25v2.2 reruns: 1/25 each.
- Verdict accuracy, v1 on test (held out)58/69 (84%)held outn = 6930/30 violations flagged, 5/39 false alarms; v2 was then changed after reading these errors.
- Verdict accuracy, word list only, test212/44 (27%)held outn = 44Baseline without the model: 4/19 violations flagged.
- Model rewrites still using a generic or neutrality term, v2 on test22/9held outn = 9Withheld from display since v2.1.
Dataset
18 synthetic marketing pieces with evidence files for fictional brands (dev 3, test 9, test2 6), each sentence labelled with the expected verdict and EU Empowering Consumers Directive rule.
Caveats
- Small and synthetic: 44 held-out sentences in 6 pieces, 2 to 4 planted per rule.
- The same author (Claude, for Decosa) wrote the cases and the prompts, so they may share blind spots.
- Prompts were changed after reading the v1 test errors, so v2 and later numbers on the test split are not held out; the shipped v2.3 on test2 is not held out either.
- No regulator decision or court case was used as ground truth; labels are our reading of the Directive.
- Text only: labels drawn as artwork, colours and imagery are not read.
- The gateway is not fully deterministic at temperature 0; reruns differ by a sentence or two.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 8.6 s
- Receipts
- 26
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.013
Self-host verification
Verified on 25 Sep 2026: fresh clone, compose up, sample against local model servers
A fresh clone of a decosa-api pre-release build (not yet merged to main), the api image built from it, the compose file from this prompt, then its smoke steps against the already-running local Qwen3.8-27B vLLM on the direct route. All three samples ran (4.5-6 s each): Fernhollow banned_claims with 5 bans, Quillbrook nothing_flagged, Northwick (UK) high_risk; the report verified and a changed status failed. Model-server startup itself not re-verified.
Rehearsal bundle: green-claims-check.zip (3 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Text only: artwork, colours and label images are not read. Describe a label in words to have it checked.
- Checks the Directive's text, not national transposing laws, and not the durability and repair bans (23d to 23j).
- The eval is small and synthetic, and the gateway is not fully deterministic: the same piece can get a different verdict on a borderline claim from run to run. A person reviews every finding.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Rulepack, sentences, evidence spans, verdicts, claim table and signed report (no model; CPU)decosa-api green module (decosa_api/verticals/green) on the promo pre-check engine, with the grounding module (decosa_api/verticals/grounding)AGPL-3.0-or-later
- Claim typing, grounding judge, evidence reader and rewriteQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Standard · one GPU for the model (hosted demo) (6)
- Held-out test2 (6 synthetic pieces, 44 sentences): verdicts right, first run / three reruns: 43/44 / 41/44 eachdocs/evals/green-claims-check.md (v2 and v2.2 result files), 25 Sep 2026
- Planted violations flagged, held out (test2 / v1 on test): 19/19 (18/19 in reruns) / 30/30docs/evals/green-claims-check.md
- False alarms on clean claims and non-claims, held out (test2 / v1 on test): 0/25 (1/25 in reruns) / 5/39docs/evals/green-claims-check.md
- Substantiated claims whose cited span holds the expected evidence (test2): 16/16docs/evals/green-claims-check.md
- Per rule on test2 (2a, 4a, 4b, 4c, 10a, 6(2)(d)), first run: precision and recall 1.00 each, on 3, 4, 2, 4, 3 and 2 planted claims; reruns: 10a recall 0.67, 2a precision 0.69docs/evals/green-claims-check.md
- Word list alone (no model), test2: 12/44 verdicts, 4/19 violations flaggeddocs/evals/green-claims-check/baseline-test2-eu.json