Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: green-claims substantiation check (use case 45)

Run 25 Sep 2026 on our server, gateway route, Qwen3.8-27B (the same gateway model as the promo pre-check), temperature 0, thinking off, under normal shared load. Script: scripts/green_eval.py; cases: scripts/green_cases.py and scripts/green_cases_test2.py; results: docs/evals/green-claims-check/*.json (every row, verdict, rule, quote and rewrite, plus token usage). Every model call was receipted (receipt count = call count in each results file).

Data

All synthetic, written for this eval: fictional brands, certifiers, schemes and consultancies; no real company, product or certificate. The only real names are EU law and the EU Ecolabel and EU energy label as legal instruments. Each piece is marketing copy (pack, web page or ad script) plus an evidence file (LCA summaries, certificates, offset contracts, plans, internal policies). Each sentence is labelled with its expected verdict, the rule ids that should fire, and, for claims the evidence backs, a short piece of the evidence text the cited span should contain.

split pieces sentences planted violations clean environmental claims not environmental compliant pieces
dev (the demo samples) 3 26 11 10 5 1
test 9 69 30 26 13 3
test2 6 44 19 17 8 2

Planted per rule (test / test2): Annex I 2a own-brand labels 4 / 3; 4a generic claims 7 / 4; 4b whole-product claims 4 / 2; 4c offset-based neutrality 6 / 4; 10a legal requirement as a feature 3 / 3; Article 6(2)(d) future targets without a verified plan 4 / 2; plus 2 / 1 specific claims the evidence contradicts (Article 6(1)).

What was held out, and what was not

  • v1. Prompts and rules were written and shaped on the dev split only. The test split was written and committed (66fb380) before any model run, and v1 was frozen (76415f5) before it was scored once.
  • v2 (2866055). After reading the v1 test errors I changed the prompts and rules: the scope of Annex I 4b (not for future targets; a named part or site is not "the whole product"; coverage judged against what the claim names), the 10a definition (a measured figure is not a legal requirement), take-back services count as environmental, and the 4c finding only when offsets are shown or unclear. So v2 and later numbers on the test split are not held out. To get a clean number for v2 I wrote a new split, test2, after v2 was frozen and before any run on it.
  • v2.1 withholds a model rewrite that still contains a generic or neutrality term (display only; no verdict changes).
  • v2.2 adds a word-list floor for well-known 10a cases (CFCs, RoHS, phosphates in laundry or dishwasher detergent), after five repeat runs of the Fernhollow demo sample missed "phosphate-free" every time (v2.1-repeat-*.json).
  • v2.3 (the shipped version) tells the classifier that a supplier certificate named as the source of a figure is not a label, after a v2.2 test2 rerun flagged one. v2.3 numbers on test2 are therefore not held out either.
  • One label fix after v1: the expected evidence text for the Tidewell steel claim was not in its evidence file; it changes only the grounding score.

Results

The gateway is not fully deterministic at temperature 0 (shared batching), so the same code can differ run to run; the test2 reruns below show how much.

v1 on test (held out) v2 on test2 (held out, first run) v2.2 on test2 (3 reruns) v2.3 on test2 (shipped; not held out) word list only, test2
Verdict accuracy (4 classes, all sentences) 58/69 (84%) 43/44 (98%) 41/44 each 43/44 12/44 (27%)
Planted violations flagged (any flag) 30/30 19/19 18/19 each 19/19 4/19
False alarms, clean claims and non-claims 5/39 0/25 1/25 each 0/25 0/25
False alarms on the compliant pieces 3/20 0/14 1/14 each 0/14 0/14
Substantiated claims whose cited span holds the expected evidence 14/19 (first span only) 16/16 15/15 each 16/16 n/a
Model rewrites that still used a generic or neutrality term (withheld since v2.1) 1/20 2/9 2/9 to 3/11 3/11 n/a

v2.3 on the (no longer held-out) test split: 61/69 (88%), 30/30 flagged, 5/39 false alarms, 18/18 grounded. On dev: 25/26.

Per rule, precision and recall (true positives, false positives, misses):

Rule v1 on test (held out) v2 on test2 (held out) v2.2 on test2, 3 reruns (summed)
Annex I 2a, label without a certification scheme P 1.00, R 1.00 (4, 0, 0) P 1.00, R 1.00 (3, 0, 0) P 0.69, R 1.00 (9, 4, 0)
Annex I 4a, generic claim P 1.00, R 1.00 (7, 0, 0) P 1.00, R 1.00 (4, 0, 0) P 1.00, R 1.00 (12, 0, 0)
Annex I 4b, whole product when only a part qualifies P 0.40, R 0.50 (2, 3, 2) P 1.00, R 1.00 (2, 0, 0) P 1.00, R 1.00 (6, 0, 0)
Annex I 4c, offset-based neutrality P 0.86, R 1.00 (6, 1, 0) P 1.00, R 1.00 (4, 0, 0) P 1.00, R 1.00 (12, 0, 0)
Annex I 10a, legal requirement as a feature P 0.75, R 1.00 (3, 1, 0) P 1.00, R 1.00 (3, 0, 0) P 1.00, R 0.67 (6, 0, 3)
Article 6(2)(d), future claim without a verified plan P 1.00, R 1.00 (4, 0, 0) P 1.00, R 1.00 (2, 0, 0) P 1.00, R 1.00 (6, 0, 0)

UK mode (CMA Green Claims Code; verdicts only, the rule ids differ): v2 on test2 43/44, 19/19 flagged, 0/25 false alarms; v2.2 41/44; v2.3 42/44 (it missed the toy-safety 10a claim in that run).

Cost and speed (v2 on test2): about 19 model calls per piece (1 per sentence, 3 per environmental claim), 1,640 generated and 17,800 prompt tokens per piece, about $0.008 per piece at the gateway list price ($0.30 / $1.50 per million tokens), median 6.3 s per piece with 6 workers when called directly. Through the HTTP route the 10-sentence smoke sample took 37 s under load, and the three Watch recordings took 6 to 8 s each.

Where it fails

  • Held-out v2 misses: "Charging the battery fully uses about 0.5 kWh of electricity" is called not an environmental claim in every run (energy use stated as a plain spec is borderline; the label is ours). In the reruns, "our toys meet the EU Toy Safety Directive, which sets them apart" was missed (10a depends on the model knowing the law applies to every toy), and a frame "with 70% recycled content, according to our supplier's certificate" was read as a label and banned under 2a (fixed in v2.3, not held out).
  • v1 errors (all on test): 4b fired on part-specific claims ("our bars are made in a solar-powered workshop", a 2027 packaging target) and on an EU Ecolabel claim, and missed two "entire business" claims the evidence contradicted outright; 10a fired on a measured VOC figure; a take-back service was not seen as an environmental claim.
  • Grounding is strict. "Substantiated" needs the grounding judge (vertical 17) to say supported with a cited span. It sometimes calls a fair paraphrase partial (a liner "not recyclable" vs "not recyclable in household collection", "avoids about 3 kg" vs a net saving of 3 kg), which turns a clean claim into "needs substantiation". That is the safe direction for this job, and it is where the remaining false alarms on test come from.
  • Rewrites are drafted by the model. On 2 of 9 held-out rewrites it kept "carbon neutral" while naming the offsets, which is still banned under 4c; those are now withheld and the report says to rewrite by hand.

Limits of this eval

  • Small and synthetic: 44 held-out sentences in 6 pieces, 2 to 4 planted per rule. A rule at 100% here can still miss in real copy. The pieces are short and each planted claim is fairly clear; real copy mixes claims in one sentence.
  • The cases and the prompts were written by the same author (Claude, for Decosa), so they may share blind spots.
  • 10a depends on the model knowing which EU rules apply to a whole product category. It was tested on well-known cases (CFCs, phosphates in laundry detergent, RoHS, toy safety); less known ones are likely to be missed.
  • Text only: labels drawn as artwork, colours and imagery are not read.
  • No real regulator decision or court case was used as ground truth. The labels are our reading of the Directive text.