Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: GPSR listing pack (72)

Run 27 Sep 2026 on our server, branch the pre-release branch, through a pre-release server on :8471: Qwen3.8-27B via the hosted gateway (receipted), Hy-MT2-7B on GPU0 (language-pack block). Runner: scripts/gpsr_eval.py. Results: docs/evals/gpsr-listing-pack/{dev,test-frozen,test-after-fix}.json.

Data

32 synthetic products from decosa_api/verticals/gpsr/synth.py (seed 988): 8 dev, 24 test, fixed before any run. Eight product kinds (toy, charger, kettle, radiator, ladder, candle, lamp, blender), manufacturers in and outside the EU (invented firms, .example addresses), EU responsible persons, and one planted problem per product at most: manufacturer e-mail or postal address left out, no responsible person for a non-EU maker, no model, no identifier (GTIN and batch), a label number that differs from the sheet, a label model that differs, or no warnings. Five target languages per pack (de, fr, es, it, pl). No real supplier documents were used.

Results

dev (8) test, frozen (24) test after one fix (24; 2 hit a gateway outage)
Missing Article 19 elements found 4 of 5, 0 false 15 of 19, 0 false 18 of 18, 0 false
Label-vs-sheet mismatches found 1 / 1, 0 false 5 / 5, 0 false 4 / 4, 0 false
Extracted fields right (model, GTIN, manufacturer name and e-mail) 29 / 30 87 / 88 80 / 81
Translations with no high flag 40 / 40 115 / 115 105 / 105
Median time per pack (5 languages; before the meaning check existed, superseded below) 6.7 s 6.0 s 6.8 s
  • The dev miss was the eval's own mistake: a dropped postal address was still printed on the label, so it was not missing. The generator was fixed before the test run.
  • The 4 frozen test misses were all "identifier" products: GTIN and batch left out, the model number kept. The pack counted the model as an identifier. Article 9(5) accepts "a type, batch or serial number or other element", so the pack now says check (not missing) when only the type is there, and the eval counts that as found. This fix was made after seeing the test results; the after-fix column is on the same products.
  • The one field miss (test): a product whose label model was planted to differ; the model took the label's model number. The label-vs-sheet flag caught it.
  • A dev product with no safety section: the model listed the specification "Maximum load: 150 kg" as safety information, so "warnings" came back found. Arguably right; not counted.

The default run, measured (30 Sep 2026, QA sweep BUG-1)

The table above timed packs before the meaning check existed. Since 28 Sep every pack runs it by default (a back-translation and one Qwen3.8 judgment per translated line), and the page still showed the old time and cost. scripts/gpsr_measure.py now measures the default run end to end through the HTTP API and writes docs/evals/gpsr-listing-pack/measure-default.json; the site copies that file (npm run sync-evals) and reads its speed, receipt count and cost from it (a site test ties them). Run on 30 Sep on production (api.decosa.ai, after the fixes below were deployed), with the 24 held-out test products at five languages (de, fr, es, it, pl) and the four page samples twice. A first run the same day on a pre-release server under other load measured 33.9 s / 57.9 s for the same set; production was faster, with about the same calls and tokens:

Default run (meaning check on) 24 test products, 5 languages
Median / 95th percentile wall time 18.7 s / 24.2 s
Model calls (receipts), median 61
Gateway LLM tokens per pack, mean (prompt / completion) 12,403 / 1,022
Cost at the gateway list price ($0.30 / $1.50 per M) $0.0053 per pack (about $0.53 per 100); the gateway charged 70% of that

The figures in this table are copied from the JSON for reading; the page takes them from the JSON. Most of the time and tokens are the meaning check: the same pack with "meaning": false makes one extraction call plus the translations (the QA sweep measured 5.0 s and about 1.2k tokens for five languages).

Meaning check in the pack: notes and the code cross-check (30 Sep 2026, QA sweep BUG-2)

The sweep found the meaning check's "check" verdicts turning most default runs amber. Two changes, both in the pack only:

  1. A meaning check is an advisory note on the line; it no longer changes the language's status or the verdict. Only an error makes the language "check". In the production measurement above, 17 of 536 translated lines came back error and 73 as notes (the earlier run: 19 and 71); with the old rule the notes alone would have turned more runs amber. The errors read in the earlier run were real: Polish "Piec" (bake) for "Burn", German "Heizer" (a stoker) for "heater", "top 2 steps" as "the last two steps".
  2. meaning.cross_check (cross.v1, opt-in, on in the pack) sets aside three kinds of issue when code sees the opposite: a negation issue on a double negative ("Do not use ... without supervision") whose negation count the translation and the back-translation keep, each negation on the same word; an "added" issue whose back-translation adds no content word ("Small parts." -> "Enthält kleine Teile"); a hedge issue whose hedge words are unchanged. An error whose issues are all set aside becomes a note; a note becomes ok; the set-aside issue stays in the segment with its reason. Replayed on the language pack's recorded meaning eval (scripts/meaning_crosscheck_eval.py, no model calls, results in docs/evals/gpsr-listing-pack/meaning-crosscheck.json): planted errors caught unchanged (dev 267 of 279, test 556 of 562; 0 lost), false flags on correct translations 52 to 50 of 192 on test (23 of 96 on dev, unchanged). The rules were written from the sweep's examples, but a first version (negation counts only) lost planted moved negations on this set ("with food, not on an empty stomach") and was tightened after seeing that, so this replay is not a held-out result. The gain is small on this set because its false flags are mostly other kinds; the change that matters for the verdict is (1).
  3. The language pack's code negation check (all callers) no longer reads English "no more than 6 hours" as a negation, and counts negative adjectives for "not suitable" (Romanian "nepotrivit", German "ungeeignet" and others) as negations: the sweep's Romanian and French "negation lost" flags came from these. The language pack tests pass unchanged.

Checkable properties (rehearsal/gpsr-listing-pack)

  1. The charger sample (UK manufacturer, no EU responsible person) comes back fail with rp_name, rp_postal and rp_electronic missing.
  2. The manufacturer's name, postal and electronic address, type, identifier and warnings are found.
  3. German and Polish safety text have no high flag.
  4. The signed record verifies at /record/verify.
  5. The kettle sample's label says 1.8 l where the sheet says 1.7 l: a label_sheet_mismatch flag, and picture missing.

Limits

  • Synthetic products written by the same author as the pack; label photos were tested on one rendered label.
  • Harmonisation legislation (toys, electrical equipment, cosmetics) is not checked.
  • Translations are checked for facts (numbers, units, codes, negations, names), not wording: the recorded German charger run has "Decken Sie den Ladegerät nicht ab" (wrong article), which no check catches.