Eval: GPSR listing pack (72)
Run 27 Sep 2026 on our server, branch the pre-release branch, through a pre-release server on :8471: Qwen3.8-27B via the hosted gateway
(receipted), Hy-MT2-7B on GPU0 (language-pack block). Runner: scripts/gpsr_eval.py. Results:
docs/evals/gpsr-listing-pack/{dev,test-frozen,test-after-fix}.json.
Data
32 synthetic products from decosa_api/verticals/gpsr/synth.py (seed 988): 8 dev, 24 test, fixed before any run. Eight
product kinds (toy, charger, kettle, radiator, ladder, candle, lamp, blender), manufacturers in and outside the EU (invented
firms, .example addresses), EU responsible persons, and one planted problem per product at most: manufacturer e-mail or
postal address left out, no responsible person for a non-EU maker, no model, no identifier (GTIN and batch), a label number
that differs from the sheet, a label model that differs, or no warnings. Five target languages per pack (de, fr, es, it,
pl). No real supplier documents were used.
Results
| dev (8) | test, frozen (24) | test after one fix (24; 2 hit a gateway outage) | |
|---|---|---|---|
| Missing Article 19 elements found | 4 of 5, 0 false | 15 of 19, 0 false | 18 of 18, 0 false |
| Label-vs-sheet mismatches found | 1 / 1, 0 false | 5 / 5, 0 false | 4 / 4, 0 false |
| Extracted fields right (model, GTIN, manufacturer name and e-mail) | 29 / 30 | 87 / 88 | 80 / 81 |
| Translations with no high flag | 40 / 40 | 115 / 115 | 105 / 105 |
| Median time per pack (5 languages; before the meaning check existed, superseded below) | 6.7 s | 6.0 s | 6.8 s |
- The dev miss was the eval's own mistake: a dropped postal address was still printed on the label, so it was not missing. The generator was fixed before the test run.
- The 4 frozen test misses were all "identifier" products: GTIN and batch left out, the model number kept. The pack counted the model as an identifier. Article 9(5) accepts "a type, batch or serial number or other element", so the pack now says check (not missing) when only the type is there, and the eval counts that as found. This fix was made after seeing the test results; the after-fix column is on the same products.
- The one field miss (test): a product whose label model was planted to differ; the model took the label's model number. The label-vs-sheet flag caught it.
- A dev product with no safety section: the model listed the specification "Maximum load: 150 kg" as safety information, so "warnings" came back found. Arguably right; not counted.
The default run, measured (30 Sep 2026, QA sweep BUG-1)
The table above timed packs before the meaning check existed. Since 28 Sep every pack runs it by default (a back-translation
and one Qwen3.8 judgment per translated line), and the page still showed the old time and cost. scripts/gpsr_measure.py
now measures the default run end to end through the HTTP API and writes docs/evals/gpsr-listing-pack/measure-default.json;
the site copies that file (npm run sync-evals) and reads its speed, receipt count and cost from it (a site test ties them).
Run on 30 Sep on production (api.decosa.ai, after the fixes below were deployed), with the 24 held-out test products at five
languages (de, fr, es, it, pl) and the four page samples twice. A first run the same day on a pre-release server under other load
measured 33.9 s / 57.9 s for the same set; production was faster, with about the same calls and tokens:
| Default run (meaning check on) | 24 test products, 5 languages |
|---|---|
| Median / 95th percentile wall time | 18.7 s / 24.2 s |
| Model calls (receipts), median | 61 |
| Gateway LLM tokens per pack, mean (prompt / completion) | 12,403 / 1,022 |
| Cost at the gateway list price ($0.30 / $1.50 per M) | $0.0053 per pack (about $0.53 per 100); the gateway charged 70% of that |
The figures in this table are copied from the JSON for reading; the page takes them from the JSON. Most of the time and
tokens are the meaning check: the same pack with "meaning": false makes one extraction call plus the translations (the QA
sweep measured 5.0 s and about 1.2k tokens for five languages).
Meaning check in the pack: notes and the code cross-check (30 Sep 2026, QA sweep BUG-2)
The sweep found the meaning check's "check" verdicts turning most default runs amber. Two changes, both in the pack only:
- A meaning check is an advisory note on the line; it no longer changes the language's status or the verdict. Only an error makes the language "check". In the production measurement above, 17 of 536 translated lines came back error and 73 as notes (the earlier run: 19 and 71); with the old rule the notes alone would have turned more runs amber. The errors read in the earlier run were real: Polish "Piec" (bake) for "Burn", German "Heizer" (a stoker) for "heater", "top 2 steps" as "the last two steps".
meaning.cross_check(cross.v1, opt-in, on in the pack) sets aside three kinds of issue when code sees the opposite: a negation issue on a double negative ("Do not use ... without supervision") whose negation count the translation and the back-translation keep, each negation on the same word; an "added" issue whose back-translation adds no content word ("Small parts." -> "Enthält kleine Teile"); a hedge issue whose hedge words are unchanged. An error whose issues are all set aside becomes a note; a note becomes ok; the set-aside issue stays in the segment with its reason. Replayed on the language pack's recorded meaning eval (scripts/meaning_crosscheck_eval.py, no model calls, results indocs/evals/gpsr-listing-pack/meaning-crosscheck.json): planted errors caught unchanged (dev 267 of 279, test 556 of 562; 0 lost), false flags on correct translations 52 to 50 of 192 on test (23 of 96 on dev, unchanged). The rules were written from the sweep's examples, but a first version (negation counts only) lost planted moved negations on this set ("with food, not on an empty stomach") and was tightened after seeing that, so this replay is not a held-out result. The gain is small on this set because its false flags are mostly other kinds; the change that matters for the verdict is (1).- The language pack's code negation check (all callers) no longer reads English "no more than 6 hours" as a negation, and counts negative adjectives for "not suitable" (Romanian "nepotrivit", German "ungeeignet" and others) as negations: the sweep's Romanian and French "negation lost" flags came from these. The language pack tests pass unchanged.
Checkable properties (rehearsal/gpsr-listing-pack)
- The charger sample (UK manufacturer, no EU responsible person) comes back
failwithrp_name,rp_postalandrp_electronicmissing. - The manufacturer's name, postal and electronic address, type, identifier and warnings are found.
- German and Polish safety text have no high flag.
- The signed record verifies at
/record/verify. - The kettle sample's label says 1.8 l where the sheet says 1.7 l: a
label_sheet_mismatchflag, andpicturemissing.
Limits
- Synthetic products written by the same author as the pack; label photos were tested on one rendered label.
- Harmonisation legislation (toys, electrical equipment, cosmetics) is not checked.
- Translations are checked for facts (numbers, units, codes, negations, names), not wording: the recorded German charger run has "Decken Sie den Ladegerät nicht ab" (wrong article), which no check catches.