80 · Healthcare · Compliance and trust · live
Label consistency across PI, SmPC, CCDS and carton
Eval results
Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)
- Planted drifts caught31 / 31test splitn = 31held-out invented medicine: 8 numbers, 5 missing warnings, 4 missing contraindications, 5 storage, 9 translations (German, French); all triaged likely error
- Quotes tight to the drift30 / 31test splitn = 31the overlapping quote is a phrase or sentence (at most 250 characters)
- Logged twins triaged deliberate6 / 6test splitn = 6the same drift with a deviation-log line that explains it
- False flags on the clean set3 / 24test splitn = 24all three from the translation number lock (German compound noun, decimal comma); 0 after a post-test fix, re-scored on the same model outputs
- Triage agreement, Qwen3.835 / 35test splitn = 35deliberate or error on the test cases; 37 / 37 on dev
- Triage agreement, blind Opus 5.5 (same prompt)35 / 35test splitn = 35Claude Code sub-agent, inputs only; 36 / 37 on dev
- Planted drifts caught, dev31 / 31dev (tuned on)n = 31the medicine the prompts were tuned on (three rounds)
Dataset
Two invented medicines, each a seven-document set (CCDS, US PI, SmPC in English, German and French, leaflet, carton): Norvexa for dev, Pelmora held out and run once on the frozen pipeline. 31 drifts planted one at a time per medicine, plus 6 twins with an explaining deviation-log line.
Caveats
- Same author wrote the documents, the plants, the prompts and the key; the drift types and conventions are the ones the prompt names.
- Synthetic only: no real label set and no comparison with a labelling reviewer's findings.
- Small n: 31 drifts and 6 twins in the test split; 31 / 31 still leaves a 95% lower bound near 89%.
- The three test false flags were fixed after the test run; the held-out figure stays 3 / 24.
- The triage comparison shows both models follow written conventions; it says nothing about unlogged differences a reviewer knows were approved.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 28 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 2.3 s
- Receipts
- 3
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.001
Self-host verification
Verified on 27 Sep 2026: fresh clone into a clean directory, api image from docker/api/Dockerfile, compose with a named volume, direct route to the local Qwen3.8-27B, language pack and document reader, local signing; torn down after
The rehearsal bundle passed 7/7 in 13.8 s; every receipt signed. The language pack and reader were the running services, not built from the compose file here.
Rehearsal bundle: label-consistency-check.zip (8 KB, 7 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted numbers are from 6 production smoke runs after the merge: 3 on 27-28 Sep under load (18.2, 21.0, 24.3 s) and 3 on a quiet gateway on 28 Sep (2.1, 2.3, 2.3 s). Cost is the median at list price, model calls included (range $0.0011 to $0.0011).
- Measured on synthetic label sets written by the same author as the prompts; not on a real company's labels or against a labelling reviewer's findings.
- Full seven-document sets take minutes on the shared gateway (415 s measured); the Watch replay shows a recorded real run.
- QRD headings are checked in English, German and French only; other EU languages get the number lock and meaning check.
- Interactions, pregnancy sections, pharmacology and artwork layout are not compared.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Section splitter (QRD numbers, 21 CFR 201.57 numbers, heading words), pair plan, quote location at character offsets, number-with-unit comparison, QRD heading check, triage rules for translations, signed record (no model; CPU)decosa-api label check (decosa_api/verticals/labelcheck), on the language-pack block (decosa_api/lang), the document reader (decosa_api/docreader) and the signed record (07)AGPL-3.0-or-later
- One call per section pair: the differences as JSON with an exact quote from each document; one triage call per document pair; the meaning judgment per translated paragraph; a full-page read of a scanned carton (image input)Qwen3.8-27B (NVFP4, vision tower on)Apache-2.0
- Back-translation of each translated paragraph into English for the meaning check (the language-pack block): German, French and the other languages it servesHy-MT2-7BApache-2.0
- Scanned or PDF cartons and labels: finds and orders the regions of each page (Docling Heron layout) and reads them (PaddleOCR-VL-1.6), with a page and box per lineDocument reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)Apache-2.0 (PaddleOCR-VL-1.6 weights, Heron layout weights); MIT (Docling)
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · text documents, meaning check off (1)
- Planted drifts caught: not measured as a separate tierthe eval ran with the meaning check on; 2 of 9 translation drifts on the test split were found only by the meaning check
Standard · Qwen3.8-27B, Hy-MT2-7B and the document reader (hosted demo) (5)
- Planted drifts caught, held-out synthetic set (numbers, missing warnings and contraindications, storage, translations): 31 / 31, all triaged likely error; 30 / 31 quoted to a phrase or sentencedecosa-api docs/evals/label-consistency-check.md, test split (Pelmora, run once)
- Same drift with an explaining deviation-log entry, triaged likely deliberate: 6 / 6decosa-api docs/evals/label-consistency-check.md, test split
- False flags on the clean held-out set: 3 of 24 flags (all from the translation number lock; 0 after a post-test fix, re-scored on the same model outputs)decosa-api docs/evals/label-consistency-check.md
- Deliberate-or-error triage against a blind frontier judge (Claude Opus 5.5, same prompt), 72 cases: Qwen3.8 72 / 72, Opus 71 / 72; 35 / 35 each on the test casesdecosa-api docs/evals/label-consistency-check.md
- Full seven-document set, hosted gateway route: 415 s under shared load, 43 Qwen3.8 calls, $0.0147 at list pricedecosa-api docs/evals/label-consistency-check.md
Best · the 30B translation model (not measured) (1)
- Planted translation drifts caught: not measured yetnot run