Skip to content
decosa

80 · Healthcare · Compliance and trust · live

Label consistency across PI, SmPC, CCDS and carton

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)

  • Planted drifts caught31 / 31test splitn = 31held-out invented medicine: 8 numbers, 5 missing warnings, 4 missing contraindications, 5 storage, 9 translations (German, French); all triaged likely error
  • Quotes tight to the drift30 / 31test splitn = 31the overlapping quote is a phrase or sentence (at most 250 characters)
  • Logged twins triaged deliberate6 / 6test splitn = 6the same drift with a deviation-log line that explains it
  • False flags on the clean set3 / 24test splitn = 24all three from the translation number lock (German compound noun, decimal comma); 0 after a post-test fix, re-scored on the same model outputs
  • Triage agreement, Qwen3.835 / 35test splitn = 35deliberate or error on the test cases; 37 / 37 on dev
  • Triage agreement, blind Opus 5.5 (same prompt)35 / 35test splitn = 35Claude Code sub-agent, inputs only; 36 / 37 on dev
  • Planted drifts caught, dev31 / 31dev (tuned on)n = 31the medicine the prompts were tuned on (three rounds)

Dataset

Two invented medicines, each a seven-document set (CCDS, US PI, SmPC in English, German and French, leaflet, carton): Norvexa for dev, Pelmora held out and run once on the frozen pipeline. 31 drifts planted one at a time per medicine, plus 6 twins with an explaining deviation-log line.

Caveats

  • Same author wrote the documents, the plants, the prompts and the key; the drift types and conventions are the ones the prompt names.
  • Synthetic only: no real label set and no comparison with a labelling reviewer's findings.
  • Small n: 31 drifts and 6 twins in the test split; 31 / 31 still leaves a 95% lower bound near 89%.
  • The three test false flags were fixed after the test run; the held-out figure stays 3 / 24.
  • The triage comparison shows both models follow written conventions; it says nothing about unlogged differences a reviewer knows were approved.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
28 Sep 2026
Latency, this run
n/a
p50 over passed runs
2.3 s
Receipts
3
Model calls
n/a
Tokens
n/a
Cost per run
$0.001

Self-host verification

Verified on 27 Sep 2026: fresh clone into a clean directory, api image from docker/api/Dockerfile, compose with a named volume, direct route to the local Qwen3.8-27B, language pack and document reader, local signing; torn down after

The rehearsal bundle passed 7/7 in 13.8 s; every receipt signed. The language pack and reader were the running services, not built from the compose file here.

Rehearsal bundle: label-consistency-check.zip (8 KB, 7 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted numbers are from 6 production smoke runs after the merge: 3 on 27-28 Sep under load (18.2, 21.0, 24.3 s) and 3 on a quiet gateway on 28 Sep (2.1, 2.3, 2.3 s). Cost is the median at list price, model calls included (range $0.0011 to $0.0011).
  • Measured on synthetic label sets written by the same author as the prompts; not on a real company's labels or against a labelling reviewer's findings.
  • Full seven-document sets take minutes on the shared gateway (415 s measured); the Watch replay shows a recorded real run.
  • QRD headings are checked in English, German and French only; other EU languages get the number lock and meaning check.
  • Interactions, pregnancy sections, pharmacology and artwork layout are not compared.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Section splitter (QRD numbers, 21 CFR 201.57 numbers, heading words), pair plan, quote location at character offsets, number-with-unit comparison, QRD heading check, triage rules for translations, signed record (no model; CPU)decosa-api label check (decosa_api/verticals/labelcheck), on the language-pack block (decosa_api/lang), the document reader (decosa_api/docreader) and the signed record (07)AGPL-3.0-or-later
  • One call per section pair: the differences as JSON with an exact quote from each document; one triage call per document pair; the meaning judgment per translated paragraph; a full-page read of a scanned carton (image input)Qwen3.8-27B (NVFP4, vision tower on)Apache-2.0
  • Back-translation of each translated paragraph into English for the meaning check (the language-pack block): German, French and the other languages it servesHy-MT2-7BApache-2.0
  • Scanned or PDF cartons and labels: finds and orders the regions of each page (Docling Heron layout) and reads them (PaddleOCR-VL-1.6), with a page and box per lineDocument reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)Apache-2.0 (PaddleOCR-VL-1.6 weights, Heron layout weights); MIT (Docling)

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · text documents, meaning check off (1)
  • Planted drifts caught: not measured as a separate tierthe eval ran with the meaning check on; 2 of 9 translation drifts on the test split were found only by the meaning check
Standard · Qwen3.8-27B, Hy-MT2-7B and the document reader (hosted demo) (5)
  • Planted drifts caught, held-out synthetic set (numbers, missing warnings and contraindications, storage, translations): 31 / 31, all triaged likely error; 30 / 31 quoted to a phrase or sentencedecosa-api docs/evals/label-consistency-check.md, test split (Pelmora, run once)
  • Same drift with an explaining deviation-log entry, triaged likely deliberate: 6 / 6decosa-api docs/evals/label-consistency-check.md, test split
  • False flags on the clean held-out set: 3 of 24 flags (all from the translation number lock; 0 after a post-test fix, re-scored on the same model outputs)decosa-api docs/evals/label-consistency-check.md
  • Deliberate-or-error triage against a blind frontier judge (Claude Opus 5.5, same prompt), 72 cases: Qwen3.8 72 / 72, Opus 71 / 72; 35 / 35 each on the test casesdecosa-api docs/evals/label-consistency-check.md
  • Full seven-document set, hosted gateway route: 415 s under shared load, 43 Qwen3.8 calls, $0.0147 at list pricedecosa-api docs/evals/label-consistency-check.md
Best · the 30B translation model (not measured) (1)
  • Planted translation drifts caught: not measured yetnot run

How we measure · All tools