Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Label consistency check (80): eval

27 Sep 2026. On a pre-release build. Runners: scripts/labelcheck_eval.py (planted drifts), scripts/labelcheck_judge_eval.py (the "deliberate or error" comparison). Per-drift rows, flags and false flags: docs/evals/label-consistency-check/. Model calls: Qwen3.8-27B through the model gateway (every call receipted) and Hy-MT2-7B on the language pack for back-translation, measured under the shared gateway's load.

Sets

Two invented medicines, each a seven-document set: CCDS, US PI, EU SmPC in English, German and French, English package leaflet and carton text (decosa_api/verticals/labelcheck/data/). Every medicine, company, NDC and EU number is invented; the texts were written for this repo (AGPL-3.0-or-later with decosa-api). No real label is used, and no DailyMed text is paired with invented EU documents.

Split Medicine Role
dev Norvexa (tavorexin tablets) the samples; prompts, rules and conventions were written and tried on it (three rounds, v1 to v3)
test Pelmora (rilzopant capsules) written before the freeze, run once on the frozen pipeline (commit 8a7c012)

Each base set carries the usual deliberate differences: US controlled room temperature with °F, US lab units and lb, QRD standard statements, lay leaflet wording, a region-only indication and a wording change that the company's deviation log records. On top of the base set, 31 drifts are planted one at a time per medicine, and 6 "twins" repeat a drift together with a deviation-log line that explains it (the right triage is then likely_deliberate):

Type n per medicine Examples (dev)
number (doses, strengths, intervals, percentages) 8 "first 7 days" to "14 days" in the US PI; 100 mg to 10 mg on the carton
missing warning 5 a warning paragraph removed from the US PI or the SmPC
missing contraindication 4 a contraindication removed from the SmPC or the leaflet
storage 5 carton 25°C where the SmPC says 30°C; a "do not freeze" dropped
translation (German, French SmPC) 9 a dropped "nicht", "4 weeks" as "2 semaines", a paragraph missing

Scoring

  • caught: a flag in a pair that includes the edited document, whose quote on either side overlaps the planted span.
  • as error: caught and triaged likely_error (twins: likely_deliberate is the right answer).
  • tight span: the overlapping quote is at most 250 characters (a phrase or a sentence, not a section).
  • false flags: flags triaged likely_error in the base run that overlap no known base difference, plus new likely_error flags in a variant's edited pairs that overlap neither the plant nor anything in the base run.

Results: planted drifts

Split Caught As error Tight span Twins as deliberate Base-set flags False flags (base) Base differences misjudged New false flags in variants
dev v1 30/31 24/31 28/31 6/6 29 1 0 20
dev v2 30/31 30/31 30/31 6/6 29 4 1 3
dev v3 (frozen) 31/31 31/31 30/31 6/6 29 3 0 2
test (run once) 31/31 31/31 30/31 6/6 24 3 0 1

Per type on the test split: numbers 8/8, missing warnings 5/5, missing contraindications 4/4, storage 5/5, translations 9/9, all triaged likely_error; tight spans 30/31 (the loose one is a German "should be measured" change quoted as its English phrase only, with no German side).

False flags on the test split. All three base false flags came from the translation number lock, not the model: "4 migraine days" against "4 Migränetagen" (the unit is the end of a German compound noun) and "mL/min/1.73" against "ml/min/1,73" (decimal comma) in German and French. Post-test fix (commit daa7243, with regression tests): the label check now accepts a unit at the end of a compound noun and an identifier written with a decimal comma. Re-scored with the same cached model outputs, the test base set has 0 false flags (pelmora-postfix.json) and every translation drift is still caught. This is a fix made after seeing the test split; the held-out figure above stays 3.

The dev false flags (3) are model flags on real wording changes the base set carries: "Do not start" as "Avoid use", "must use contraception" as "Advise ... to use", and a leaflet that says "talk to your doctor" where the SmPC says "must not be initiated". A reviewer may well want these; they count as false here because they were not planted.

Results: "deliberate or error" against a frontier judge

72 cases (37 dev, 35 test): every caught planted drift in a same-language pair (not translations, which are triaged by rule), its twin with the explaining deviation-log line, and the base differences with one right answer (regional conventions, deviation-log entries). Each case holds only the inputs the production triage sees: both roles, the deviation log and the difference with both quotes. The frontier judge is a blind Claude Code sub-agent (Opus 5.5) that read only the cases and judge_brief.md (the production triage system prompt, one case at a time), never the key; we scored its verdicts.

Judge All Dev Test Numbers Missing warnings Missing contraindications Storage Twins Base differences
Qwen3.8-27B (production prompt v2) 72/72 37/37 35/35 16/16 10/10 8/8 10/10 12/12 16/16
Opus 5.5, blind, same prompt 71/72 36/37 35/35 16/16 10/10 8/8 10/10 12/12 15/16

The one disagreement is a dev case where the key is itself soft: the US PI's "plaque psoriasis" without "chronic", keyed likely_deliberate because it sits in the logged US-only indication difference; Opus said unclear ("no log entry says so"). Qwen cited the right deviation-log line in 12/12 twins.

A first judge run gave Opus the v1 prompt by mistake (the brief was written before the v2 prompt); it scored 68/72, with three of the four misses on storage lines missing from the carton, which v2's carton rule settles. It is kept as judge_opus_v1prompt.jsonl and not used above.

What this shows and what it doesn't: on these synthetic cases the triage is a rule-following task that both models do nearly perfectly once the regional conventions and the deviation log are written out. It says nothing about the hard real cases (an unlogged difference that a reviewer knows was approved), which synthetic data can't contain.

Time and cost (hosted gateway route, shared load, load average about 65)

Run Documents Qwen3.8 calls Prompt / completion tokens LLM cost at list price Wall time
norvexa-drift (the full set) 7 43 (+57 back-translation and meaning attestations) 30,512 / 3,667 $0.0147 415 s
norvexa-quick 2 (CCDS, US PI) 8 5,825 / 1,638 $0.0042 148 s
norvexa-scan (carton as an image) 2 3 (+ the document reader) 2,915 / 292 $0.0013 107 s
smoke (two short documents) 2 3 2,321 total $0.0011 81 s

Times are dominated by queueing on the shared gateway; the self-hosted rehearsal (4 documents, direct route to the local Qwen3.8) took 14 s.

Rehearsal expectations (rehearsal/label-consistency-check/expected.json)

  1. The set is reported as drifted.
  2. The US PI against the CCDS has at least two likely errors in its warnings (the 6-month interval, the missing warning).
  3. The dropped German negation is flagged in the translation.
  4. The logged US-only indication difference is not called an error.
  5. The German SmPC's QRD headings are all right.
  6. Every model call has a signed receipt, and the signed record verifies.

Self-host check: a fresh clone (8a7c012) built from docker/api/Dockerfile, compose with a named volume, direct route to the local Qwen3.8-27B, the running language pack (:8491) and document reader (:8497): 7/7 checks passed in 13.8 s; torn down.

Limits

  • Same author wrote the documents, the plants, the prompts and the key; the test medicine is new text, but the drift types and conventions are the ones the prompt names. Real label sets are longer and messier.
  • Synthetic only: no real CCDS or EU label was used, and no comparison with a labelling reviewer's findings.
  • English, German and French only in the QRD heading pack; other EU languages get the number lock and the meaning check but no heading check.
  • Interactions, pregnancy sections, pharmacology and artwork layout are not compared.
  • Test split: n = 31 drifts and 6 twins; 100% on 31 cases still leaves a 95% lower bound near 89%.