Label consistency check (80): eval
27 Sep 2026. On a pre-release build. Runners: scripts/labelcheck_eval.py (planted drifts),
scripts/labelcheck_judge_eval.py (the "deliberate or error" comparison). Per-drift rows, flags and false flags:
docs/evals/label-consistency-check/. Model calls: Qwen3.8-27B through the model gateway (every call receipted) and
Hy-MT2-7B on the language pack for back-translation, measured under the shared gateway's load.
Sets
Two invented medicines, each a seven-document set: CCDS, US PI, EU SmPC in English, German and French, English package
leaflet and carton text (decosa_api/verticals/labelcheck/data/). Every medicine, company, NDC and EU number is
invented; the texts were written for this repo (AGPL-3.0-or-later with decosa-api). No real label is used, and no DailyMed text is
paired with invented EU documents.
| Split | Medicine | Role |
|---|---|---|
| dev | Norvexa (tavorexin tablets) | the samples; prompts, rules and conventions were written and tried on it (three rounds, v1 to v3) |
| test | Pelmora (rilzopant capsules) | written before the freeze, run once on the frozen pipeline (commit 8a7c012) |
Each base set carries the usual deliberate differences: US controlled room temperature with °F, US lab units and lb, QRD standard statements, lay leaflet wording, a region-only indication and a wording change that the company's deviation log records. On top of the base set, 31 drifts are planted one at a time per medicine, and 6 "twins" repeat a drift together with a deviation-log line that explains it (the right triage is then likely_deliberate):
| Type | n per medicine | Examples (dev) |
|---|---|---|
| number (doses, strengths, intervals, percentages) | 8 | "first 7 days" to "14 days" in the US PI; 100 mg to 10 mg on the carton |
| missing warning | 5 | a warning paragraph removed from the US PI or the SmPC |
| missing contraindication | 4 | a contraindication removed from the SmPC or the leaflet |
| storage | 5 | carton 25°C where the SmPC says 30°C; a "do not freeze" dropped |
| translation (German, French SmPC) | 9 | a dropped "nicht", "4 weeks" as "2 semaines", a paragraph missing |
Scoring
- caught: a flag in a pair that includes the edited document, whose quote on either side overlaps the planted span.
- as error: caught and triaged likely_error (twins: likely_deliberate is the right answer).
- tight span: the overlapping quote is at most 250 characters (a phrase or a sentence, not a section).
- false flags: flags triaged likely_error in the base run that overlap no known base difference, plus new likely_error flags in a variant's edited pairs that overlap neither the plant nor anything in the base run.
Results: planted drifts
| Split | Caught | As error | Tight span | Twins as deliberate | Base-set flags | False flags (base) | Base differences misjudged | New false flags in variants |
|---|---|---|---|---|---|---|---|---|
| dev v1 | 30/31 | 24/31 | 28/31 | 6/6 | 29 | 1 | 0 | 20 |
| dev v2 | 30/31 | 30/31 | 30/31 | 6/6 | 29 | 4 | 1 | 3 |
| dev v3 (frozen) | 31/31 | 31/31 | 30/31 | 6/6 | 29 | 3 | 0 | 2 |
| test (run once) | 31/31 | 31/31 | 30/31 | 6/6 | 24 | 3 | 0 | 1 |
Per type on the test split: numbers 8/8, missing warnings 5/5, missing contraindications 4/4, storage 5/5, translations 9/9, all triaged likely_error; tight spans 30/31 (the loose one is a German "should be measured" change quoted as its English phrase only, with no German side).
False flags on the test split. All three base false flags came from the translation number lock, not the model: "4
migraine days" against "4 Migränetagen" (the unit is the end of a German compound noun) and "mL/min/1.73" against
"ml/min/1,73" (decimal comma) in German and French. Post-test fix (commit daa7243, with regression tests): the label
check now accepts a unit at the end of a compound noun and an identifier written with a decimal comma. Re-scored with the
same cached model outputs, the test base set has 0 false flags (pelmora-postfix.json) and every translation drift is still
caught. This is a fix made after seeing the test split; the held-out figure above stays 3.
The dev false flags (3) are model flags on real wording changes the base set carries: "Do not start" as "Avoid use", "must use contraception" as "Advise ... to use", and a leaflet that says "talk to your doctor" where the SmPC says "must not be initiated". A reviewer may well want these; they count as false here because they were not planted.
Results: "deliberate or error" against a frontier judge
72 cases (37 dev, 35 test): every caught planted drift in a same-language pair (not translations, which are triaged by rule),
its twin with the explaining deviation-log line, and the base differences with one right answer (regional conventions,
deviation-log entries). Each case holds only the inputs the production triage sees: both roles, the deviation log and the
difference with both quotes. The frontier judge is a blind Claude Code sub-agent (Opus 5.5) that read only the cases
and judge_brief.md (the production triage system prompt, one case at a time), never the key; we scored its verdicts.
| Judge | All | Dev | Test | Numbers | Missing warnings | Missing contraindications | Storage | Twins | Base differences |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3.8-27B (production prompt v2) | 72/72 | 37/37 | 35/35 | 16/16 | 10/10 | 8/8 | 10/10 | 12/12 | 16/16 |
| Opus 5.5, blind, same prompt | 71/72 | 36/37 | 35/35 | 16/16 | 10/10 | 8/8 | 10/10 | 12/12 | 15/16 |
The one disagreement is a dev case where the key is itself soft: the US PI's "plaque psoriasis" without "chronic", keyed likely_deliberate because it sits in the logged US-only indication difference; Opus said unclear ("no log entry says so"). Qwen cited the right deviation-log line in 12/12 twins.
A first judge run gave Opus the v1 prompt by mistake (the brief was written before the v2 prompt); it scored 68/72, with
three of the four misses on storage lines missing from the carton, which v2's carton rule settles. It is kept as
judge_opus_v1prompt.jsonl and not used above.
What this shows and what it doesn't: on these synthetic cases the triage is a rule-following task that both models do nearly perfectly once the regional conventions and the deviation log are written out. It says nothing about the hard real cases (an unlogged difference that a reviewer knows was approved), which synthetic data can't contain.
Time and cost (hosted gateway route, shared load, load average about 65)
| Run | Documents | Qwen3.8 calls | Prompt / completion tokens | LLM cost at list price | Wall time |
|---|---|---|---|---|---|
| norvexa-drift (the full set) | 7 | 43 (+57 back-translation and meaning attestations) | 30,512 / 3,667 | $0.0147 | 415 s |
| norvexa-quick | 2 (CCDS, US PI) | 8 | 5,825 / 1,638 | $0.0042 | 148 s |
| norvexa-scan (carton as an image) | 2 | 3 (+ the document reader) | 2,915 / 292 | $0.0013 | 107 s |
| smoke (two short documents) | 2 | 3 | 2,321 total | $0.0011 | 81 s |
Times are dominated by queueing on the shared gateway; the self-hosted rehearsal (4 documents, direct route to the local Qwen3.8) took 14 s.
Rehearsal expectations (rehearsal/label-consistency-check/expected.json)
- The set is reported as drifted.
- The US PI against the CCDS has at least two likely errors in its warnings (the 6-month interval, the missing warning).
- The dropped German negation is flagged in the translation.
- The logged US-only indication difference is not called an error.
- The German SmPC's QRD headings are all right.
- Every model call has a signed receipt, and the signed record verifies.
Self-host check: a fresh clone (8a7c012) built from docker/api/Dockerfile, compose with a named volume, direct route to the local Qwen3.8-27B, the running language pack (:8491) and document reader (:8497): 7/7 checks passed in 13.8 s; torn down.
Limits
- Same author wrote the documents, the plants, the prompts and the key; the test medicine is new text, but the drift types and conventions are the ones the prompt names. Real label sets are longer and messier.
- Synthetic only: no real CCDS or EU label was used, and no comparison with a labelling reviewer's findings.
- English, German and French only in the QRD heading pack; other EU languages get the number lock and the meaning check but no heading check.
- Interactions, pregnancy sections, pharmacology and artwork layout are not compared.
- Test split: n = 31 drifts and 6 twins; 100% on 31 cases still leaves a 95% lower bound near 89%.