Eval: pharmacovigilance intake (use case 78)
Measured on our server on 27 Sep 2026, branch the pre-release branch, gateway route (Qwen3.8-27B NVFP4, every call receipted), with
the narrative off (draft_narrative: false). Scores: docs/evals/pv-intake/results.json, raw outputs runs.json and
runs-audio.json, scorer scripts/pv_eval.py.
The set
- 80 synthetic adverse event reports about four fictional products (two drugs, two biologics), each with its reference
label, written and labelled by a separate labeller (a Claude Code sub-agent that never saw the prompts or the code),
from written rules (
pv-intake/label-rules.md). 56 in English, 24 not (German, Dutch and Polish 4 each; French, Spanish, Italian and Portuguese 3 each). 28 emails, 20 web forms, 16 call transcripts, 16 fax texts. - Planted: 20 invalid reports (5 each missing patient, reporter, product or event); seriousness nuances (ER visits with no admission, prolonged admissions, life threatening in the reporter's view, disability, congenital anomaly, important medical events); expectedness nuances (synonyms, more severe or more specific than the label, a fatal outcome of a listed event); 15 reports where a company employee or partner was told before the received date; absent fields.
- Gold deadlines were computed by the labeller's own script (
deadlines.py, plaintimedelta). - Calls: the 10 English two-voice call transcripts were voiced in the audio drama studio (use case 51) with Kokoro-82M
house voices, every line through the consent ledger (
pv-intake/audio/consent.json), then sent as MP3 audio. - Run once. No prompt was changed after the run. Two rule changes were made after the run, from errors found here (so the "after" numbers are not held out): (1) a reporter qualification of "consumer" no longer counts as an identifiable reporter on its own (GVP VI.B.2: consumer is the default when no qualification is given); recomputed from the stored outputs, since the criteria are code; (2) numeric dates in a report not in English are read day first.
Results (80 reports, run once)
| Measure | All | English (56) | Not English (24) |
|---|---|---|---|
| Case validity right | 76/80 | 52/56 | 24/24 |
| Missing minimum criteria caught | 17/20 (Wilson 95% 0.64-0.95) | 13/16 | 4/4 |
| Criteria wrongly called missing | 1/300 | ||
| Seriousness right (valid cases with an event) | 60/60 | 40/40 | 20/20 |
| Seriousness criteria set exact | 59/60 | ||
| Expectedness (case "unexpected") right | 56/60 (3 wrong, 1 unclear) | 38/40 | 18/20 |
| Day 0 exact (both valid) | 55/59 | 36/39 | 19/20 |
| Deadlines exact, end to end | 78/86 | 51/56 | 27/30 |
| Patient/reporter fields kept | 410 | 278 | 132 |
| ... of which not in the gold's stated fields | 12 (2.9%, Wilson 1.7-5.0%) | 10 | 2 |
| Fields dropped by the quote check | 72 |
- Missing criteria: patient 5/5, product 5/5, event 5/5, reporter 2/5. The 3 misses were anonymous consumers ("I don't
want to give my name") where the model filled
qualification: consumer, which met the reporter criterion. After rule change (1), recomputed from the same outputs: 20/20 caught, validity 79/80, nothing else changes. The one false "missing" (pv-e019): "the mother of a teenage girl" was not read as a patient descriptor. - The 12 fields outside the gold, read one by one: 1 truly invented (sex from a first name, pv-e067); 4 a "consumer" qualification inferred from the text rather than written; 4 the company's own staff (a nurse educator, a patient support programme) recorded as the reporter's organisation or qualification, where the labeller records only the original reporter; 3 written in the report but not listed by the labeller ("Who is reporting: Patient"). So values not in the text: 5 of 410 (1.2%).
- Expectedness errors: pv-e047 (a GI haemorrhage called life threatening, label lists GI haemorrhage without a severity qualifier: the labeller's own judgement call; the frontier judge also said unexpected); pv-e043 (the model split one GI bleed into "weakness" and "upper GI bleeding" and called the parts unlisted); pv-e014 ("very dry lips" against "dry mouth"); pv-e002 unclear (itchy skin vs pruritus).
- Day 0 misses: pv-e022 and pv-e055 (the report came from the company's medical information vendor, "our agent took the call on 16-Jan": the model did not treat the vendor as the company); pv-e059 ("when she visited on the 16th of November", no year: dates without a year are not accepted); pv-e052 (Portuguese "08/07/2026" read month first as 7 Aug, after the received date, so unused; rule change (2) now reads it as 8 July). The remaining deadline misses follow from these and from the expectedness errors.
- Clock code (the 15/90-day arithmetic, no model): 180/180 against the labeller's dates from gold inputs, and 60,000/60,000 against a second implementation (ordinal day arithmetic) on 20,000 random day-0s from 2024 to 2028.
Calls (10 voiced reports, MOSS-Transcribe-Diarize)
Validity 10/10, seriousness 7/7, expectedness 6/7, day 0 6/7, fields outside the gold 1/43; the same as the same 10 cases sent as text transcripts. The diarizer shares GPU0 and failed to allocate on some tries (CUBLAS_STATUS_ALLOC_FAILED); those runs were retried and the recogniser now retries itself. Qwen3-ASR (non-English calls) was not measured: its GPU service could not load while GPU0 was full.
Frontier comparison (seriousness and expectedness)
Judge: a Claude Code sub-agent (Opus 5.5), blind: it read only the 80 reports, the labels, the definitions and the output
schema (pv-intake/frontier-input.json), never the gold or the plants, and wrote frontier-verdicts.json; scored here
against the same gold, on the 60 valid cases with an event. No API spend.
| Open (Qwen3.8-27B) | Frontier (blind Opus) | |
|---|---|---|
| Seriousness right | 60/60 | 60/60 |
| Seriousness criteria set exact | 59/60 (1 extra) | 58/60 (2 extra) |
| Expectedness right | 56/60 | 59/60 |
| Expectedness errors | 3 false "unexpected", 1 unclear | 1 false "unexpected" (pv-e047, shared) |
The gap is in expectedness (3 cases on 60): splitting one event into parts and near-synonyms. Small n; not a large gap. Not an own-model priority on this evidence; a larger judge (the best tier) is the first thing to try.
Latency and cost
Median 145 s per report (90th percentile 275 s) with 4 reports in parallel on the shared, busy gateway; about 4 model calls per report without the narrative. Smoke run (the angioedema email, 4 calls): 132 s, 4,652 tokens, $0.0030 at the gateway list price.
Checkable properties of the demo (rehearsal/pv-intake)
POST /pv/clockwith day 0 2026-09-02, serious and unexpected: US and EU 15-day due 2026-09-17.- The angioedema email: valid, serious, unexpected, day 0 2026-09-02 (the sales representative), first deadline 2026-09-17.
- The signed case record verifies; changing its seriousness breaks it.
- The recorded call: valid, serious, unexpected (febrile neutropenia vs neutropenia), due 2026-09-30.
- The German email: not valid (no identifiable patient), no day 0, follow-up questions in German.
Limits
Synthetic reports from one labeller, not real case files or a safety physician's labels; 80 reports, run once; the narrative was off; Qwen3-ASR and non-English calls not measured; scanned forms were measured on the demo scan only (the fax_text cases are typed text).