Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: pharmacovigilance intake (use case 78)

Measured on our server on 27 Sep 2026, branch the pre-release branch, gateway route (Qwen3.8-27B NVFP4, every call receipted), with the narrative off (draft_narrative: false). Scores: docs/evals/pv-intake/results.json, raw outputs runs.json and runs-audio.json, scorer scripts/pv_eval.py.

The set

  • 80 synthetic adverse event reports about four fictional products (two drugs, two biologics), each with its reference label, written and labelled by a separate labeller (a Claude Code sub-agent that never saw the prompts or the code), from written rules (pv-intake/label-rules.md). 56 in English, 24 not (German, Dutch and Polish 4 each; French, Spanish, Italian and Portuguese 3 each). 28 emails, 20 web forms, 16 call transcripts, 16 fax texts.
  • Planted: 20 invalid reports (5 each missing patient, reporter, product or event); seriousness nuances (ER visits with no admission, prolonged admissions, life threatening in the reporter's view, disability, congenital anomaly, important medical events); expectedness nuances (synonyms, more severe or more specific than the label, a fatal outcome of a listed event); 15 reports where a company employee or partner was told before the received date; absent fields.
  • Gold deadlines were computed by the labeller's own script (deadlines.py, plain timedelta).
  • Calls: the 10 English two-voice call transcripts were voiced in the audio drama studio (use case 51) with Kokoro-82M house voices, every line through the consent ledger (pv-intake/audio/consent.json), then sent as MP3 audio.
  • Run once. No prompt was changed after the run. Two rule changes were made after the run, from errors found here (so the "after" numbers are not held out): (1) a reporter qualification of "consumer" no longer counts as an identifiable reporter on its own (GVP VI.B.2: consumer is the default when no qualification is given); recomputed from the stored outputs, since the criteria are code; (2) numeric dates in a report not in English are read day first.

Results (80 reports, run once)

Measure All English (56) Not English (24)
Case validity right 76/80 52/56 24/24
Missing minimum criteria caught 17/20 (Wilson 95% 0.64-0.95) 13/16 4/4
Criteria wrongly called missing 1/300
Seriousness right (valid cases with an event) 60/60 40/40 20/20
Seriousness criteria set exact 59/60
Expectedness (case "unexpected") right 56/60 (3 wrong, 1 unclear) 38/40 18/20
Day 0 exact (both valid) 55/59 36/39 19/20
Deadlines exact, end to end 78/86 51/56 27/30
Patient/reporter fields kept 410 278 132
... of which not in the gold's stated fields 12 (2.9%, Wilson 1.7-5.0%) 10 2
Fields dropped by the quote check 72
  • Missing criteria: patient 5/5, product 5/5, event 5/5, reporter 2/5. The 3 misses were anonymous consumers ("I don't want to give my name") where the model filled qualification: consumer, which met the reporter criterion. After rule change (1), recomputed from the same outputs: 20/20 caught, validity 79/80, nothing else changes. The one false "missing" (pv-e019): "the mother of a teenage girl" was not read as a patient descriptor.
  • The 12 fields outside the gold, read one by one: 1 truly invented (sex from a first name, pv-e067); 4 a "consumer" qualification inferred from the text rather than written; 4 the company's own staff (a nurse educator, a patient support programme) recorded as the reporter's organisation or qualification, where the labeller records only the original reporter; 3 written in the report but not listed by the labeller ("Who is reporting: Patient"). So values not in the text: 5 of 410 (1.2%).
  • Expectedness errors: pv-e047 (a GI haemorrhage called life threatening, label lists GI haemorrhage without a severity qualifier: the labeller's own judgement call; the frontier judge also said unexpected); pv-e043 (the model split one GI bleed into "weakness" and "upper GI bleeding" and called the parts unlisted); pv-e014 ("very dry lips" against "dry mouth"); pv-e002 unclear (itchy skin vs pruritus).
  • Day 0 misses: pv-e022 and pv-e055 (the report came from the company's medical information vendor, "our agent took the call on 16-Jan": the model did not treat the vendor as the company); pv-e059 ("when she visited on the 16th of November", no year: dates without a year are not accepted); pv-e052 (Portuguese "08/07/2026" read month first as 7 Aug, after the received date, so unused; rule change (2) now reads it as 8 July). The remaining deadline misses follow from these and from the expectedness errors.
  • Clock code (the 15/90-day arithmetic, no model): 180/180 against the labeller's dates from gold inputs, and 60,000/60,000 against a second implementation (ordinal day arithmetic) on 20,000 random day-0s from 2024 to 2028.

Calls (10 voiced reports, MOSS-Transcribe-Diarize)

Validity 10/10, seriousness 7/7, expectedness 6/7, day 0 6/7, fields outside the gold 1/43; the same as the same 10 cases sent as text transcripts. The diarizer shares GPU0 and failed to allocate on some tries (CUBLAS_STATUS_ALLOC_FAILED); those runs were retried and the recogniser now retries itself. Qwen3-ASR (non-English calls) was not measured: its GPU service could not load while GPU0 was full.

Frontier comparison (seriousness and expectedness)

Judge: a Claude Code sub-agent (Opus 5.5), blind: it read only the 80 reports, the labels, the definitions and the output schema (pv-intake/frontier-input.json), never the gold or the plants, and wrote frontier-verdicts.json; scored here against the same gold, on the 60 valid cases with an event. No API spend.

Open (Qwen3.8-27B) Frontier (blind Opus)
Seriousness right 60/60 60/60
Seriousness criteria set exact 59/60 (1 extra) 58/60 (2 extra)
Expectedness right 56/60 59/60
Expectedness errors 3 false "unexpected", 1 unclear 1 false "unexpected" (pv-e047, shared)

The gap is in expectedness (3 cases on 60): splitting one event into parts and near-synonyms. Small n; not a large gap. Not an own-model priority on this evidence; a larger judge (the best tier) is the first thing to try.

Latency and cost

Median 145 s per report (90th percentile 275 s) with 4 reports in parallel on the shared, busy gateway; about 4 model calls per report without the narrative. Smoke run (the angioedema email, 4 calls): 132 s, 4,652 tokens, $0.0030 at the gateway list price.

Checkable properties of the demo (rehearsal/pv-intake)

  1. POST /pv/clock with day 0 2026-09-02, serious and unexpected: US and EU 15-day due 2026-09-17.
  2. The angioedema email: valid, serious, unexpected, day 0 2026-09-02 (the sales representative), first deadline 2026-09-17.
  3. The signed case record verifies; changing its seriousness breaks it.
  4. The recorded call: valid, serious, unexpected (febrile neutropenia vs neutropenia), due 2026-09-30.
  5. The German email: not valid (no identifiable patient), no day 0, follow-up questions in German.

Limits

Synthetic reports from one labeller, not real case files or a safety physician's labels; 80 reports, run once; the narrative was off; Qwen3-ASR and non-English calls not measured; scanned forms were measured on the demo scan only (the fax_text cases are typed text).