78 · Healthcare · Compliance and trust · live
Pharmacovigilance intake
Eval results
Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)
- Missing minimum criteria caught17 / 20held outn = 20Run once. Reporter 2 of 5 (anonymous consumers passed); patient, product and event 5 of 5 each. A rule change found on this set makes it 20 of 20 (not held out).
- Criteria wrongly called missing1 / 300held outn = 300'The mother of a teenage girl' not read as identifying the patient.
- Seriousness right60 / 60held outn = 60Valid reports with an event; case level. Criteria set exact 59 of 60.
- Expectedness right56 / 60held outn = 603 false unexpected, 1 unclear. Blind frontier judge (Claude Code Opus sub-agent): 59 of 60.
- Day 0 exact55 / 59held outn = 59Misses: a vendor-sent report (2), a date without a year, a day-first date.
- Deadlines exact, end to end78 / 86held outn = 86US 15-day, EU 15- and 90-day; misses follow day 0 and expectedness errors.
- Clock code against the labeller's dates180 / 180held outn = 180And 60,000 of 60,000 random day-0s against a second implementation.
- Kept fields outside the gold12 / 410 (2.9%)held outn = 4105 of 410 not written in the report at all (inferred consumer, sex from a name); 72 fields dropped by the quote check.
- Reports not in English: expectedness right18 / 20held outn = 2024 reports in 7 languages; seriousness 20 of 20; validity 24 of 24.
- Voiced calls: validity right10 / 10held outn = 10MOSS-Transcribe-Diarize on house-voice calls; seriousness 7 of 7, expectedness 6 of 7, day 0 6 of 7.
Dataset
80 synthetic adverse event reports (24 not in English) about four fictional products, with planted missing criteria, seriousness and expectedness nuances and earlier company-awareness dates, written and labelled by a separate agent that never saw the prompts; 10 call transcripts voiced with house voices.
Caveats
- Synthetic reports from one labeller, not real case files or a safety physician's labels.
- Run once; 80 reports, so the intervals are wide.
- Two rule changes were made after the run from errors it found (anonymous consumers, day-first dates); the headline numbers are before them.
- The narrative was off; calls not in English and scans beyond the demo form were not measured.
- The frontier judge is a Claude Code sub-agent, blind to labels, not the API.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 28 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 7.1 s
- Receipts
- 4
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.003
Self-host verification
Verified on 27 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile (ffmpeg present), compose api service with a named data volume, direct route, local signing; torn down after
The assembly prompt's smoke test passed against the already-running local servers (network_mode host instead of the compose llm, mt, diarize and docreader services): clock 2026-09-17 for us_15 and eu_15; the email valid, serious, unexpected, day 0 2026-09-02, due 2026-09-17; the call transcribed by the diarizer, valid, serious, unexpected, day 0 2026-09-15, due 2026-09-30; the German email not valid (missing patient) with the follow-up question in German; the record verified (19 entries); all receipts attested. The rehearsal bundle passed 18/18. Model-server startup was not re-run.
Rehearsal bundle: pv-intake.zip (1.0 MB, 18 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted numbers are from 6 production smoke runs after the merge: 3 on 27-28 Sep under load (29.8, 39.8, 92.0 s) and 3 on a quiet gateway on 28 Sep (7.0, 7.0, 7.1 s). Cost is the median at list price, model calls included (range $0.0028 to $0.0030).
- Run once, 3 of 20 anonymous-reporter reports were passed as valid (fixed by a rule change found on the same set, so that fix is not held out).
- Day 0 is missed when a medical information vendor sends the report ('our agent took the call on ...') or the earlier date has no year.
- Measured on 80 synthetic reports from one labeller, not real case files or a safety physician's labels; calls not in English and Qwen3-ASR not measured; scans measured on the demo form only.
- Speech recognition and the document reader share a busy GPU on the hosted demo; a call can fail to transcribe when that GPU is full (it retries, then says so).
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Intake, the quote checks, the four minimum criteria, day 0 and the clocks, follow-up questions, the E2B(R3)-shaped draft, signing and the HTTP API (/pv/*)decosa-api pharmacovigilance intake (decosa_api/verticals/pv), importing the language pack, the document reader, grounding (17), typed judgment (24) and the dates block (54)AGPL-3.0-or-later
- Reads the fields with quotes, judges the seriousness criteria per event and expectedness against the label, drafts the narrative and judges its sentences (grounding)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
- Call recordings in English to a timed, speaker-labelled transcript (the fallback for other languages)MOSS-Transcribe-Diarize 0.9BApache-2.0
- Call recordings in German, French, Spanish, Italian, Dutch, Polish, Portuguese or Czech (the fallback for English)Qwen3-ASR-1.7B (language pack)Apache-2.0
- Reports not in English, translated line by line to English, and follow-up questions back into the reporter's languageHy-MT2-7B (language pack)Apache-2.0
- Finds and orders the regions of a scanned form (text, boxes, checkboxes) with their positionsDocling 2.130 with the Heron layout model (document reader block)MIT (Docling) + Apache-2.0 (weights)
- Reads each region of a scanned formPaddleOCR-VL-1.6 (0.9B, document reader block)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
How well does it draft a case?
- Missing minimum criteria caught: 17 of 20 (20 of 20 after a rule change found on this set)
- Seriousness / expectedness right: 60/60 · 56/60 (blind Opus: 60/60 · 59/60)
- Day 0 exact: 55 of 59 (clock arithmetic 180 of 180 against the labeller)
- Fields kept that the report does not state: 12 of 410 (5 of 410 not written at all)
Source: decosa-api docs/evals/pv-intake.md, 27 Sep 2026
Lite · one 48-80 GB card (1)
- the planted set: not measured yet
Standard · the hosted demo (8)
- planted set, 80 reports (24 not in English), run once: missing minimum criteria caught: 17 of 20 run once (reporter 2 of 5; patient, product, event 5 of 5); 20 of 20 after a rule change found on this setdecosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
- invented fields (patient and reporter fields kept that the report does not state): 12 of 410 kept fields outside the gold (2.9%); 5 of 410 not written in the report at alldecosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
- seriousness right / expectedness right (valid cases with an event, case level): seriousness 60 of 60; expectedness 56 of 60decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
- day 0 exact / deadlines exact (end to end): day 0 55 of 59; deadlines 78 of 86decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
- clock code against the labeller's own dates and a second implementation: 180 of 180 labeller dates; 60,000 of 60,000 random day-0s against a second implementationdecosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
- not in English: seriousness / expectedness / invented fields: 24 reports in 7 languages: seriousness 20 of 20, expectedness 18 of 20, 2 of 132 fields outside the golddecosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
- calls (voiced from 10 call transcripts, transcribed by MOSS-Transcribe-Diarize): validity 10 of 10, seriousness 7 of 7, expectedness 6 of 7, day 0 6 of 7decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
- frontier comparison (blind Claude Code Opus sub-agent, same cases): seriousness / expectedness: open 60/60 and 56/60; blind Opus 60/60 and 59/60decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
Best · a larger judge (1)
- the planted set: not measured yet
Wanted: the best setup · a large judge with room to spare (1)
- the planted set: not measured yet