Skip to content
decosa

78 · Healthcare · Compliance and trust · live

Pharmacovigilance intake

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)

  • Missing minimum criteria caught17 / 20held outn = 20Run once. Reporter 2 of 5 (anonymous consumers passed); patient, product and event 5 of 5 each. A rule change found on this set makes it 20 of 20 (not held out).
  • Criteria wrongly called missing1 / 300held outn = 300'The mother of a teenage girl' not read as identifying the patient.
  • Seriousness right60 / 60held outn = 60Valid reports with an event; case level. Criteria set exact 59 of 60.
  • Expectedness right56 / 60held outn = 603 false unexpected, 1 unclear. Blind frontier judge (Claude Code Opus sub-agent): 59 of 60.
  • Day 0 exact55 / 59held outn = 59Misses: a vendor-sent report (2), a date without a year, a day-first date.
  • Deadlines exact, end to end78 / 86held outn = 86US 15-day, EU 15- and 90-day; misses follow day 0 and expectedness errors.
  • Clock code against the labeller's dates180 / 180held outn = 180And 60,000 of 60,000 random day-0s against a second implementation.
  • Kept fields outside the gold12 / 410 (2.9%)held outn = 4105 of 410 not written in the report at all (inferred consumer, sex from a name); 72 fields dropped by the quote check.
  • Reports not in English: expectedness right18 / 20held outn = 2024 reports in 7 languages; seriousness 20 of 20; validity 24 of 24.
  • Voiced calls: validity right10 / 10held outn = 10MOSS-Transcribe-Diarize on house-voice calls; seriousness 7 of 7, expectedness 6 of 7, day 0 6 of 7.

Dataset

80 synthetic adverse event reports (24 not in English) about four fictional products, with planted missing criteria, seriousness and expectedness nuances and earlier company-awareness dates, written and labelled by a separate agent that never saw the prompts; 10 call transcripts voiced with house voices.

Caveats

  • Synthetic reports from one labeller, not real case files or a safety physician's labels.
  • Run once; 80 reports, so the intervals are wide.
  • Two rule changes were made after the run from errors it found (anonymous consumers, day-first dates); the headline numbers are before them.
  • The narrative was off; calls not in English and scans beyond the demo form were not measured.
  • The frontier judge is a Claude Code sub-agent, blind to labels, not the API.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
28 Sep 2026
Latency, this run
n/a
p50 over passed runs
7.1 s
Receipts
4
Model calls
n/a
Tokens
n/a
Cost per run
$0.003

Self-host verification

Verified on 27 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile (ffmpeg present), compose api service with a named data volume, direct route, local signing; torn down after

The assembly prompt's smoke test passed against the already-running local servers (network_mode host instead of the compose llm, mt, diarize and docreader services): clock 2026-09-17 for us_15 and eu_15; the email valid, serious, unexpected, day 0 2026-09-02, due 2026-09-17; the call transcribed by the diarizer, valid, serious, unexpected, day 0 2026-09-15, due 2026-09-30; the German email not valid (missing patient) with the follow-up question in German; the record verified (19 entries); all receipts attested. The rehearsal bundle passed 18/18. Model-server startup was not re-run.

Rehearsal bundle: pv-intake.zip (1.0 MB, 18 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted numbers are from 6 production smoke runs after the merge: 3 on 27-28 Sep under load (29.8, 39.8, 92.0 s) and 3 on a quiet gateway on 28 Sep (7.0, 7.0, 7.1 s). Cost is the median at list price, model calls included (range $0.0028 to $0.0030).
  • Run once, 3 of 20 anonymous-reporter reports were passed as valid (fixed by a rule change found on the same set, so that fix is not held out).
  • Day 0 is missed when a medical information vendor sends the report ('our agent took the call on ...') or the earlier date has no year.
  • Measured on 80 synthetic reports from one labeller, not real case files or a safety physician's labels; calls not in English and Qwen3-ASR not measured; scans measured on the demo form only.
  • Speech recognition and the document reader share a busy GPU on the hosted demo; a call can fail to transcribe when that GPU is full (it retries, then says so).

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Intake, the quote checks, the four minimum criteria, day 0 and the clocks, follow-up questions, the E2B(R3)-shaped draft, signing and the HTTP API (/pv/*)decosa-api pharmacovigilance intake (decosa_api/verticals/pv), importing the language pack, the document reader, grounding (17), typed judgment (24) and the dates block (54)AGPL-3.0-or-later
  • Reads the fields with quotes, judges the seriousness criteria per event and expectedness against the label, drafts the narrative and judges its sentences (grounding)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
  • Call recordings in English to a timed, speaker-labelled transcript (the fallback for other languages)MOSS-Transcribe-Diarize 0.9BApache-2.0
  • Call recordings in German, French, Spanish, Italian, Dutch, Polish, Portuguese or Czech (the fallback for English)Qwen3-ASR-1.7B (language pack)Apache-2.0
  • Reports not in English, translated line by line to English, and follow-up questions back into the reporter's languageHy-MT2-7B (language pack)Apache-2.0
  • Finds and orders the regions of a scanned form (text, boxes, checkboxes) with their positionsDocling 2.130 with the Heron layout model (document reader block)MIT (Docling) + Apache-2.0 (weights)
  • Reads each region of a scanned formPaddleOCR-VL-1.6 (0.9B, document reader block)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

How well does it draft a case?

  • Missing minimum criteria caught: 17 of 20 (20 of 20 after a rule change found on this set)
  • Seriousness / expectedness right: 60/60 · 56/60 (blind Opus: 60/60 · 59/60)
  • Day 0 exact: 55 of 59 (clock arithmetic 180 of 180 against the labeller)
  • Fields kept that the report does not state: 12 of 410 (5 of 410 not written at all)

Source: decosa-api docs/evals/pv-intake.md, 27 Sep 2026

Lite · one 48-80 GB card (1)
  • the planted set: not measured yet
Standard · the hosted demo (8)
  • planted set, 80 reports (24 not in English), run once: missing minimum criteria caught: 17 of 20 run once (reporter 2 of 5; patient, product, event 5 of 5); 20 of 20 after a rule change found on this setdecosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
  • invented fields (patient and reporter fields kept that the report does not state): 12 of 410 kept fields outside the gold (2.9%); 5 of 410 not written in the report at alldecosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
  • seriousness right / expectedness right (valid cases with an event, case level): seriousness 60 of 60; expectedness 56 of 60decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
  • day 0 exact / deadlines exact (end to end): day 0 55 of 59; deadlines 78 of 86decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
  • clock code against the labeller's own dates and a second implementation: 180 of 180 labeller dates; 60,000 of 60,000 random day-0s against a second implementationdecosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
  • not in English: seriousness / expectedness / invented fields: 24 reports in 7 languages: seriousness 20 of 20, expectedness 18 of 20, 2 of 132 fields outside the golddecosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
  • calls (voiced from 10 call transcripts, transcribed by MOSS-Transcribe-Diarize): validity 10 of 10, seriousness 7 of 7, expectedness 6 of 7, day 0 6 of 7decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
  • frontier comparison (blind Claude Code Opus sub-agent, same cases): seriousness / expectedness: open 60/60 and 56/60; blind Opus 60/60 and 59/60decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route
Best · a larger judge (1)
  • the planted set: not measured yet
Wanted: the best setup · a large judge with room to spare (1)
  • the planted set: not measured yet

How we measure · All tools