Skip to content
decosa

29 · Healthcare · Compliance and trust · live

Clinical AI assurance monitor

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 28 Sep 2026Eval write-up (decosa-api, access required)

  • Invented fact caught as an error49 (98%)test splitn = 50
  • Changed detail caught as an error47 (94%)test splitn = 50With the detail checker (M17). The judge alone: 36 (72%). False error flags on faithful sentences unchanged: 4 of 1,135
  • Key item left out caught45 (90%)test splitn = 5048/50 after a guard was removed post-test; that number is not held out
  • Wrong speaker caught / typed correctly47 (94%) / 31 (62%)test splitn = 50
  • Invented exam finding caught as an error49 (98%)test splitn = 50
  • False error flags on faithful notes (sentences)4 (0.35%)test splitn = 1,135Key items called left out: 3 of 368 (0.8%); notes with at least one error finding: 6 of 50 (12%)
  • Tighten: words removed (median per note)15%test splitn = 50Range 0-31%; 15,835 to 13,499 words over 50 clean notes; 0 words added (checked in code on every output)
  • Tighten: key items that stopped being fully recorded5 of 362 (1.4%)test splitn = 362All 5 came back 'partly recorded', none missing; e.g. 'over the counter' dropped from a medicine
  • Changed detail caught as an error, 200 more blind plants188 (94.0%)test splitn = 20095% CI 89.8-96.5%; judge alone 147 (73.5%); written by a blind sub-agent on the 50 held-out visits, never trained on
  • Detail checker alone on CPU: changed details caught / faithful sentences flagged33 of 50 (66%) / 3 of 1,135 (0.26%)test splitn = 50No language model, windows by word overlap, p(changed) >= 0.9; median 7.1 s per note on 8 CPU threads

Dataset

All 57 PriMock57 mock primary-care consultations (CC BY 4.0) with reference transcripts; one synthetic cited scribe note per visit plus five copies each with one planted error; 7 visits dev, 50 held out.

Caveats

  • The notes are synthetic and the plants are clean single errors; real scribe errors are subtler, so these detection rates are an upper bound.
  • Reference transcripts: no speech-recognition error enters this eval.
  • One judge model family; a second judge is not run. Plants were checked by a validator, not reviewed by a clinician.
  • Primary care, English, remote visits, 50 test visits; no specialty or inpatient data and no clinician agreement study.
  • The detail checker's thresholds were fixed on the dev split before the test run. On ACI-Bench's human notes (real speech-recognition transcripts) it turned 9 of 1,469 sentences into errors: 2 real conflicts, 3 details never said aloud, 4 not errors. On ACI-Bench plants the judge alone caught 71/78 and with the checker 72/78.
  • A blind frontier model (Claude Opus via Claude Code) proofreading 40 of the same notes caught 20/20 changed details with 1 flag on 20 faithful notes; this tool caught 19/20 with 2 flags, on open weights you can run yourself.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
28 Sep 2026
Latency, this run
n/a
p50 over passed runs
28 s
Receipts
30
Model calls
n/a
Tokens
n/a
Cost per run
$0.021

Self-host verification

Verified on 25 Sep 2026: assemble-prompt.md on our server: fresh clone into a clean directory, api image built from docker/api/Dockerfile, compose with a named volume, pointed at the already-running Qwen3.8-27B vLLM (127.0.0.1:8114) and MOSS diarizer (127.0.0.1:8092) instead of starting new ones; then torn down

Images build, the service starts, and the smoke steps pass: the wrong-speaker sample flagged (17 s), the report verifies with transcript and note hashes, a changed count fails verification, the summary signs, and 170 s of audio came back as 22 speaker turns. Model-server startup itself was not re-verified. /healthz says ok: false on this stack because it also checks the live-scribe recogniser, which the monitor does not use.

Rehearsal bundle: clinical-ai-monitor.zip (8 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Measured on synthetic notes with one clear planted error each; real scribe errors are subtler.
  • Changed details: 94% caught as errors with the detail checker (47/50 and 188/200 blind plants), 72-74% by the judge alone; the checker only upgrades the judge's own 'detail not in the visit' flags.
  • On real speech-recognition transcripts the detail checker adds some false error flags (4 in 1,469 ACI-Bench sentences, e.g. 'type i' vs 'type 1'); each comes with its transcript line.
  • Wrong-speaker errors are caught but typed correctly only 62% of the time.
  • Tighten shortened the 50 test notes by a median 15%; 5 of 362 recorded key items came back 'partly recorded' after tightening (none missing).
  • Measures against the transcript it is given; speech-recognition errors pass into the measure.
  • Clinicians' own notes get flagged for things never said aloud; use it on AI drafts.
  • Audio input is API-only (POST /monitor/transcribe); the page takes text.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Monitor: transcript and note parsing, sentence and section offsets, finding types, the signed report and the summary with intervals and drift (no model; CPU)decosa-api clinical AI assurance monitor (decosa_api/verticals/monitor)AGPL-3.0-or-later
  • Judge: one call per note sentence, one checklist extraction, one coverage check; also the speaker role map for anonymous labelsQwen3.8-27B (NVFP4)Apache-2.0
  • Detail checker (M17): after the judge, each drug, dose, frequency, route, date, duration, side and number in a sentence is read against its transcript lines and labelled same / changed / absent; a judge 'detail not in the visit' flag becomes a changed-detail error when p(changed) >= 0.5decosa-note-detail-checker-modernbert-large (M17, our own model; Apache-2.0 on Hugging Face)Apache-2.0
  • Audio in (optional): one-pass speaker-attributed transcript of the whole visitMOSS-Transcribe-Diarize 0.9BApache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · text transcripts, one 32 GB card (1)
  • Same judge and prompts as standard, so the same detection and false-flag rates: see standarddocs/evals/clinical-ai-monitor.md
Standard · judge plus diarizer (hosted demo) (6)
  • Planted errors caught at error severity (50 held-out PriMock57 visits, one error per note): invented fact / invented exam finding / wrong speaker / key item left out / changed detail: 49/50 / 49/50 / 47/50 / 45/50 / changed detail 47/50 with the detail checker (36/50 judge alone); every plant flagged at least as review (50/50 each)docs/evals/clinical-ai-monitor.md
  • Changed details, 200 more blind plants on the same 50 held-out visits (8 detail types): 188/200 (94%) with the detail checker, 147/200 (73.5%) judge alone; false error flags on faithful notes unchanged (4/1,135)docs/evals/clinical-ai-monitor.md (M17)
  • Wrong-speaker errors typed as wrong speaker: 31/50 (62%); the rest flagged as contradicts or not in the visitdocs/evals/clinical-ai-monitor.md
  • Error flags on faithful notes (1,135 sentences, 368 key items): 4 sentences (0.35%; 1 real error in the note, 2 judge mistakes, 1 debatable), 3 key items (0.8%; 1 real)docs/evals/clinical-ai-monitor.md
  • Clinicians' own PriMock57 notes: lines with an error finding: 12.1%; in a sample of 20, 16 were statements the transcript does not support (names, routine negatives never asked)docs/evals/clinical-ai-monitor.md
  • Speech recognition, if you send audio (MOSS-Transcribe-Diarize, PriMock57): 10.3% WERscribe-bench RESULTS.md (the clinical scribe's pass 2)
Wanted · a second judge from another family (1)
  • This eval, same protocol: not measured yet

How we measure · All tools