Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Clinical AI assurance monitor: eval (25 Sep 2026)

Questions: (1) does the monitor catch each kind of scribe error: an invented fact, a changed detail, a key item left out, the wrong speaker, an invented exam finding? (2) how often does it flag a faithful note? (3) what does it do with real clinicians' notes?

Data

  • Visits: all 57 PriMock57 mock primary-care consultations (Papadopoulos Korfiatis et al., 2022; github.com/babylonhealth/primock57; CC BY 4.0), remote GP consultations role-played by clinicians and actors. The reference transcripts come from the per-speaker annotations (consecutive turns by one speaker merged; 40-134 turns, median 8,400 characters), so speaker labels are exact and there is no speech-recognition error in this eval.
  • Clean notes: for each visit, a US-style ambient-scribe note (chief complaint, HPI, history, medications, allergies, social and family history, exam, assessment and plan; 16-32 sentences) written by a different model family (Claude) than the judge, with every sentence cited to transcript lines, and required to include every medication, allergy status, plan item and safety-net instruction. The spec is in docs/evals/clinical-ai-monitor/encounters.jsonl (each row holds the transcript, the note as cited sentences, and the plants).
  • Planted notes: five copies of each clean note, each with exactly one error:
    • fabricated: one inserted sentence with a fact nowhere in the visit (a past diagnosis, a drug, a symptom, a pet);
    • altered: one sentence with one detail changed to conflict with the visit (dose, duration, side, number, a flipped negation);
    • omission: every mention of one key item deleted (a prescribed drug, a referral, an allergy, safety-net advice);
    • speaker: a fact from the visit given to the wrong person (the doctor's suggestion as the patient's belief, a relative's condition as the patient's own);
    • exam: an exam finding or vital sign nobody measured (these were remote visits).
  • Human notes: each visit's own PriMock57 clinician note (terse UK GP style: "3/7 hx of diarrhoea. PMH: asthma. Imp: gastroenteritis."). Not a false-flag test: clinicians write things that were never said aloud.
  • Split: 7 visits for development (hash of the id), 50 held out. Nothing was tuned on dev either: the prompts were written before any run and not changed; dev was a pipeline check. The test run started after the code was frozen (commit f73b4e0). One change was made after the test run, reported below.

Method

Each note goes through the same code the API runs (Monitor in decosa_api/verticals/monitor/check.py): one Qwen3.8-27B call per checked sentence against the whole numbered transcript, one checklist extraction and one coverage call. The dev split and the first two test visits used the model gateway route; the gateway then slowed to a 12.6 s median per call under other load, so the rest of the test split called the same vLLM server directly (same weights, prompts and temperature 0). Identical prompts are answered from a disk cache, so sentences shared by the clean and planted copies of a note are judged once.

Scoring:

  • A sentence plant is caught when the planted sentence gets an error finding (not in the visit, contradicts the visit, wrong speaker, invented exam); flagged when it gets any finding (including "detail not in the visit", which is review level); typed when the finding matches the plant (fabricated and altered: not in the visit or contradicts; speaker: wrong speaker; exam: invented exam).
  • An omission is caught when a checklist item at the omitted lines (or sharing a word with the omitted item) is "left out"; flagged when it is "left out" or "partly recorded".
  • False flags are findings on the clean notes.

Results: planted errors (test, 50 held-out visits)

Error Planted Caught (error) Flagged (any) Typed correctly
Invented fact 50 49 (98%) 50 49
Changed detail 50 36 (72%) 50 36
Key item left out 50 45 (90%) 50 45
Wrong speaker 50 47 (94%) 50 31 (62%)
Invented exam finding 50 49 (98%) 50 49

Wilson 95% intervals: 49/50 0.90-1.00; 47/50 0.84-0.98; 45/50 0.79-0.96; 36/50 0.58-0.83.

  • Every plant was flagged. The misses as errors are all "review" flags instead:
    • Changed detail: 14 of 50 came back "detail not in the visit" (PARTIAL) rather than "contradicts". The judge reads "two weeks" against "four days" as a changed detail; both are flagged, at a lower severity.
    • Left out: 5 came back "partly recorded". Three of those were caused by a guard that turned "left out" into "partly recorded" when a word of the medicine's name appeared anywhere in the note ("dizziness", "anti-inflammatories"). The guard never fired on a clean note in dev or test, so it was removed after the test run; without it the three would have been caught (48/50), but that number is not held out.
    • Wrong speaker: 16 of the 47 caught were typed "contradicts" or "not in the visit" instead of "wrong speaker" (the judge sees the conflict but not whose it is); 3 were only "detail", all doctor's suggestions written as the patient's (the takeaway as the cause, named antihistamines, sore ears the patient never reported).
    • Invented fact: one came back "detail" (a daily iron supplement appended to a real medication sentence).
    • Exam: one came back "detail" (a photo finding in a visit where the doctor did look at a photo).
  • The key item behind every omission plant was in the extracted checklist (50/50), so the checklist step itself missed nothing here.

Results: false flags on the clean notes (test, 50 notes, 1,135 checked sentences, 368 key items)

Count Rate
Sentences with an error finding 4 0.35% (95% 0.14-0.90%)
Sentences with a review finding ("detail") 9 0.8%
Key items called "left out" 3 0.8% of items (95% 0.3-2.4%)
Key items called "partly recorded" 3 0.8%
Notes with at least one error finding 6 12%

What the 7 error-level false flags were, on review:

  • Sentences (4): one real error in the "clean" note (it said vomiting was less prominent than diarrhoea; the patient said the opposite); two judge mistakes ("Metformin." when the doctor says "Metformin, OK"; a mother staying until the ambulance comes, which the doctor asked for); one debatable (the note calls an inconsistent inhaler history unclear; the judge reads one answer as clear).
  • Left out (3): one real (the doctor said the consultant would call back; the note does not say so); two not real: the checklist listed "no known drug allergies" and "allergy status not confirmed" where the patient's answer is not in the transcript.
  • The nine "detail" flags are mostly the judge refusing a word the transcript does not quite say ("video" appointment, "mild" hay fever, "stayed on the line until the ambulance arrived").

So on faithful notes the judge's own error rate is about 2 in 1,135 sentences and 2 in 368 key items, plus a few debatable calls. Six of the 50 clean notes got an error finding: two for real errors in the note, three for judge mistakes and one debatable, so roughly 1 faithful note in 12 carries a wrong error flag.

Results: the clinicians' own notes (test, 50 notes, 1,075 checked lines)

  • 12.1% of lines get an error finding (130; 95% 10.3-14.2%), another 11% a "detail" flag; 17% of checklist items are "left out"; 46 of 50 notes have at least one error finding.
  • A random sample of 20 flagged lines: 16 are statements the transcript does not support (names and results never said, such as "DH: Mercilon" or "(Low T3/4)"; routine negatives never asked, such as "No altered taste/smell" or "NKDA"; a duration of one week where the patient said three; "no frequency" where the patient reported frequency); 2 are judge errors ("LMP 2/52" read against "three weeks, I mean two weeks"; a paraphrase of vertigo); 2 are unclear.
  • This is the expected difference between a clinician's note and a scribe's draft: clinicians record their own knowledge and habits, and a transcript-grounded check calls those unsupported. Use the monitor on AI drafts; do not read these rates as clinician error rates.

Dev split (7 visits, for reference)

Caught: invented fact 7/7, changed detail 7/7, left out 7/7, wrong speaker 7/7 (5 typed), exam 7/7. Clean notes: 1 error flag in 175 sentences (the diabetes type the note calls unclear, debatable) and 1 of 55 items called left out.

Latency and cost

  • Direct to the model server (4 sentence calls in flight, the eval's setting): median 12.6 s per clean note of about 23 sentences, 2-3 s for a planted copy with the rest cached.
  • Through the model gateway on our server, shared with other workloads: 24 s median per note on the dev split; 75-172 s for the four recorded demo runs later in the day, when the gateway's median call took 12.6 s.
  • Generated tokens: about 107 per call (311,125 over 2,919 calls). A 27-sentence note is 29 calls, about 3,100 generated and about 95,000 prompt tokens (the transcript repeats in every call; the prefix cache makes most of it cheap on the server, but list prices count it). At the gateway list price used by the nightly checks ($0.30 per million prompt tokens, $1.50 per million generated) that is about $0.03 per note, an estimate from token counts.

Rework as "Check an AI note" (28 Sep 2026): tighten, buyer scorecard, time and cost

The sentence judge, checklist and coverage prompts are unchanged (same prompts_sha256), so every detection and false-flag number above still describes the check. What was added, and what it measured:

Tighten (only removes or condenses)

  • Data: the 50 held-out clean notes above (15,835 words). The prompt was written and tried on the three demo samples, which are dev-split visits.
  • Method: scripts/monitor_tighten_eval.py: tighten each note (one proposal call; one repair call when code refuses a proposed sentence), check in code that no word was added, then re-run the coverage call on the tightened note with the same key items the check had extracted from the transcript, and compare each item's status. Direct route to the model server (same weights as the gateway), temperature 0, 129 calls.
  • Results:
Result
Words removed median 15% per note (range 0-31%); 15,835 → 13,499 words in total
Notes with any word not in the original 0 of 50 (checked in code on every output)
Key items recorded before tightening 362 of 368
... still fully recorded after 357; 5 (1.4%) became "partly recorded", none missing
Guard 66 proposed sentences refused and repaired in the second call, 2 replaced by the original words, 6 sentences put back because a detail went missing
Time median 12 s per note (direct route)

The five: "over the counter" dropped from Dioralyte; a safety-net list shortened; "in the shower" dropped from Dermol (twice); and a sentence on diarrhoea frequency condensed. Words that are not numbers, doses, frequencies, sides, routes or medicine names can be deleted, so a qualifier like "over the counter" can go; the tightened note is a draft for the clinician, and the original stays.

  • The guard is strict on purpose: in-order words mean no paraphrase at all, so the notes shrink less than a free rewrite would, but a tightened note can only say less, never something different.

Buyer scorecard

POST /monitor/scorecard runs the same check per visit and signs a summary; POST /monitor/summary renders it as Markdown. The rates are what the checker found, printed next to its own held-out catch and false-flag rates above, so a reader can see that rates near 0.35 per 100 sentences are within its noise. No new accuracy claim.

Time and cost, hosted (28 Sep 2026)

  • 9 runs through the gateway (the three note samples of the demo, three times each, tighten on, the gateway shared with other work): p50 28.0 s, max 41.4 s end to end. The check alone: median 23 s. The tighten proposal now runs alongside the check; before that change, 9 runs gave p50 38.7 s (max 71.4 s). The 116.9 s p50 of 25 Sep was measured under much heavier gateway load.
  • Cost at list price from metered tokens: $0.010 for the 22-sentence demo note (18 calls, ~21,000 prompt tokens), $0.051-0.053 for a 27-sentence PriMock57 note against a 2,500-word transcript (30 calls, ~156,000 prompt tokens: the transcript is sent with every sentence call). Tighten adds one or two calls.

M17: the detail checker (28 Sep 2026)

The judge's weakest class was changed details: 14 of the 50 changed-detail plants came back as "a detail not in the visit" (review) instead of an error. M17 is our own small model that settles those flags.

The model

  • decosa-note-detail-modernbert-large: ModernBERT-large (Apache-2.0), initialised from our published decosa-grounding-modernbert-large and fine-tuned on our server's GPU (7-9 minutes per run, three runs; v3 kept). Input: the note sentence with one detail marked, and the transcript lines picked for it (word overlap, a line that says the same value, and the judge's cited lines). Output: same / changed / absent, plus token-level evidence.
  • Details come from rules plus a medicine lexicon (decosa_api/verticals/monitor/details.py: drug, dose, frequency, route, date, duration, side, number; openFDA label names, CC0).
  • Service: services/detail_checker (CPU, 127.0.0.1:8495, DECOSA_DETAIL_URL); each call gets a signed model-call receipt. The API calls it per judged sentence that has a detail, while the judge's other calls run.
  • The rule (frozen on dev before the test run): a sentence the judge flagged PARTIAL becomes a "changed detail" error when one of its details has p(changed) >= 0.5. SUPPORTED sentences are never upgraded (on dev that added 4-5 false flags in 175 faithful sentences; the PARTIAL-only rule added none). It never removes or lowers a judge flag.

Training data (no test visit, no Claude-written text)

  • 260 synthetic visits and scribe notes written by Qwen3.8-27B (direct route to our model server), plus Qwen notes for the 7 PriMock57 dev visits and the 87 ACI-Bench train and valid encounters (CC BY 4.0), plus ACI-Bench's own human notes. 4 note generations failed JSON and were skipped. Training uses the synthetic visits (9 in 10) and ACI-Bench train; dev is the other synthetic visits, ACI-Bench valid and the PriMock57 dev visits.
  • Pairs labelled by code: same when the detail's value is found among the same kind of values in its window (after normalising number words, units and misspelt medicines); changed when a same detail is swapped for another value of its kind; absent when the evidence lines are removed (also with other values of the kind left in the window) or a dose, frequency or side is inserted. Train 26,144 pairs, dev 4,338 (synthetic visits by hash, ACI-Bench valid, PriMock57 dev).
  • What changed between runs, all on dev: v1 called faithful details "changed" when the window missed the evidence but held another value of the same kind (9/175 dev sentences); v2 added value-aware windows and distractor absents; v3 fixed the number normaliser ("sixty five kilograms", "thirty seven point five" were being dropped from the same class, so the model had learned "words against digits means changed").

Results (test; one run after freezing)

Set Judge alone Judge + M17
The 50 changed-detail plants above 36 (72%) 47 (94%) (95% CI 84-98%)
200 more changed-detail plants on the same 50 visits, written by a blind Claude Code sub-agent (8 types, never trained on) 147 (73.5%) 188 (94.0%) (95% CI 89.8-96.5%)
Error flags on the 50 faithful notes (1,135 sentences) 4 (0.35%) 4 (0.35%): M17 upgraded none
Invented fact / wrong speaker / invented exam 49 / 47 / 49 unchanged

By type on the 200 (judge alone → with M17): date 28 → 41 of 44, duration 30 → 38 of 41, number 26 → 33 of 36, drug 26 → 29 of 29, side 17 → 18 of 19, frequency 9 → 14 of 16, dose 6 → 9 of 9, route 5 → 6 of 6. Twelve plants stay missed: in all twelve the judge flagged the sentence for review and M17 read the detail as the same or absent (below 0.5).

The integrated code path (scripts/monitor_eval.py --detail-url, the API's own Monitor with the service) reproduces the first and third rows exactly: 47/50 caught, 4/1,135 faithful sentences flagged, 11 upgrades, all on planted sentences.

M17 alone, no language model (windows by word overlap only, the CPU-only mode a practice could run without a GPU): at p(changed) >= 0.9, 33 of 50 changed details caught (66%) and 3 of 1,135 faithful sentences flagged (0.26%); on the 200 extra plants 121 (60.5%). With the judge's lines in the window it catches more (39/50) but flags 14/1,135. Time on CPU: median 7.1 s per note (8 threads, ~18 details per note, our server under other load; 12 notes), against the judge's 12.6 s on a GPU and 116.9 s hosted under load on 25 Sep.

Against a frontier proofreader (blind, 28 Sep 2026)

A Claude Code sub-agent (Opus, blind: it saw only the transcript and the note, never the labels) proofread 40 of the held-out notes, 20 with a changed-detail plant and 20 faithful, as a careful clinician would before signing:

Frontier proofreader Judge alone Judge + M17
Changed details caught (20) 20 15 19
Faithful notes with an error flag (20) 1 2 2
The frontier model is better and should be; ours runs on open weights on the practice's own hardware in 20-30 seconds
for about $0.01-0.05. Its own estimate of the time a clinician needs to proofread each note that carefully was 3-8
minutes (median 6), an AI's estimate, not a timed human study. Judge: Claude Code Opus 5.5, blind, n = 40.

ACI-Bench (CC BY 4.0, licence checked 28 Sep 2026)

ACI-Bench test set 1 (39 of its 40 encounters; one note is over the 80-sentence limit): in-person visits with dictated exams, US English, human-written notes, transcripts with the speech-recognition and speaker-tag errors the dataset keeps on purpose. M17 was trained on ACI-Bench train only.

  • 78 changed-detail plants (2 per note, 8 types, written by a blind Claude Code sub-agent on sentences it checked against the transcript): judge alone 71 (91%), judge + M17 72 (92.3%). The judge does much better here than on PriMock57 (in-person visits with clear dictation), so M17 has less to add.
  • The human notes as written (1,469 checked sentences) are not a false-flag test, because clinicians write things never said aloud. The judge raised error flags on 97 sentences (6.6%), and M17 turned 9 more "detail not in the visit" flags into errors. Read by hand:
    • 2 real conflicts: "the past several months" against "six weeks", and Lasix 40 mg against "four milligrams".
    • 3 details never said aloud: ages worked out from a date of birth, and a prednisone strength.
    • 4 that are not errors: "type 1 diabetes" against the recogniser's "type i diabetes"; "yesterday around dinnertime" against "last night around supper time"; "04/2022" against "starting in April"; and "right arm" against "my arm hurts right here". So on real transcripts M17 adds some false error flags (about 4 in 1,500 sentences here). They come with the transcript lines, so a reader can dismiss them in seconds, but they count.

Limits

  • The plants are single clean swaps on synthetic notes; the extra 200 were written by one model family (Claude), the training plants by code. Real scribe errors on real transcripts are not measured.
  • The rule depends on the judge flagging the sentence first: a changed detail the judge calls SUPPORTED stays missed.
  • Standalone M17 is not a replacement for the judge: it only sees details, not invented facts, speakers or exams.

Limits of this eval

  • The notes are synthetic and the plants are clean single errors. Real scribe errors are subtler and cluster (a wrong drug in the HPI and the plan). The detection rates are an upper bound for errors this clear.
  • Reference transcripts. No speech-recognition error: in production the monitor measures against the hospital's own transcript (the vendor's, or the diarizer's), and transcript errors pass into the measure.
  • One judge. The same model family judges every note; its blind spots set the rate for every vendor alike, which is fair for comparison but not an absolute measure. A second judge (DeepSeek V4 Flash, "wanted" tier) is not run.
  • Primary care, English, remote visits, 50 test visits. No specialty, inpatient or multi-party visits (interpreter, family) beyond what PriMock57 has. Rates per specialty would need far more visits than this.
  • Plant authorship. One model family wrote the clean notes and plants; the judge is another. Plants were checked by a validator (one error of each type, anchored to transcript lines), not reviewed by a clinician.
  • No clinician time-and-agreement study.

Verdict

Would a buyer pay? A CMIO or AI governance committee that has deployed a scribe: yes, for this as a monitoring and vendor-comparison tool. On held-out visits it flags every planted error, catches 90-98% of four error types at error severity (72% for changed details, the rest at review severity), and on faithful notes raises a wrong error flag on about 0.35% of sentences and 0.8% of key items. It runs on one GPU inside the hospital, keeps no text, and produces signed per-visit reports and summaries a committee can file. What is missing before it is a product: an eval on real scribe output with clinician-labelled errors; wrong-speaker typing (62% typed correctly); more visits per specialty for per-specialty rates; measurement from audio through the diarizer, where speech errors enter; a second judge to cross-check rates; and a sampling plan (which visits, how many per week) with the consent and contract checks.