01 · Healthcare · live
Visit copilot
Eval results
Scored on a held-out or test splitRun 26 Sep 2026
- Medical-term miss rate, live ASR (Voxtral Mini 4B Realtime)8.4%held outn = 57146 of 1,741 lexicon terms; Nemotron-3.5 streaming 12.7%; WER 13.2 (Nemotron-3.5 streaming 11.8; MOSS-TD pass 2 10.3)
- Speaker labels during the visit, word-level role accuracy (rolling MOSS windows)99.71%held outn = 11PriMock57, 110 min of audio; whole-recording pass 99.97%; worst consultation 98.1%
- Considerations: warning-feature encounter recall (synthetic, held-out)14/15 and 13/15held outn = 15two runs; dev 7/7 both; topic recall held-out 13/16 and 12/16
- Guidance: false-alarm encounters1/16 and 1/16held outn = 16both flags outside the guideline pack (tea-coloured urine on a statin; drowsy driving), shown as 'no bundled guideline'
- Guidance: directive wording shown0held outall runs, the live lint
- Guidance items judged useful at that moment (blind clinician judge)checklist 95/145 (66%); on real GP consultations 20/31 (65%)held outn = 425blind clinician judge: 76 moments in 38 synthetic visits and 24 in 12 PriMock57 consultations; warning-feature items 4/10 and 1/5 useful, the rest already covered; history elements 11% and 18% useful, so they fold away during the visit; 3 of 523 items judged harmful, all synthetic
- Note sentences supported by the human transcript (blind judge, held-out)289/308 (93.8%)held outn = 308PriMock57, 11 new consultations; the 28 Sep writer 229/254 (90.2%) on the same visits; writer-caused misses 12 -> 3; 16 misheard by speech recognition, which a transcript self-check cannot see
- WH-380-E boxes filled right (held-out FMLA visits written blind)91.4% (85/93)held outn = 828 Sep code 79.6% on the same visits; 85.3% vs 75.7% by code scoring alone (free text judged blind otherwise); false fills 10 -> 2. Dev: work/school note 90%, instructions 97%
- Signatures filled by a paperwork draft0held outn = 28signature and signing-date boxes
- Guidance checklist items the GP went on to ask (PriMock57, real GPs)25/43 (58%)held outn = 43history elements 88/107 (82%); what the lane listed that the clinician asked later in the same visit
Dataset
PriMock57 (57 recorded mock primary-care consultations, CC BY 4.0) for speech, speaker labels and note grounding; 38 synthetic US visits with planted warning-feature labels (written and labelled by separate agents) for guidance; 12 synthetic visits with paperwork truth written blind by a separate agent. Judges: Claude Code Opus 5.5 as blind sub-agents.
Caveats
- Synthetic visits and mock consultations, not real clinic audio.
- Two-speaker visits only.
- The judges are models (Claude Code Opus as blind sub-agents), not clinicians.
- The guidance was tested on a 26-topic pack.
- Costs are at list price; eval runs used the direct route (same weights), the e2e runs the gateway.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 29 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 60 s
- Receipts
- 80
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.070
Self-host verification
Verified on 29 Sep 2026: fresh clone of the branch into a clean directory, api image built from docker/api/Dockerfile, compose with a named volume and DECOSA_CLINIC_PROFILE, pointed at the model servers already running on our server (Qwen3.8-27B, Voxtral, MOSS diarizer, M17) instead of starting new ones; then torn down
The image builds and starts; /healthz ok with asr, llm and diarize true. The copilot sample at 2x passed end to end: 28 speaker turns, guidance, a self-checked note (28 sentences, M17 on), codes from the assessment, three paperwork drafts with the clinic profile from the environment, 0 signatures; done 44.5 s after the audio on a GPU shared with our evals. Model-server startup itself was not re-verified.
Rehearsal bundle: clinical.zip (660 KB, 21 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted is for synthetic visits only; real visits must be self-hosted (no BAA yet).
- Speaker labels were tested on two-speaker visits; a third speaker is labelled Other.
- The self-check compares the note with the visit's own transcript, so a word the recogniser misheard passes it. The note marks sentences that may rest on one (where the two recognisers disagree, or a word is unknown): 8 of 17 such sentences on held-out visits, with 8% of good sentences also marked. Check names, numbers and yes/no answers against the cited turn.
- WH-380-E drafts still need every box checked (91% of filled boxes right on held-out visits; vague essential-function wording and incomplete date lists are the usual misses); third-party insurer FMLA forms are not bundled yet.
- No EHR write-back yet: copy the note and the drafts, or use the API.
- Guidance is a 26-topic public-domain pack plus the CMS history elements; a missing item does not mean nothing is missing.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Pass 1: live streaming transcript for the in-visit view (no speakers)Voxtral Mini 4B RealtimeApache-2.0
- Speaker labels during the visit: rolling windows (every 15 s of new audio, 6 s overlap) re-transcribed with speaker labels; the committed turns become the visit transcript the note citesMOSS-Transcribe-Diarize 0.9BApache-2.0
- Language model: live SOAP draft, guidance report, practitioner lanes, window role map, cited note, the self-check's sentence judge, assessment codes and the paperwork field mappingQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
- Note self-check, detail step (M17): each drug, dose, frequency, date, side and number in a note sentence read against its transcript lines; a 'detail not in the visit' flag becomes a changed-detail errordecosa-note-detail-checker-modernbert-large (M17, our own model)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 48 GB card, live pass only (3)
- ACI-Bench ROUGE-L (Qwen3.8-27B FP8, human transcript): 34.2scribe-bench wiki models.md / RESULTS.md
- PriMock57 note composite, vanilla pipeline (streaming ASR -> Qwen3.8-27B, official weights, test 37): 41.01scribe-bench wiki vanilla-vs-best
- Medical-term miss / WER, live ASR (Voxtral Mini 4B Realtime, PriMock57, 57 visits): 8.4% / 13.2 (Nemotron-3.5 streaming: 12.7% / 11.8)scribe-bench asr_score on PriMock57 (57 visits) through the live realtime endpoint; eval results file asr-voxtral-primock57.json (2026-09-23)
Standard · one 96 GB Blackwell card, two passes (4)
- Medical-term miss / WER, pass 2 MOSS-TD vs the Voxtral live pass (PriMock57, 57 visits): 8.4% / 10.3 vs 8.4% / 13.2scribe-bench RESULTS.md (MOSS-TD); eval results file asr-voxtral-primock57.json (Voxtral, 2026-09-23)
- DER / word speaker misattribution, MOSS-TD: 11.4 / 1.0%scribe-bench RESULTS.md, wiki decoder-finding
- Composite, two-pass + role map + Qwen3.8-27B minus vanilla (official weights, test 37): +2.0 [-1.0, +5.7], not resolved; misattributions -0.05scribe-bench wiki vanilla-vs-best (measured with Sortformer + Parakeet as pass 2)
- Verifier recall on injected errors / flags on clean, Qwen3.8-27B judge: 99.1% / 6.5%scribe-bench wiki verifier
Best · adds DeepSeek V4 Flash as the note writer on 2x 96 GB (3)
- ACI-Bench base ROUGE-L / term precision / plan recall: 35.8 / 70.3 / 93 (Qwen3.8-27B ROUGE-L 34.2)scribe-bench wiki models.md
- PriMock57 cited note, official weights (test 37): grounded / term precision: 91.33% / 31.06 (vanilla 88.62% / 26.41)scribe-bench wiki vanilla-vs-best, citations-and-verifiability
- Composite vs vanilla, official weights: +1.25 [-3.76, +6.23], not resolved; plan recall -8.1 [-13.8, -1.7]scribe-bench wiki vanilla-vs-best