145 · Healthcare · preview
Put the visit into your EHR as drafts
Eval results
Not held outRun 29 Sep 2026
- Typed or picked values wrong (read back from the EHR's database)0 / 419syntheticn = 41995% CI 0-0.91%. 32 synthetic visits, clean run 2 (quantity and refills only when said). Run 1: 0 / 534.
- Items entered, of those the rules allow the agent to enter101 / 104syntheticn = 104Misses: 2 inhalers wrongly held by the medicine check, 1 lab order the agent could not finish. Left for you by rule: 14 prescriptions with no quantity said, 2 controlled substances, stops and follow-ups.
- Wrong-patient attempts that wrote to another chart0 / 20syntheticn = 2016 wrong charts open at the start (near-duplicate name, date of birth or MRN; another patient): stopped with 0 entries. 4 chart switches mid-run: caught before the next save.
- Forced sign, transmit, send, fax, e-mail, bill and delete requests held80 / 80syntheticn = 808 kinds x 10 visits, sent from inside the page after the drafts; plus 16 of 16 clicks on the app's own eSign, Transmit Order and Fax buttons held.
- Planted medicine errors flagged by the medicine check90 / 90syntheticn = 90Dose, frequency, drug (sound-alikes) and duration; sound-alike pairs written after the fix: 28 / 28. End to end: 8 of 8 never entered.
- Planted instructions in the chart that caused harm0 / 10syntheticn = 10Encounter reason or an intake note, English and Spanish; 9 of 10 runs paused on the text.
- Correct medicines wrongly held by the medicine check2 / 28syntheticn = 28Both an inhaler with a spoken route.
Dataset
32 synthetic primary-care visits (18 templates, made-up patients): items fixed in code as the ground truth, transcripts written by Qwen3.8-27B around them, on a self-hosted OpenEMR 7.0.3 test instance.
Caveats
- Synthetic visits and transcripts; the model that wrote the transcripts is the model that drives the agent.
- One EHR (OpenEMR). The agent's instructions per form (the EHR profile) were written while testing on the same forms.
- The medicine checker was trained on Qwen-written visits; its numbers on these transcripts are likely optimistic.
- Values were checked against the visit output, not against what a clinician would have wanted: a wrong visit output would be copied faithfully.
- Transcripts were patched twice before the measured runs, both recorded: amounts per dose, and quantities and refills said in half the visits.
- Two blind cold users (a solo physician: maybe / maybe; a practice manager: no / no) read a one-page description; neither used it on real work.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 29 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 90 s
- Receipts
- n/a
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.034
Self-host verification
Not yet verified on a fresh self-host setup.
No rehearsal bundle published yet.
Known limits
- Measured on one EHR only: self-hosted OpenEMR 7.0.3 with made-up patients. A second open-source EHR (OpenMRS O3) was tried and did not start cleanly; not measured.
- Not yet packaged as the browser extension; the engine ran in a test browser against the test EHR.
- A prescription whose quantity was not said goes to your list (OpenEMR requires a quantity, and the agent never guesses one).
- The medicine check wrongly held 2 of 28 correct medicines (an inhaler whose route was said).
- The entries are saved under your login, so the EHR's audit log shows you; each draft's comment or note says it came from ehr-drafts, and the signed record shows every step.
- About 1.5 to 3 minutes per visit on a shared model; no batch mode yet.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- The agent that fills the EHR's forms: one receipted decision per step, values only from the visit, the one draft save released after the chart and form are checked in codeQwen3.8-27B on the Decosa computer-use engineApache-2.0
- Reads each medicine (drug, dose, frequency, route when said, duration) against the transcript lines it came fromdecosa-note-detail-checker (M17)Apache-2.0
- Holds sign, finalise, transmit, send, fax, e-prescribe, bill and delete requests in the browser; releases each draft save oncedecosa-api ehr-drafts and the computer-use network gateAGPL-3.0-or-later
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Standard · one GPU for the model (3)
- Typed and picked values wrong, read back from the EHR's database (32 synthetic visits): 0 / 419decosa-api docs/evals/ehr-drafts.md, clean run 2, 2026-09-29 (run 1: 0 / 534)
- Wrong-patient attempts that wrote to another chart: 0 / 20decosa-api docs/evals/ehr-drafts.md, 2026-09-29
- Planted wrong doses, drugs, frequencies and durations flagged by the medicine check: 90 / 90decosa-api docs/evals/ehr-drafts.md, 2026-09-29