Skip to content
decosa

145 · Healthcare · preview

Put the visit into your EHR as drafts

Open the toolJSON

Eval results

Not held outRun 29 Sep 2026

  • Typed or picked values wrong (read back from the EHR's database)0 / 419syntheticn = 41995% CI 0-0.91%. 32 synthetic visits, clean run 2 (quantity and refills only when said). Run 1: 0 / 534.
  • Items entered, of those the rules allow the agent to enter101 / 104syntheticn = 104Misses: 2 inhalers wrongly held by the medicine check, 1 lab order the agent could not finish. Left for you by rule: 14 prescriptions with no quantity said, 2 controlled substances, stops and follow-ups.
  • Wrong-patient attempts that wrote to another chart0 / 20syntheticn = 2016 wrong charts open at the start (near-duplicate name, date of birth or MRN; another patient): stopped with 0 entries. 4 chart switches mid-run: caught before the next save.
  • Forced sign, transmit, send, fax, e-mail, bill and delete requests held80 / 80syntheticn = 808 kinds x 10 visits, sent from inside the page after the drafts; plus 16 of 16 clicks on the app's own eSign, Transmit Order and Fax buttons held.
  • Planted medicine errors flagged by the medicine check90 / 90syntheticn = 90Dose, frequency, drug (sound-alikes) and duration; sound-alike pairs written after the fix: 28 / 28. End to end: 8 of 8 never entered.
  • Planted instructions in the chart that caused harm0 / 10syntheticn = 10Encounter reason or an intake note, English and Spanish; 9 of 10 runs paused on the text.
  • Correct medicines wrongly held by the medicine check2 / 28syntheticn = 28Both an inhaler with a spoken route.

Dataset

32 synthetic primary-care visits (18 templates, made-up patients): items fixed in code as the ground truth, transcripts written by Qwen3.8-27B around them, on a self-hosted OpenEMR 7.0.3 test instance.

Caveats

  • Synthetic visits and transcripts; the model that wrote the transcripts is the model that drives the agent.
  • One EHR (OpenEMR). The agent's instructions per form (the EHR profile) were written while testing on the same forms.
  • The medicine checker was trained on Qwen-written visits; its numbers on these transcripts are likely optimistic.
  • Values were checked against the visit output, not against what a clinician would have wanted: a wrong visit output would be copied faithfully.
  • Transcripts were patched twice before the measured runs, both recorded: amounts per dose, and quantities and refills said in half the visits.
  • Two blind cold users (a solo physician: maybe / maybe; a practice manager: no / no) read a one-page description; neither used it on real work.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
29 Sep 2026
Latency, this run
n/a
p50 over passed runs
90 s
Receipts
n/a
Model calls
n/a
Tokens
n/a
Cost per run
$0.034

Self-host verification

Not yet verified on a fresh self-host setup.

No rehearsal bundle published yet.

Known limits

  • Measured on one EHR only: self-hosted OpenEMR 7.0.3 with made-up patients. A second open-source EHR (OpenMRS O3) was tried and did not start cleanly; not measured.
  • Not yet packaged as the browser extension; the engine ran in a test browser against the test EHR.
  • A prescription whose quantity was not said goes to your list (OpenEMR requires a quantity, and the agent never guesses one).
  • The medicine check wrongly held 2 of 28 correct medicines (an inhaler whose route was said).
  • The entries are saved under your login, so the EHR's audit log shows you; each draft's comment or note says it came from ehr-drafts, and the signed record shows every step.
  • About 1.5 to 3 minutes per visit on a shared model; no batch mode yet.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • The agent that fills the EHR's forms: one receipted decision per step, values only from the visit, the one draft save released after the chart and form are checked in codeQwen3.8-27B on the Decosa computer-use engineApache-2.0
  • Reads each medicine (drug, dose, frequency, route when said, duration) against the transcript lines it came fromdecosa-note-detail-checker (M17)Apache-2.0
  • Holds sign, finalise, transmit, send, fax, e-prescribe, bill and delete requests in the browser; releases each draft save oncedecosa-api ehr-drafts and the computer-use network gateAGPL-3.0-or-later

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Standard · one GPU for the model (3)
  • Typed and picked values wrong, read back from the EHR's database (32 synthetic visits): 0 / 419decosa-api docs/evals/ehr-drafts.md, clean run 2, 2026-09-29 (run 1: 0 / 534)
  • Wrong-patient attempts that wrote to another chart: 0 / 20decosa-api docs/evals/ehr-drafts.md, 2026-09-29
  • Planted wrong doses, drugs, frequencies and durations flagged by the medicine check: 90 / 90decosa-api docs/evals/ehr-drafts.md, 2026-09-29

How we measure · All tools