Put the visit into your EHR as drafts
The visit's codes, orders, prescriptions and note entered into your own EHR as unsigned drafts, the chart checked first, each medicine checked against the transcript, and a checklist of what went where; you sign.
Built on: Agent flight recorder, Signed record
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the medicine check and the commit detector run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the put the visit into your ehr as drafts API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real patient data belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- ehr-drafts
Use the hosted API
# Decosa "Put the visit into your EHR as drafts": use the hosted API (made-up patients only)
You are wiring Decosa's EHR-drafts check into this project. It takes a visit's output (diagnosis codes, lab and imaging
orders, medicine changes, the note, the transcript) and returns the review for a given EHR: which items an agent may
enter there as drafts, which go to a copy list in the EHR's field order (each value with its transcript line), and which
are held (a medicine whose dose, frequency, route or drug differs from the transcript, or a controlled substance). The
hosted API never touches an EHR: the agent itself runs in the clinician's own browser. Use only what is listed below. If
you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **Made-up patients only.** Chart data is protected health information; no business associate agreement is signed yet.
Never send real patient information here.
- Nothing here signs, sends, transmits or e-prescribes, and nothing should be built that does.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page, in `DECOSA_API_KEY`, sent as
`Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "ehr-drafts"}` returns `{"token", "expires_at", "budget"}`.
## Endpoints
- `GET /ehr-drafts/info`: what it does and never does.
- `GET /ehr-drafts/gate`: per EHR, whether a tool may type in the clinician's session, whether a write API exists, and
the route (`draft`, `api_or_copy`, `copy`).
- `GET /ehr-drafts/samples`: made-up visits in the input format `decosa.visit-drafts.v1`.
- `POST /ehr-drafts/check` (token). Body: `{"sample_id": "wrong-dose-planted", "ehr": "openemr"}` or
`{"visit": {...}, "ehr": "therapynotes"}`. A visit: `patient {name, dob YYYY-MM-DD, mrn?}`, `encounter {date}`,
`transcript [{n, role, text}]`, `items [{id, kind: problem|order|med_new|med_change|med_stop|note|follow_up, lines: [n],
...}]`. Answer: `{review: {headline, counts, items: [{id, kind, label, status, group: ready|check|yours, where, why, lines,
source, flags}], copy: {title, note, items: [{label, value, cite, quote}]}, still_active, never, agent_may_type}, ehr,
medcheck, receipt}`.
- `POST /ehr-drafts/ack` (token) `{panel: <review>, person: {name, role}}`: the clinician's signed acknowledgement of the
review (not an EHR signature).
## Build
Show the review grouped: "check first" (held medicines, with the transcript line), "ready as drafts", "for you to enter"
(the copy list with a copy button per value). Show the headline verbatim. Keep the transcript line next to every value.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa "Put the visit into your EHR as drafts": self-host
You are setting up Decosa's EHR-drafts tool on this practice's own machine. Use only what is listed below. If something
is missing, stop and ask me.
What it is: decosa-api's `ehr-drafts` module (`decosa_api/verticals/ehrdrafts`) on the shared computer-use engine
(`decosa_api/cu`), the note-detail checker (a CPU model, `services/detail_checker`), and Qwen3.8-27B served by vLLM. The
agent drives a tab of the clinician's own browser (Chrome DevTools Protocol), types the visit's items into the open chart
as drafts, and stops before sign. It types only into EHRs whose terms allow a tool in the session (today: a self-hosted
OpenEMR); for other EHRs, use `POST /ehr-drafts/check` for the copy list.
1. Check the GPU (`nvidia-smi`; about 20 GB for Qwen3.8-27B NVFP4) and Docker with the NVIDIA container toolkit.
2. Serve the model with vLLM (image `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`)
on `127.0.0.1:8114`, OpenAI-compatible.
3. Clone decosa-api (ask me for access), create a venv, `pip install -e '.[flight]'` and `playwright install chromium`.
4. Start the note-detail checker: `python services/detail_checker/server.py` on `127.0.0.1:8495` (CPU).
5. Start the API: `DECOSA_HOST=127.0.0.1 DECOSA_PORT=8445 DECOSA_LLM_ROUTE=direct DECOSA_DETAIL_URL=http://127.0.0.1:8495 python -m decosa_api`.
6. Smoke test: `GET /ehr-drafts/samples`, then `POST /ehr-drafts/check {"sample_id": "wrong-dose-planted", "ehr": "openemr"}`:
the metformin item must be in the "check" group with the transcript line that says 1000 mg.
7. The agent: `decosa_api.verticals.ehrdrafts.runner.run_visit(visit, OpenEMR(<your OpenEMR URL>), chat, run, emit,
surface=CDPSurface(<Chrome's DevTools endpoint>, <your OpenEMR URL>))` with the chart open in that tab. Read
`runner.py` first: it checks the chart, the medicines, and releases each draft save once; sign, transmit, send and
e-prescribe are held.
Never point it at an EHR whose terms forbid tools in the session, and never at real patients before your own review of
the rails on a test copy.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
(no mock bundle is published for this tool yet, so use the synthetic sample from the smoke test below),
show me what is in it, and run the rehearsal against the local API:
the smoke test from the steps below, on that synthetic sample only.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property. Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B on the Decosa computer-use engine needs a GPU.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B on the Decosa computer-use engine with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). Qwen3.8-27B on the Decosa computer-use engine needs about 32 GB at its smallest setting. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B on the Decosa computer-use engine: run it at its smallest setting (about 32 GB instead of 33.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).
- L40Sstandard tierRuns
The standard tier fits (33.6 of 48 GB).
- H100 80 GB (SXM)standard tierRuns
The standard tier fits (33.6 of 80 GB).
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (33.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (33.6 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B on the Decosa computer-use engine with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B on the Decosa computer-use engine with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"ehr-drafts"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py ehr-drafts
No mock-data bundle is published for this tool yet. Rehearse with the synthetic sample the prompt's smoke test uses.
Prompt for your coding agent
# Decosa "Put the visit into your EHR as drafts": self-host
You are setting up Decosa's EHR-drafts tool on this practice's own machine. Use only what is listed below. If something
is missing, stop and ask me.
What it is: decosa-api's `ehr-drafts` module (`decosa_api/verticals/ehrdrafts`) on the shared computer-use engine
(`decosa_api/cu`), the note-detail checker (a CPU model, `services/detail_checker`), and Qwen3.8-27B served by vLLM. The
agent drives a tab of the clinician's own browser (Chrome DevTools Protocol), types the visit's items into the open chart
as drafts, and stops before sign. It types only into EHRs whose terms allow a tool in the session (today: a self-hosted
OpenEMR); for other EHRs, use `POST /ehr-drafts/check` for the copy list.
1. Check the GPU (`nvidia-smi`; about 20 GB for Qwen3.8-27B NVFP4) and Docker with the NVIDIA container toolkit.
2. Serve the model with vLLM (image `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`)
on `127.0.0.1:8114`, OpenAI-compatible.
3. Clone decosa-api (ask me for access), create a venv, `pip install -e '.[flight]'` and `playwright install chromium`.
4. Start the note-detail checker: `python services/detail_checker/server.py` on `127.0.0.1:8495` (CPU).
5. Start the API: `DECOSA_HOST=127.0.0.1 DECOSA_PORT=8445 DECOSA_LLM_ROUTE=direct DECOSA_DETAIL_URL=http://127.0.0.1:8495 python -m decosa_api`.
6. Smoke test: `GET /ehr-drafts/samples`, then `POST /ehr-drafts/check {"sample_id": "wrong-dose-planted", "ehr": "openemr"}`:
the metformin item must be in the "check" group with the transcript line that says 1000 mg.
7. The agent: `decosa_api.verticals.ehrdrafts.runner.run_visit(visit, OpenEMR(<your OpenEMR URL>), chat, run, emit,
surface=CDPSurface(<Chrome's DevTools endpoint>, <your OpenEMR URL>))` with the chart open in that tab. Read
`runner.py` first: it checks the chart, the medicines, and releases each draft save once; sign, transmit, send and
e-prescribe are held.
Never point it at an EHR whose terms forbid tools in the session, and never at real patients before your own review of
the rails on a test copy.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
(no mock bundle is published for this tool yet, so use the synthetic sample from the smoke test below),
show me what is in it, and run the rehearsal against the local API:
the smoke test from the steps below, on that synthetic sample only.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property. Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsPut the visit into your EHR as drafts on GeForce RTX 5090: use the Standard · one GPU for the model tier
The standard tier fits with changes: Qwen3.8-27B on the Decosa computer-use engine: run it at its smallest setting (about 32 GB instead of 33.6 GB), with a shorter context and fewer parallel sessions.
Standard · one GPU for the model: what changesuses estimates
- Qwen3.8-27B on the Decosa computer-use engine: run it at its smallest setting (about 32 GB instead of 33.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- The agent that fills the EHR's forms: Qwen3.8-27B on the Decosa computer-use engine. ~33.6 GB (at least ~32 GB), weights 29 GB (from stack.json). Qwen3.8-27B FP8: 33.6 GB is the sales lite tier's allotment (stack.json). Weights of about 29 GB are an estimate (27.8B parameters at one byte, plus higher-precision layers). (stack.json lists 20 GB for this component.)
- Reads each medicine: decosa-note-detail-checker (M17). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Holds sign, finalise, transmit, send, fax, e-...: decosa-api ehr-drafts and the computer-use network gate. CPU. Runs on CPU (vram_gb 0 in stack.json).
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Put the visit into your EHR as drafts, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Put the visit into your EHR as drafts on my hardware Fetch https://decosa.ai/prompts/ehr-drafts-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=ehr-drafts) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · one GPU for the model (standard). Fit check: runs with changes, about 32 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - The agent that fills the EHR's forms: Qwen3.8-27B on the Decosa computer-use engine (Qwen/Qwen3.8-27B-FP8), 33.6 GB. Change: Qwen3.8-27B on the Decosa computer-use engine: run it at its smallest setting (about 32 GB instead of 33.6 GB), with a shorter context and fewer parallel sessions. - Reads each medicine: decosa-note-detail-checker (M17) (decosaai/decosa-note-detail-checker-modernbert-large), CPU - Holds sign, finalise, transmit, send, fax, e-...: decosa-api ehr-drafts and the computer-use network gate, CPU GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B on the Decosa computer-use engine ~32 GB (100%); about 0 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/ehr-drafts-assemble.md
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 29 Sep 2026 · measured 29 Sep 2026: · p50 90 s · p95 163 s (32 runs) · ~$0.034 per run
Loading the nightly status…
Self-host: not yet verified
Measured cost to run: about $0.34 per 10 visits (hosted, 29 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
Known limits (6)
- Measured on one EHR only: self-hosted OpenEMR 7.0.3 with made-up patients. A second open-source EHR (OpenMRS O3) was tried and did not start cleanly; not measured.
- Not yet packaged as the browser extension; the engine ran in a test browser against the test EHR.
- A prescription whose quantity was not said goes to your list (OpenEMR requires a quantity, and the agent never guesses one).
- The medicine check wrongly held 2 of 28 correct medicines (an inhaler whose route was said).
- The entries are saved under your login, so the EHR's audit log shows you; each draft's comment or note says it came from ehr-drafts, and the signed record shows every step.
- About 1.5 to 3 minutes per visit on a shared model; no batch mode yet.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the medicine check and the commit detector run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the put the visit into your ehr as drafts API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real patient data belongs on your own hardware.
Works today with self-hosted OpenEMR: the visit's codes, orders, prescriptions and note typed in as unsigned drafts. You sign.
For clinicians in small practices who re-type every visit into the EHR. An agent in your own browser enters what the visit produced (ICD-10-CM problems, lab and imaging orders, new prescriptions and dose changes, the SOAP note) into the chart you have open as drafts: problems unconfirmed, prescriptions saved but not sent, orders saved but not transmitted, the note unsigned. It reads the patient's name and date of birth off the screen before anything and before every save, reads each medicine against the transcript and holds any that differ, types a quantity or refill count only if it was said, never drafts a controlled substance, and stops before sign: signing, transmitting, sending and e-prescribing are held in the browser's network layer whatever a button says. It types only where the EHR's terms allow a tool in your session (today: self-hosted OpenEMR); every other EHR gets the same checklist as a copy list.
- Deployment
- Self-host first
- Regulatory
- Checked 28 Sep 2026; not legal advice. FDA: suggested codes and orders a clinician can review, each with the transcript line behind it and nothing time-critical, fit the non-device clinical decision support criteria (FD&C Act 520(o)(1)(E); FDA guidance of 29 Jan 2026). DEA: controlled-substance e-prescriptions are signed only by the prescriber through their own two-factor authentication (21 CFR 1311.135, 1311.140); this tool never drafts one and never touches a sign screen. HIPAA: hosting or relaying chart data makes Decosa a business associate (45 CFR 160.103); no BAA is signed yet, so the hosted demo is synthetic only and real use runs the model in the practice's network. EHR terms: eClinicalWorks, DrChrono and TherapyNotes forbid tools in the customer's session in their own terms (read 28 Sep 2026); Practice Fusion, Tebra and SimplePractice restrict it or could not be read; those EHRs get the copy list. CPT codes are not shown or typed (AMA licence pending); diagnoses are ICD-10-CM and tests carry LOINC codes. Texas SB 1188 requires telling patients when AI is used for diagnosis or treatment recommendations.
At a glance
- What it gives you
- Drafts in your EHR (unsigned), a checklist of what went where with the transcript line behind each value, a copy list for what it could not enter, every request the gate held, and a signed record of the run.
- What it never does
- Sign, finalise, transmit an order, send, fax, e-mail or e-prescribe; draft a controlled substance; type a value the visit did not give (a quantity or refill count nobody said stays blank); edit or delete an existing entry.
- Where it may type
- Works today with self-hosted OpenEMR (open source). Commercial EHRs (eClinicalWorks, DrChrono, athenaOne and others) come through each vendor's official API, not yet available; until then they get the copy list. The hosted demo uses made-up patients only.
- Data retention
- The hosted check keeps nothing: the visit lives in memory for the request, logs carry counts only. The agent's signed record holds hashes of the patient identifiers, never their text, and stays with the practice.
- What leaves the box
- Self-hosted, nothing: the agent, the model and the checker run in your network. Our hosted model would see chart screens, so it waits for a BAA.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Standard
one GPU for the model
Qwen3.8-27B drives the forms; the medicine checker runs on CPU. What the eval measured.
- Models
- Qwen3.8-27B on the Decosa computer-use engine
- decosa-note-detail-checker (M17)
- decosa-api ehr-drafts and the computer-use network gate
- Hardware
- 1x RTX PRO 6000 96 GB (measured)
- Quality evidence
- Typed and picked values wrong, read back from the EHR's database (32 synthetic visits)0 / 419decosa-api docs/evals/ehr-drafts.md, clean run 2, 2026-09-29 (run 1: 0 / 534)
- Wrong-patient attempts that wrote to another chart0 / 20decosa-api docs/evals/ehr-drafts.md, 2026-09-29
- Planted wrong doses, drugs, frequencies and durations flagged by the medicine check90 / 90decosa-api docs/evals/ehr-drafts.md, 2026-09-29
- Verification
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
The agent that fills the EHR's forms: one receipted decision per step, values only from the visit, the one draft save released after the chart and form are checked in codeQwen3.8-27B on the Decosa computer-use engineQwen/Qwen3.8-27B-FP8 on Hugging Face (opens in a new tab) 27B · 20 GBNo proof yet | Standard | 27B · 20 GB | No proof yet | |
| ||||
Reads each medicine (drug, dose, frequency, route when said, duration) against the transcript lines it came fromdecosa-note-detail-checker (M17)decosaai/decosa-note-detail-checker-modernbert-large on Hugging Face (opens in a new tab) 0.4B · 0 GBProof: partial | Standard | 0.4B · 0 GB | Proof: partial | |
| ||||
Holds sign, finalise, transmit, send, fax, e-prescribe, bill and delete requests in the browser; releases each draft save oncedecosa-api ehr-drafts and the computer-use network gate 0 GBProof: partial | Standard | 0 GB | Proof: partial | |
| ||||
Tools, services and hardware
Tools
- ICD-10-CM 2026 code set (CMS/NCHS order file) (opens in a new tab)US government publication
Diagnosis codes on the problem list; the test EHR's code picker.
- LOINC (opens in a new tab)LOINC licence (free, with notice)
Codes for the lab and imaging tests; names shown in our own words.
- DEA controlled substance schedules (21 CFR 1308) (opens in a new tab)US government publication
Controlled substances are never drafted.
- RxNorm Current Prescribable Content (NLM) (opens in a new tab)Public; no licence required
Every medicine name is resolved to its ingredients (brand, salt and combination names), from a snapshot bundled with the code. A name that resolves to no known drug is held, not drafted.
- openFDA NDC directory (FDA) (opens in a new tab)Public domain (CC0)
The DEA schedule of each ingredient: a medicine is held when any ingredient is scheduled.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /ehr-drafts/info, /ehr-drafts/gate, /ehr-drafts/samples; POST /ehr-drafts/check (the medicine check and the review for your EHR; touches no EHR), POST /ehr-drafts/ack (your signed acknowledgement).
- vLLM (model):8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B for the agent, in your network (self-host) or through our gateway once a BAA is signed.
- note-detail checker:8495
The medicine check, on CPU: a small FastAPI service in decosa-api (services/detail_checker, systemd unit decosa-detail-checker); no published image yet.
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured: the eval ran the agent against the model on our server, shared with other workloads.
- Any laptop or desktop for the browser side Fits
The agent drives a browser tab; the medicine checker runs on CPU.
Latency per lane
- one visit (4-5 items) entered by the agent on the synthetic OpenEMR, 3 visits in parallel on a shared model90.3 s
Measuredmeasured on our server 2026-09-29: p50 over 32 visits (p95 162.9 s), direct route; 188.6 s p50 in an earlier run under heavier load
- the hosted check (medicine check and the review for your EHR)2.0 s
Measuredmeasured on our server 2026-09-29: smoke module on the pre-release server
The assembly prompt for this stack is coming soon. The Self-host tab has the general steps.
Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- Works today with self-hosted OpenEMR: the visit's codes, orders, prescriptions and note typed in as unsigned drafts. You sign.
- Who it's for
- Clinicians and practice staff in small practices who re-type each visit into the EHR.
- Where it runs
- Works today with self-hosted OpenEMR, in the clinician's own browser; commercial EHRs through each vendor's official API (not yet available). Hosted demo: made-up patients only
- Key numbers
- 0 / 419 Typed or picked values wrong (read back from the EHR's database) (synthetic, n = 419)
- 101 / 104 Items entered, of those the rules allow the agent to enter (synthetic, n = 104)
- 0 / 20 Wrong-patient attempts that wrote to another chart (synthetic, n = 20)
- 90.3 s Median end-to-end run, hosted (QA sweep 2026-09-29)
- Models
- Qwen3.8-27B (drives the forms) · decosa-note-detail-checker (M17, reads each medicine against the transcript, CPU) · decosa-commit-detector (flags buttons that commit, CPU)
- Where
- Works today with self-hosted OpenEMR, in the clinician's own browser; commercial EHRs through each vendor's official API (not yet available). Hosted demo: made-up patients only
- Checks
- Chart identifiers read off the screen before any entry and before every save; every typed value from the visit output; the form compared with the visit in code before its one save; sign/send/transmit held at the network layer; signed flight record of every step
- Industry
- Healthcare
- Input
- Text and documents
- Runs
- Self-host
- Output
- Notes, reports and drafts · Signed record or verdict
- Data
- Patient data (PHI)
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Self-host
- Built from
- Agent flight recorder · Signed record
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Put the visit into your EHR as drafts
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…