Skip to content
decosa

Run the visit with a copilot

One screen for the whole visit: who said what, what is still worth asking with its source, a note checked sentence by sentence before you see it, codes and prior-auth criteria, and the work note, FMLA sections and patient instructions drafted, never signed or sent. On your own GPU for real patients.

60 s from the end of the visit to the checked note and paperwork (median, measured 2026-09-29)~$0.080 per visitDetails, API and self-host

  • Transcribes, with who said what

  • Guides while the patient is there

  • Writes a note it has checked

  • Codes and does the paperwork

On your own GPU for real patients (no patient data leaves the practice). The hosted demo takes made-up visits only, and nothing from a demo visit is kept.

Demo only: synthetic visits. Do not say real patient information. Real patients run on your own hardware.

Back injury at work: sciatica, a work note and FMLA (synthetic US visit, 2:54). Recorded from the live pipeline; it plays at 4x, with everything appearing when it did.

0:00 / 3:56

Guidance

For the clinician: things to consider covering, each with what triggered it and its source. Not a diagnosis or an instruction.

Listening. Guidance appears once the reason for the visit is clear.

Proof · signed records appear as each step finishes

Each step is signed: which model ran, and a fingerprint of what went in and came out, so it can be checked later.

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB), or 2× RTX 5090.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the visit copilot API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
clinical

Use the hosted API

# Decosa Clinical scribe: use the hosted API

You are adding Decosa's Clinical scribe to this project. Decosa runs open models (Voxtral Mini 4B Realtime for speech,
Qwen3.8-27B for text) through the Decosa API. Every model output comes with a signed receipt.
Use only the endpoints below. If you need something that is not listed, stop and ask me; do not guess endpoints.

- Base URL: `https://api.decosa.ai`
- WebSocket base: `wss://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz` returns `{"ok": true, "asr": bool, "llm": bool, "live_sessions": n, "queue": n}`.

## Auth: API key (or a demo session)
1. Preferred: an API key. Create one on the tool page with "Get an API key"; it looks like `dk_…` and is shown
   only once. Keep it in an environment variable, never in code: `DECOSA_API_KEY=dk_…`. Send
   `Authorization: Bearer $DECOSA_API_KEY` on calls that need auth; WebSockets take `?token=$DECOSA_API_KEY` in the URL.
2. Without a key, use a short demo session: `POST https://api.decosa.ai/demo/session` with JSON `{"vertical": "clinical"}` returns
   `{"token": "<opaque>", "expires_at": <unix seconds>, "budget": {"seconds_audio": 300, "llm_tokens": 20000}}`.
3. Send `Authorization: Bearer <token>` on calls that need it (session-bound calls such as live audio, replay, chat and
   render jobs). WebSockets take `?token=<token>` in the URL instead. These need no token: `GET /healthz`,
   `GET /demo/scripts`, `GET /demo/recordings`, `GET /demo/recordings/{id}/events`, `GET /studio/gallery`,
   `GET /studio/jobs/{id}`, `GET /receipts/{id}`.
4. Demo-session limits: a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz), and a global cap on concurrent live audio sessions. Over a limit the API answers
   HTTP 429 with a `Retry-After` header (seconds): wait that long, then retry. Reuse a token until `expires_at`.
5. The API keeps no PII; session transcripts live in memory and are deleted when the session ends.

## Live audio
`WS wss://api.decosa.ai/ws/live?vertical=clinical&token=<token>`

Client to server:
- binary frames: 16 kHz mono PCM16 little-endian, about 100 ms each (1600 samples, 3200 bytes);
- a text frame `{"type":"stop"}` when the speaker is done. The server then writes final lanes, sends `done` and closes.

Server to client, JSON text frames:
- `{"type":"ready","vertical":"clinical","models":{"asr":"...","llm":"..."}}`
- `{"type":"transcript","t":12.4,"text":"...","final":true}` (partials have `final:false` and are replaced by the next update)
- `{"type":"lane","lane":"<lane id>","title":"<Title>","body":"<markdown or text>","data":{...optional...},"latency_ms":850}`
  Each lane event replaces the previous content of that lane.
- `{"type":"receipt","id":"<completion id>","model":"qwen3.8-27b","gateway_sig":"<hex>","provider":"<provider id>"}`
  
- `{"type":"budget","seconds_audio_left":240,"llm_tokens_left":15000}`
- `{"type":"error","message":"..."}`
- `{"type":"done","summary":{...final artifact...}}`

Lanes for `clinical`:
- `note`: SOAP draft note, updated during the visit
- `codes`: ICD-10-CM and RxNorm codes, validated; `data` holds the structured codes
- `priorauth`: prior-authorization criteria met and what to ask next
- `visit_level`: suggested E/M visit level from medical decision making

Final artifact in `done.summary`: the note plus codes.

Close codes the server uses:
- `4401` bad or expired token: get a new session, then reconnect.
- `4409` token already has a live socket: close the other one, or get a new session.
- `4429` busy (live-session cap): wait, then retry. Back off at least 15 s.
- `4400` bad parameters, `4402` budget exhausted, `4403` vertical or origin mismatch: do not retry; fix the cause.

On any other drop before `done`, reconnect with backoff (0.5 s, 1 s, 2 s, ... up to 8 s) using the same token while it
is valid. Stop sending audio when `seconds_audio_left` reaches 0.

## Testing without a microphone
`POST https://api.decosa.ai/demo/replay` with `Authorization: Bearer <token>` and JSON
`{"vertical":"clinical","script_id":"<script id>"}` streams the same event types over SSE (`text/event-stream`, one JSON
event per `data:` line). It runs a canned audio script through the real pipeline, so the output is live.
Script ids for `clinical`: `clinical-back-pain`, `clinical-diabetes-followup`. `GET https://api.decosa.ai/demo/scripts` lists all scripts (no token).

## Receipts (public, read-only)
`GET https://api.decosa.ai/receipts/{id}` returns
`{"id", "model", "weights_root", "request_hash", "output_hash", "provider": {"miner_id", "pubkey", "sig"}, "gateway": {"pubkey", "sig"}, "proof": {"format", "verified"}, "checks": [{"name", "ok", "detail"}]}`.

## Patient and client data
The hosted API is a public demo. It must not receive PHI or any real patient or client information. For real
data, self-host (see the "Run it yourself" prompt) so audio and text never leave the site.

## Audio format
Convert a recording to the wire format with ffmpeg:
`ffmpeg -i input.wav -ar 16000 -ac 1 -f s16le input.pcm`

## Example: TypeScript (Node 22+, global fetch and WebSocket)
```ts
import { readFileSync } from "node:fs";

const BASE = "https://api.decosa.ai";
// Prefer your API key (dk_…); fall back to a short demo session.
let token = process.env.DECOSA_API_KEY;
if (!token) {
  const session = await fetch(`${BASE}/demo/session`, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ vertical: "clinical" }),
  });
  if (session.status === 429) throw new Error(`busy, retry after ${session.headers.get("Retry-After")} s`);
  ({ token } = await session.json());
}

const ws = new WebSocket(
  `${BASE.replace(/^http/, "ws")}/ws/live?vertical=clinical&token=${encodeURIComponent(token)}`,
);
ws.binaryType = "arraybuffer";
ws.onmessage = (m) => {
  const ev = JSON.parse(String(m.data));
  if (ev.type === "transcript" && ev.final) console.log("heard:", ev.text);
  if (ev.type === "lane") console.log(`[${ev.lane}]`, ev.body);
  if (ev.type === "receipt") console.log("receipt:", `${BASE}/receipts/${ev.id}`);
  if (ev.type === "done") { console.log("final:", ev.summary); ws.close(); }
};
ws.onopen = async () => {
  const pcm = readFileSync("input.pcm"); // 16 kHz mono PCM16 LE
  for (let i = 0; i < pcm.length; i += 3200) {
    ws.send(pcm.subarray(i, i + 3200));
    await new Promise((r) => setTimeout(r, 100)); // real time
  }
  ws.send(JSON.stringify({ type: "stop" }));
};
```

## Example: Python (`pip install requests websockets`)
```python
import asyncio, json, os, requests, websockets

BASE = "https://api.decosa.ai"
token = os.environ.get("DECOSA_API_KEY")  # dk_… API key preferred
if not token:
    r = requests.post(f"{BASE}/demo/session", json={"vertical": "clinical"})
    if r.status_code == 429:
        raise SystemExit(f"busy, retry after {r.headers.get('Retry-After')} s")
    token = r.json()["token"]

async def main():
    url = BASE.replace("http", "ws", 1) + f"/ws/live?vertical=clinical&token={token}"
    async with websockets.connect(url) as ws:
        async def send_audio():
            with open("input.pcm", "rb") as f:  # 16 kHz mono PCM16 LE
                while chunk := f.read(3200):
                    await ws.send(chunk)
                    await asyncio.sleep(0.1)
            await ws.send(json.dumps({"type": "stop"}))
        sender = asyncio.create_task(send_audio())
        async for msg in ws:
            ev = json.loads(msg)
            if ev["type"] == "lane":
                print(f"[{ev['lane']}]", ev["body"])
            elif ev["type"] == "receipt":
                print("receipt:", f"{BASE}/receipts/{ev['id']}")
            elif ev["type"] == "done":
                print("final:", ev["summary"])
                break
        await sender

asyncio.run(main())
```

## What to build
1. A small client module for the calls above (session, socket, reconnect, 429 handling).
2. UI or CLI output that shows the transcript, each lane, and a link to every receipt.
3. A config value for the base URL, so it can point at a self-hosted box later with no code change.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Clinical scribe: run it yourself (containers)

You are setting up Decosa Clinical scribe to run entirely on this machine's NVIDIA GPU(s). Nothing is sent to Decosa's
hosted API and there are no Decosa charges. The local service speaks the same API as the hosted one, so apps built
against the hosted API only need a new base URL.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/clinical.zip (660 KB, 21 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py clinical` (the api image carries the same bundle under /app/rehearsal/clinical/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py clinical --bundle clinical.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the captions carry the complaint and the MRI order"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
   Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
   NVIDIA's official instructions (docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html),
   then run `sudo nvidia-ctk runtime configure --runtime=docker` and `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. If it references an `.env` file or variables, ask me for any values. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health. Find the host port that compose.yaml publishes for the API (`docker compose ps`), then poll
   `curl -fsS http://localhost:<PORT>/healthz` every 10 s until it returns `"ok": true` with `"asr": true` and
   `"llm": true`. The first start can take a while as weights download. Show me `docker compose logs --tail=50` if it
   has not come up after 20 minutes.
7. Smoke test: `curl -fsS -X POST http://localhost:<PORT>/demo/session -H 'Content-Type: application/json' -d '{"vertical":"clinical"}'`
   should return a token.
8. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.

## Patient data stays here
This box will handle PHI. Keep it that way:
- Do not expose the service to the internet.Reach it only from the clinic network or a private VPN.
- Put TLS in front of it if other machines on the network use it.
- Keep disk encryption on, and keep the machine patched.
- Clinical workloads stay local-only.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/clinical-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Self-host: the recommended way to run the scribe

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

Why patient data should stay on site

  • Visit audio and notes are protected health information. Running the scribe on a GPU in the clinic means that audio and text never cross the internet to a vendor.
  • The hosted demo keeps sessions in memory and deletes them when they end, but it is a public demo. It must not receive PHI.
  • A local box keeps working when the internet link is down, and latency stays low.
  • A self-hosted box signs its own receipts with a key it generates on first start: each note carries the model name, its weights root and the hashes of what went in and came out, without sending anything anywhere. That is an attestation by the clinic’s own box, not a proof that the model ran.
  • CPU only, 64 GB RAMDoesn't fit

    Voxtral Mini 4B Realtime needs a GPU.

  • GeForce RTX 4090Doesn't fit

    Needs about 40 GB of GPU memory at the smallest settings; 24 GB available.

  • GeForce RTX 5090Doesn't fit

    Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions.

  • L40Slite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 53.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (85.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (85.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (48 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (48 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API, the speech model and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with both models loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "asr": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"clinical"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py clinical

Download the mock-data bundle (660 KB, 21 checks)expected.json

Two cuts from a synthetic primary-care visit (the patient's complaint of eight weeks of back pain shooting down the right leg, and the doctor's assessment: S1 radiculopathy, MRI of the lumbar spine ordered), streamed over the live WebSocket as if from a microphone. The session must produce captions and a SOAP note whose claims are checked against the transcript, with validated ICD-10 codes for lumbar radiculopathy and a visit level. Then, without the speech model, the clinical considerations lane runs on the typed transcript of a second synthetic visit (chest burning after meals, and tightness on the stairs that goes up into the jaw, which the clinician puts down to reflux): it must show the cardiac warning feature with the patient's own words and a cited CDC source, carry the 'not a diagnosis' label with no directive wording, and one accept decision must seal a signed record that verifies and that the clinical AI monitor's audit finds conforming.

What the rehearsal checks
  • the session ends with a done event
  • the session reports no errors
  • the captions carry the complaint and the MRI order
  • the final SOAP note names the radiculopathy and the MRI
  • the note is self-checked before it is final: at least 3 sentences checked and supported by the transcript
  • the transcript comes back with speaker labels: the clinician
  • and the patient
  • the final note is marked self-checked
  • the paperwork drafts never fill a signature
  • a validated ICD-10 code for radiculopathy is suggested
  • a visit level (E/M) is proposed from the problems addressed
  • the live session also returns the clinical considerations lane with its label
  • the chest-pain visit shows the cardiac warning feature as a consideration
  • the cardiac item quotes the patient's words about the jaw
  • the cardiac item cites a CDC or NHLBI page
  • everything the lane wrote passes the directive-language check
  • the output is for the clinician only and is never inserted into the note
  • the accept decision seals a record that verifies
  • the record verifier accepts it
  • the clinical AI monitor's audit finds the output conforming (quotes, citations, wording, labels, signature)
  • every speech and model call has a signed receipt

Licence: Synthetic: a script written for Decosa (no real people, patients or companies) read by Decosa house voices (Kokoro-82M stock voicepacks, Apache-2.0), each allowed by the consent ledger for project decosa-clinical-demo. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa Clinical scribe: run it yourself (containers)

You are setting up Decosa Clinical scribe to run entirely on this machine's NVIDIA GPU(s). Nothing is sent to Decosa's
hosted API and there are no Decosa charges. The local service speaks the same API as the hosted one, so apps built
against the hosted API only need a new base URL.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/clinical.zip (660 KB, 21 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py clinical` (the api image carries the same bundle under /app/rehearsal/clinical/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py clinical --bundle clinical.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the captions carry the complaint and the MRI order"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
   Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
   NVIDIA's official instructions (docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html),
   then run `sudo nvidia-ctk runtime configure --runtime=docker` and `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. If it references an `.env` file or variables, ask me for any values. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health. Find the host port that compose.yaml publishes for the API (`docker compose ps`), then poll
   `curl -fsS http://localhost:<PORT>/healthz` every 10 s until it returns `"ok": true` with `"asr": true` and
   `"llm": true`. The first start can take a while as weights download. Show me `docker compose logs --tail=50` if it
   has not come up after 20 minutes.
7. Smoke test: `curl -fsS -X POST http://localhost:<PORT>/demo/session -H 'Content-Type: application/json' -d '{"vertical":"clinical"}'`
   should return a token.
8. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.

## Patient data stays here
This box will handle PHI. Keep it that way:
- Do not expose the service to the internet.Reach it only from the clinic network or a private VPN.
- Put TLS in front of it if other machines on the network use it.
- Keep disk encryption on, and keep the machine patched.
- Clinical workloads stay local-only.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/clinical-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

Doesn't fitVisit copilot on GeForce RTX 5090

Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.

Lite · one 48 GB card, live pass only: what changesuses estimates

  • Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.
Memory per component
  • Pass 1: Voxtral Mini 4B Realtime. ~24 GB (at least ~16 GB), weights 8.3 GB (from stack.json). Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more.
  • Language model for the lite tier: Qwen3.8-27B (official FP8). ~33.6 GB (at least ~32 GB), weights 29 GB (from stack.json). Qwen3.8-27B FP8: 33.6 GB is the sales lite tier's allotment (stack.json). Weights of about 29 GB are an estimate (27.8B parameters at one byte, plus higher-precision layers).

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Visit copilot, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Visit copilot on my hardware

Fetch https://decosa.ai/prompts/clinical-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=clinical)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Lite · one 48 GB card, live pass only (lite). Fit check: doesn't fit; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Pass 1: Voxtral Mini 4B Realtime (mistralai/Voxtral-Mini-4B-Realtime-2602), 24 GB
- Language model for the lite tier: Qwen3.8-27B (official FP8) (Qwen/Qwen3.8-27B-FP8), 33.6 GB

Warning: the fit check says this tier does not fit: Needs about 48 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further.

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/clinical-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 48 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh --profile live

Mac prompt for your coding agent

# Decosa Visit copilot: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Visit copilot on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 48 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/clinical.zip (660 KB, 21 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py clinical` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the captions carry the complaint and the MRI order"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Pass 1: live streaming transcript for the in-visit view (no speakers) | vLLM realtime WebSocket | MLX 4-bit (mlx-community/Voxtral-Mini-4B-Realtime-2602-4bit) on mlx-audio 0.5.6, behind scripts/mac/asr_server.py | Runs, measured |
| Speaker labels during the visit: rolling windows (every 15 s of new audio, 6 s overlap) re-transcribed with speaker labels; the committed turns become the visit transcript the note cites | transformers on CUDA | MLX 8-bit (vanch007/mlx-MOSS-Transcribe-Diarize-8bit) on mlx-audio, scripts/mac/diarize_server.py | Runs, measured |
| Language model: live SOAP draft, guidance report, practitioner lanes, window role map, cited note, the self-check's sentence judge, assessment codes and the paperwork field mapping | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |
| Note self-check, detail step (M17): each drug, dose, frequency, date, side and number in a note sentence read against its transcript lines; a 'detail not in the visit' flag becomes a changed-detail error | PyTorch on CPU (8 threads) | The same service with PyTorch on CPU (about 2.5 GB of RAM); expected to run, not timed on a Mac | Runs, not measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 48 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh --profile live`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model, plus about 5 GB for speech recognition and diarization), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`, `"asr": true` and `"diarize": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py clinical`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --profile live --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Measured end to end on 25 Sep (live captions, lanes, the after-visit pass). The visit copilot's rolling speaker labels, self-check and paperwork (28 Sep) use the same models plus the M17 detail checker on CPU; not yet re-measured on a Mac.
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 29 Sep 2026 · measured 29 Sep 2026: · p50 60 s · p95 72 s · ~$0.070 per run · 80 receipts

Loading the nightly status…

Self-host: verified 29 Sep 2026 · fresh clone of the branch into a clean directory, api image built from docker/api/Dockerfile, compose with a named volume and DECOSA_CLINIC_PROFILE, pointed at the model servers already running on our server (Qwen3.8-27B, Voxtral, MOSS diarizer, M17) instead of starting new ones; then torn down

Measured cost to run: about $0.080 per visit (hosted, 29 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The image builds and starts; /healthz ok with asr, llm and diarize true. The copilot sample at 2x passed end to end: 28 speaker turns, guidance, a self-checked note (28 sentences, M17 on), codes from the assessment, three paperwork drafts with the clinic profile from the environment, 0 signatures; done 44.5 s after the audio on a GPU shared with our evals. Model-server startup itself was not re-verified.

Known limits (6)
  • Hosted is for synthetic visits only; real visits must be self-hosted (no BAA yet).
  • Speaker labels were tested on two-speaker visits; a third speaker is labelled Other.
  • The self-check compares the note with the visit's own transcript, so a word the recogniser misheard passes it. The note marks sentences that may rest on one (where the two recognisers disagree, or a word is unknown): 8 of 17 such sentences on held-out visits, with 8% of good sentences also marked. Check names, numbers and yes/no answers against the cited turn.
  • WH-380-E drafts still need every box checked (91% of filled boxes right on held-out visits; vague essential-function wording and incomplete date lists are the usual misses); third-party insurer FMLA forms are not bundled yet.
  • No EHR write-back yet: copy the note and the drafts, or use the API.
  • Guidance is a 26-topic public-domain pack plus the CMS history elements; a missing item does not mean nothing is missing.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB), or 2× RTX 5090.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the visit copilot API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.
The open stack

The visit copilot: transcribes with speaker labels, reminds you what is still worth asking (each item with its source), writes a note checked sentence by sentence before you see it, codes it, and drafts the paperwork, on your own GPU.

One screen for the whole visit. Live captions come first, then speaker-labelled turns a few seconds behind (the MOSS diarizer on rolling windows). While the patient is in the room, a guidance lane lists history and checklist items still worth covering and considerations when a warning feature comes up, each with the words that triggered it and a cited public source; nothing there is an instruction or an alarm, and none of it enters the note. After the visit the note is written from the speaker turns with a citation on every sentence and checked (Check an AI note plus our M17 detail checker) before it is shown as final; codes follow the clinician's stated assessment; and the patient instructions, a work or school note and the FMLA WH-380-E provider sections are drafted from the visit, never signed or sent. It runs on the clinic's own GPU.

Deployment
Self-host firstHosted demo on synthetic visits; real patients: self-host (confidential on request, after a BAA).
Regulatory
HIPAA: run it on the clinic's own hardware so audio, transcripts and notes stay on-site; there is no hosted BAA yet, so the hosted demo takes synthetic visits only. Drafts for clinician review; not a diagnostic device. The guidance lane is designed to stay outside FDA's device definition under FD&C Act 520(o)(1)(E) (FDA guidance on software functions, January 2026): no image or signal analysis (the audio is only transcribed), clinician only, considerations rather than directives, the basis (transcript words and a public source) shown on every item, no alerting or time pressure, nothing written into the note. That is our reading, not a legal opinion: counsel review of the in-visit use is pending before the guidance lane is marketed. Paperwork drafts never fill signatures, signing dates or attestations and are never sent; the clinician reviews, signs and sends. California AB 3030 exempts only clinician-reviewed AI messages to patients. CPT codes are AMA-licensed and hidden in the demo.
Architecture
Text description

Architecture of the self-hosted two-pass clinical scribe. A dashed boundary labelled on-site, PHI never leaves, contains everything clinical. Pass 1, live: the microphone streams 16 kHz audio over a WebSocket to decosa-api on port 8445, which sends it to Voxtral Mini 4B Realtime (Apache-2.0; alternative Nemotron-3.5-ASR-Streaming 0.6B, NVIDIA Open Model License) and sends prompts to Qwen3.8-27B NVFP4 (Apache-2.0). The practitioner lanes note, codes, priorauth and visit_level update live. Pass 2, after the visit: the saved recording goes to MOSS-Transcribe-Diarize 0.9B (Apache-2.0; alternative Sortformer plus Parakeet-TDT 0.6B v3), giving a transcript with S01 and S02 labels; an LLM role map turns them into Doctor and Patient; the note writer (Qwen3.8-27B, or DeepSeek V4 Flash, MIT) writes a SOAP note with [[n]] line citations; a claim verifier checks each sentence before the clinician reviews and signs, and the note goes to the EHR. Badges mark that the hosted demo runs both passes (DeepSeek writer excluded). Outside the boundary, the codes lane queries NLM Clinical Tables, NLM RxNav and openFDA with only a condition name or drug name. Clinic visits use the direct LLM route, so their receipts are signed by the box's own key (attested). Along the bottom, the receipt path of the hosted demo (synthetic visits only): signed receipt, gateway countersignature, hashes only, re-checkable at GET /receipts/{id}.

Architecture

At a glance

Measured latency
The note, its self-check, the assessment codes and the paperwork drafts arrive about a minute after the audio ends on the shared gateway (measured), sooner on the direct route. Speaker labels follow the captions after a short delay; guidance updates every few seconds of speech.
Typical run cost
Several cents per visit for the synthetic sample visit: model calls at the gateway list price plus a little speech-recognition GPU time (measured). The "in model calls" figure after a run is the first part only. Cost grows with visit length (the guidance and the live view reread the transcript every tick).
Data retention
Audio and transcript live in server memory for the session and are dropped when it ends; paperwork PDFs are rendered on request and not stored. The receipt store keeps hashes of each model call, not the text.
What leaves the box (self-host)
Only the codes lane's lookups: a condition name to NLM Clinical Tables and a drug name to NLM RxNav and openFDA. Block egress for the api container and the lanes report no match; the note is still written.
Input
16 kHz mono PCM16 over WebSocket (about 100 ms frames), or a canned sample through /demo/replay. US English visits.
Clinical considerations
After the visit: up to 6 items (warning features first, then differentials) from a 26-topic public-domain pack, each with the patient's words, the source excerpt and URL, and what the note does not record. Clinician only; nothing enters the note; each accept or dismiss is sealed in a signed record. About 10 model calls and USD 0.006 for the chest-pain sample transcript (text route, 2026-09-26). Not a diagnosis and not for urgent decisions.
Guidance during the visit
Warning-feature considerations, checklist items of a triggered topic not yet asked, and history elements not yet covered (CMS 1995 Documentation Guidelines). Every item shows the transcript words and a public source, in fixed wording, and turns 'covered' when asked. No sound, no alarm colours, no deadlines. A blind clinician judge found 66% of checklist items useful at that moment but only 11% of history items, so history folds away until you ask ('Only when I ask' mode).
Paperwork
Patient instructions (the clinician's own words), a work or school note (dates said in the visit are worked out in code: 'through Friday, October 2nd' becomes 10/02/2026, citing the turn), and the Department of Labor's WH-380-E provider sections. Signatures, signing dates and the employer's section are never filled; nothing is sent. WH-380-E drafts need every box checked: 91.4% of the boxes it filled were right on 8 held-out visits.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 48 GB card, live pass only

    Live transcript, lanes and a final note written from the streaming transcript. No speaker labels, no citations, no verifier; the pipeline the benchmark calls vanilla.

    Models
    • Voxtral Mini 4B Realtime
    • Qwen3.8-27B (official FP8)
    Hardware
    1x L40S or RTX 6000 Ada 48 GB (FP8 weights; starting-point settings, not measured)
    Quality evidence
    • ACI-Bench ROUGE-L (Qwen3.8-27B FP8, human transcript)34.2scribe-bench wiki models.md / RESULTS.md
    • PriMock57 note composite, vanilla pipeline (streaming ASR -> Qwen3.8-27B, official weights, test 37)41.01scribe-bench wiki vanilla-vs-best
    • Medical-term miss / WER, live ASR (Voxtral Mini 4B Realtime, PriMock57, 57 visits)8.4% / 13.2 (Nemotron-3.5 streaming: 12.7% / 11.8)scribe-bench asr_score on PriMock57 (57 visits) through the live realtime endpoint; eval results file asr-voxtral-primock57.json (2026-09-23)
    Latency
    not measured yet on a 48 GB card; decode 49 tok/s single-stream for FP8 on an RTX PRO 6000 (measured, benchmark page 17)
    Verification
    Proof: strongSelf-host only
  • In the hosted demo

    Standard

    one 96 GB Blackwell card, two passes

    The hosted demo: live captions, speaker labels, guidance and lanes during the visit; then the cited note, its self-check (with M17 on CPU), assessment codes and paperwork, all on one Qwen3.8-27B.

    Models
    • Voxtral Mini 4B Realtime
    • MOSS-Transcribe-Diarize 0.9B
    • Qwen3.8-27B (NVIDIA NVFP4)
    • decosa-note-detail-checker-modernbert-large (M17, our own model)
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • Medical-term miss / WER, pass 2 MOSS-TD vs the Voxtral live pass (PriMock57, 57 visits)8.4% / 10.3 vs 8.4% / 13.2scribe-bench RESULTS.md (MOSS-TD); eval results file asr-voxtral-primock57.json (Voxtral, 2026-09-23)
    • DER / word speaker misattribution, MOSS-TD11.4 / 1.0%scribe-bench RESULTS.md, wiki decoder-finding
    • Composite, two-pass + role map + Qwen3.8-27B minus vanilla (official weights, test 37)+2.0 [-1.0, +5.7], not resolved; misattributions -0.05scribe-bench wiki vanilla-vs-best (measured with Sortformer + Parakeet as pass 2)
    • Verifier recall on injected errors / flags on clean, Qwen3.8-27B judge99.1% / 6.5%scribe-bench wiki verifier
    Latency
    measured on our server: live lanes arrive seconds apart; after audio ends, pass 2 (diarize, role map, cited note, verifier; the verifier is slowest) finishes and the done event arrives about a minute later (decosa-api smoke runs on our server; contract Changes: after-visit pass 2)
    Verification
    Proof: strong
  • Best

    adds DeepSeek V4 Flash as the note writer on 2x 96 GB

    Standard stack plus the loop's best writer: higher term precision and, with citations, the most grounded notes; costs two extra 96 GB cards and some plan recall.

    Models
    • Voxtral Mini 4B Realtime
    • MOSS-Transcribe-Diarize 0.9B
    • Qwen3.8-27B (NVIDIA NVFP4)
    • DeepSeek V4 Flash
    • decosa-note-detail-checker-modernbert-large (M17, our own model)
    Hardware
    2x RTX PRO 6000 96 GB for the writer (about 83 GB weights per GPU) plus the standard card for the live pass
    Quality evidence
    • ACI-Bench base ROUGE-L / term precision / plan recall35.8 / 70.3 / 93 (Qwen3.8-27B ROUGE-L 34.2)scribe-bench wiki models.md
    • PriMock57 cited note, official weights (test 37): grounded / term precision91.33% / 31.06 (vanilla 88.62% / 26.41)scribe-bench wiki vanilla-vs-best, citations-and-verifiability
    • Composite vs vanilla, official weights+1.25 [-3.76, +6.23], not resolved; plan recall -8.1 [-13.8, -1.7]scribe-bench wiki vanilla-vs-best
    Latency
    measured on our server 2026-09-12: about 220 tok/s decode with the DSpark draft, TP2 (wiki live-synthesis); full-stack latency not measured yet
    Verification
    Proof: strongSelf-host only
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Live, while it happens
Pass 1: live streaming transcript for the in-visit view (no speakers)Voxtral Mini 4B Realtimemistralai/Voxtral-Mini-4B-Realtime-2602 on Hugging Face (opens in a new tab)
4.4B · 24 GBProof: partialIn the hosted demo
Pass 1 alternative: streaming ASR (cache-aware, 1.1 s chunks)Nemotron-3.5-ASR-Streaming 0.6Bnvidia/nemotron-speech-streaming-en-0.6b on Hugging Face (opens in a new tab)Alternative to Voxtral Mini 4B Realtime
0.6BProof: partialSelf-host only
Speaker labels during the visit: rolling windows (every 15 s of new audio, 6 s overlap) re-transcribed with speaker labels; the committed turns become the visit transcript the note citesMOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab)
0.9BProof: partialIn the hosted demo
After the session
Pass 2 alternative, part 1: offline transcriptParakeet-TDT 0.6B v3nvidia/parakeet-tdt-0.6b-v3 on Hugging Face (opens in a new tab)Alternative to MOSS-Transcribe-Diarize 0.9B
0.6BProof: partialSelf-host only
Pass 2 alternative, part 2: speaker diarization (up to 4 speakers)Sortformer diarizer 4spk v1nvidia/diar_sortformer_4spk-v1 on Hugging Face (opens in a new tab)Alternative to MOSS-Transcribe-Diarize 0.9B
0.12BProof: partialSelf-host only
Language model for the lite tier: live lanes and the final note from the live transcriptQwen3.8-27B (official FP8)Qwen/Qwen3.8-27B-FP8 on Hugging Face (opens in a new tab)
27.8BProof: strongSelf-host only
Language model: live SOAP draft, guidance report, practitioner lanes, window role map, cited note, the self-check's sentence judge, assessment codes and the paperwork field mappingQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
Note writer alternative: the loop's best writerDeepSeek V4 Flashdeepseek-ai/DeepSeek-V4-Flash on Hugging Face (opens in a new tab)Alternative to Qwen3.8-27B (NVIDIA NVFP4)
284B (13B active) · 166 GBProof: strongSelf-host only
Note self-check, detail step (M17): each drug, dose, frequency, date, side and number in a note sentence read against its transcript lines; a 'detail not in the visit' flag becomes a changed-detail errordecosa-note-detail-checker-modernbert-large (M17, our own model)decosaai/decosa-note-detail-checker-modernbert-large on Hugging Face (opens in a new tab)
395M · 0 GBProof: partialIn the hosted demo

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    Lane engine and HTTP/WS API (/ws/live, /demo/replay, /healthz). No GPU. Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B, served as qwen3.8-27b. Internal to the compose network.

  • decosa-asr:8000
    ${DECOSA_REGISTRY}/decosa-asr:0.1.0

    vLLM realtime endpoint for Voxtral Mini 4B Realtime (served as voxtral-realtime). Internal to the compose network.

  • decosa-diarize:8092

    MOSS-Transcribe-Diarize 0.9B pass-2 service (decosa-api services/diarize, GPU0, loopback only). decosa-api runs pass 2 on stop: diarize, role map, cited note, verifier. No published image yet; self-host builds it (see the assemble prompt).

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB (or B200/GB200) Fits

    Default compose split: LLM 0.60 of the card (~57 GB), ASR 0.25 (~24 GB), about 4 simultaneous visits. Measured on our server with the two models on separate cards (LLM 0.92 of GPU1, ASR 0.40 of GPU0 using 34 GB); the one-card split itself is not yet measured. The after-visit pass (0.9B, BF16 weights 1.8 GB) is meant to share the card after the visit; that fit is not measured.

  • 2x RTX PRO 6000 96 GB, DeepSeek V4 Flash as writer

    DeepSeek V4 Flash ran on our server across both cards (TP2, weights about 83 GB per GPU, 32k context) with little room left, so the live models would need another card. Not measured as a full stack.

  • 1x H100 80 GB / H200 141 GB (Hopper, no NVFP4)

    Not measured. Use Qwen/Qwen3.8-27B-FP8 (Apache-2.0) with LLM_GPU_UTIL=0.62, ASR_GPU_UTIL=0.25 as a starting point.

  • 1x L40S / RTX 6000 Ada 48 GB

    Not measured. FP8 checkpoint, LLM_MAX_LEN=16384, LLM_GPU_UTIL=0.70, ASR_GPU_UTIL=0.22, DECOSA_LIVE_CAP=2 as a starting point.

  • Cards under 48 GB Does not fit

    Both models do not fit with useful KV cache (decosa-api self-host guide).

Latency per lane

  • transcript1.4 s

    Measuredmeasured on our server 2026-09-23: first transcript event 1.4 s after audio start on both clinical replays (ops/record-all.json)

  • note7.4 s

    Measuredmeasured on our server 2026-09-23: per-replay median 7354 ms (clinical-back-pain, n=10) and 4724 ms (clinical-diabetes-followup, n=9); gateway route, LLM alone on its own GPU, one session

  • codes1.4 s

    Measuredmeasured on our server 2026-09-23: median 1449 ms (clinical-diabetes-followup, n=8); back-pain median was 4 ms because repeat lookups hit the in-memory cache, max 5302 ms

  • priorauth5.6 s

    Measuredmeasured on our server 2026-09-23: per-replay median 5606 ms (clinical-back-pain, n=4) and 2036 ms (clinical-diabetes-followup, n=1)

  • visit_level6.8 s

    Measuredmeasured on our server 2026-09-23: per-replay median 6819 ms (clinical-back-pain, n=2) and 2684 ms (clinical-diabetes-followup, n=2)

  • final note after stop, single pass (before pass 2 was added)24.1 s

    Measuredmeasured on our server 2026-09-23: done event 24.1 s (clinical-back-pain) and 12.1 s (clinical-diabetes-followup) after the audio ended

  • note, 4 live sessions in parallel13.7 s

    Measuredmeasured on our server 2026-09-23: median 13698 ms, max 26594 ms (clinical-back-pain beside three other verticals, ops/smoke-parallel-1.json)

  • pass 2 transcript (9-minute visit)39.0 s

    Measuredmeasured on our server 2026-09-12: MOSS-Transcribe-Diarize, GPU1 (wiki live-synthesis)

  • final_transcript (diarize, after audio ends)9.0 s

    Measuredmeasured, median of 3 clinical runs, range 6.3-11.3 s over 3 clinical runs; decosa-api smoke runs 2026-09-23/24 on our server (ops/record-pass2.json, ops/smoke-final.json; contract Changes: after-visit pass 2)

  • role map2.1 s

    Measuredmeasured, median of 3 clinical runs, range 1.3-3.3 s; decosa-api smoke runs 2026-09-23/24 on our server (ops/record-pass2.json, ops/smoke-final.json; contract Changes: after-visit pass 2)

  • final_note (cited)4.1 s

    Measuredmeasured, median of 3 clinical runs, range 3.3-4.4 s; decosa-api smoke runs 2026-09-23/24 on our server (ops/record-pass2.json, ops/smoke-final.json; contract Changes: after-visit pass 2)

  • verifier30.0 s

    Measuredmeasured, median of 3 clinical runs, range 18.6-51.3 s, one judge call per claim (14-21 claims); decosa-api smoke runs 2026-09-23/24 on our server (ops/record-pass2.json, ops/smoke-final.json; contract Changes: after-visit pass 2)

  • considerations (text route, one encounter)18.1 s

    Measuredmeasured on our server 2026-09-26: median over 26 held-out encounters, 3 in parallel on the shared gateway, max 25.4 s (decosa-api docs/evals/clinical-considerations.md); in the live scribe it runs beside pass 2 after the visit

  • all lanes, one-card self-host (direct route)n/a

    Estimateestimate: not measured; the numbers above used the gateway route with the LLM on its own card

  • done after audio ends (note, self-check, codes, paperwork)60.0 s

    Measuredmeasured on our server 2026-09-29: 3 gateway runs of the 2:52 synthetic visit under shared load, 33.7-71.6 s; the direct route took 20 s (28 Sep)

  • turns (speaker labels behind the captions)15.0 s

    Measuredmeasured 2026-09-28/29: a window every 15 s of new audio, 2.3 s median diarizer time per window on a shared GPU

  • guidance4.3 s

    Measuredmeasured 2026-09-28: the guidance call beside the live view on each tick (every 8-15 s of speech), median model latency on the gateway 4.3 s

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

clinical/assemble-prompt.md237 lines
# Assemble the Decosa visit copilot on this machine

You are setting up the self-hosted visit copilot on a Linux box with NVIDIA GPUs that the clinic controls.
- **During the visit:**
  - Voxtral Mini 4B Realtime gives live captions.
  - The MOSS-Transcribe-Diarize 0.9B service labels who said what, over rolling windows (the `turns` lane).
  - Qwen3.8-27B keeps a SOAP draft, ICD-10-CM and RxNorm codes, prior-auth prep and the E/M level current.
  - The `guidance` lane shows questions and history still worth covering, and considerations when a warning feature comes up. Each shows the transcript words and a cited public source; it is for the clinician and never written into the note.
- **After the visit:**
  - The LLM writes the note from the speaker turns, with `[[n]]` citations (reasoning off).
  - Check an AI note checks every sentence before the note is shown as final, optionally with our M17 detail checker on CPU. Failing sentences are held back and listed.
  - The paperwork drafts are filled from the visit and never signed or sent: patient instructions, a work or school note, and the FMLA WH-380-E provider sections.

In PriMock57 tests, pass 2 missed 8.4% of medical terms against 12.7% for Nemotron streaming ASR (scribe-bench RESULTS.md); the Voxtral live pass also missed 8.4%, at a higher WER (13.2 vs 10.3). Work step by step, show each command before running it, and stop to ask me before anything that needs sudo or changes the firewall.

**Hard rules**
- Real patient audio goes ONLY to the API you are installing here. The hosted Decosa API and demo site must never receive PHI.
- Keep `DECOSA_LLM_ROUTE=direct`. Clinic visits then get receipts with `status: "attested"`; signed by the clinic box's own key, since nothing is sent to a gateway.
- Keep the API on `127.0.0.1` unless I ask for LAN access. Never expose it to the internet.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/clinical.zip (660 KB, 21 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py clinical` (the api image carries the same bundle under /app/rehearsal/clinical/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py clinical --bundle clinical.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the captions carry the complaint and the MRI order"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 1. Check the machine
```bash
nvidia-smi --query-gpu=name,memory.total,driver_version,compute_cap --format=csv
df -h /var/lib/docker 2>/dev/null || df -h /     # need ~80 GB free for images and weights
docker --version && docker compose version
docker run --rm --gpus all ubuntu:24.04 nvidia-smi -L   # proves the NVIDIA container toolkit works
```
- Driver 580 or newer is required for the images as built. If older, tell me and stop.
- Pick the model build from the GPU:
  - Blackwell (compute_cap 10.x or 12.x, e.g. RTX PRO 6000 96 GB, B200): defaults below (NVFP4).
  - Hopper 80 GB+ (H100/H200): `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=017b9c7af6b5689d5dd426a76e0bc077eb5ca20a`, `LLM_GPU_UTIL=0.62`, `ASR_GPU_UTIL=0.25`.
  - 48 GB (L40S, RTX 6000 Ada): the FP8 settings plus `LLM_MAX_LEN=16384`, `LLM_GPU_UTIL=0.70`, `ASR_GPU_UTIL=0.22`, `DECOSA_LIVE_CAP=2`.
  - Under 48 GB: both models don't fit with useful KV cache. Stop and tell me.
  - Only the Blackwell setup has been measured; the others are starting points.
- Pick a quality tier with me:
  - **lite**: a 48 GB card, pass 1 only; skip step 4b.
  - **standard**: one 96 GB Blackwell card, both passes.
  - **best**: standard plus DeepSeek V4 Flash (`deepseek-ai/DeepSeek-V4-Flash`, MIT) as the writer, on two more 96 GB cards (see 4b).
- If Docker or the toolkit is missing (Ubuntu), install them, with my OK for sudo:
```bash
curl -fsSL https://get.docker.com | sudo sh
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -fsSL https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list \
  | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#' \
  | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update && sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker && sudo systemctl restart docker
```

## 2. Images and models
Work in `~/decosa-clinical`. The images are `${DECOSA_REGISTRY}/decosa-{llm,asr,api}:0.1.0` (**publishing soon**). Try `docker pull` first. If a pull fails, build them:
- `decosa-llm`: `FROM vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1` (vLLM 0.29.0) with `ENTRYPOINT ["vllm","serve"]`.
- `decosa-asr`: `FROM vllm/vllm-openai:v0.27.1`, then `RUN pip install --no-cache-dir "mistral-common[audio]" soundfile librosa`, with `ENTRYPOINT ["vllm","serve"]`.
- `decosa-api`: needs the decosa-api source (`docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .` in the checkout). If I haven't given you the source, stop and ask me for it.

The weights (both Apache-2.0, about 25 GB) download from Hugging Face on first start into the `hf-cache` volume:
- `nvidia/Qwen3.8-27B-NVFP4` at revision `482ca0f3832238542f8f5295dde86b5f22711d80` (27.8B dense; NVFP4 MLP, FP8 attention);
- `mistralai/Voxtral-Mini-4B-Realtime-2602` (4.4B, BF16).

## 3. Write `.env` and `docker-compose.yml`
`.env` (adjust per step 1):
```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.60
ASR_GPU_UTIL=0.25
DECOSA_LIVE_CAP=4
DECOSA_BIND=127.0.0.1
DECOSA_PORT=8445
# the clinic web app's origin, e.g. https://scribe.clinic.local (keep comments on their own line: compose reads an
# inline comment after an empty value as the value)
DECOSA_CORS_ORIGINS=
# pass 2 (standard and best tiers, step 4b): the diarize service's URL as the api container sees it; empty = pass 1 only
DECOSA_DIARIZE_URL=
HF_TOKEN=
```
`docker-compose.yml`:
```yaml
name: decosa
x-gpu: &gpu
  deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
x-health: &health
  test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"]
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    <<: *gpu
    ipc: host
    restart: unless-stopped
    environment: { HF_TOKEN: "${HF_TOKEN:-}" }
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
      "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
      "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
      "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, start_period: 900s }
  asr:
    image: ${DECOSA_REGISTRY}/decosa-asr:${DECOSA_TAG}
    <<: *gpu
    ipc: host
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }   # start after llm so the memory split is stable
    environment: { HF_TOKEN: "${HF_TOKEN:-}" }
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["--model", "mistralai/Voxtral-Mini-4B-Realtime-2602", "--tokenizer-mode", "mistral", "--config-format", "mistral",
      "--load-format", "mistral", "--compilation-config", '{"cudagraph_mode":"PIECEWISE"}', "--max-model-len", "45000",
      "--max-num-batched-tokens", "8192", "--max-num-seqs", "16", "--gpu-memory-utilization", "${ASR_GPU_UTIL}",
      "--served-model-name", "voxtral-realtime", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, start_period: 600s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy }, asr: { condition: service_healthy } }
    environment:
      DECOSA_ASR_WS: ws://asr:8000/v1/realtime
      DECOSA_LLM_ROUTE: direct            # never "gateway" with PHI
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LIVE_CAP: ${DECOSA_LIVE_CAP}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_AUDIO_S: "3600"
      DECOSA_BUDGET_LLM_TOKENS: "200000"
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_CORS_ORIGINS: ${DECOSA_CORS_ORIGINS:-}
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
      DECOSA_STUDIO_WORKER: none
      DECOSA_DIARIZE_URL: ${DECOSA_DIARIZE_URL:-}   # empty: pass 1 only (lite); see 4b
    ports: ["${DECOSA_BIND}:${DECOSA_PORT}:8445"]
    volumes: [decosa-data:/data]
    healthcheck:
      test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"]
      interval: 30s
      timeout: 5s
      retries: 5
      start_period: 20s
volumes: { hf-cache: {}, decosa-data: {} }
```
Start it: `docker compose up -d`. The first start downloads weights; `llm` takes about 5–10 minutes to turn healthy, then `asr` 1–3 minutes more. Watch `docker compose ps` and `docker compose logs -f llm asr`.

## 4. Smoke test, pass 1 (synthetic visit, no microphone)
```bash
curl -s localhost:8445/healthz   # expect "ok": true, "asr": true, "llm": true, "llm_route": "direct"
TOKEN=$(curl -s localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"clinical"}' | jq -r .token)
curl -sN localhost:8445/demo/replay -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
  -d '{"vertical":"clinical","script_id":"clinical-back-pain"}' | tee /tmp/clinical-smoke.sse | grep -o '"type": *"[a-z]*"' | sort | uniq -c
grep -o '"lane": *"[a-z_]*"' /tmp/clinical-smoke.sse | sort -u
```
The replay runs a synthetic 2-minute visit in real time through the real ASR and LLM. Pass criteria:
- the stream starts with `ready` (with `"receipts": "attested"` and the demo banner);
- `transcript` events arrive within a few seconds;
- `lane` events appear for `guidance`, `note`, `codes`, `priorauth` and `visit_level`;
- the `guidance` lane's data has `audience: "clinician"`, `inserted_into_note: false`, `alarm: false` and `lint.ok: true`. Every item has `basis.quote` and `basis.source.url`, and the `note` lane has no suggestions in it;
- `receipt` events carry `"status": "attested"`;
- it ends with `done`, whose `summary` has `note`, `note_markdown`, `codes`, `priorauth` and `visit_level` (and `"pass2": "unavailable"` while `DECOSA_DIARIZE_URL` is empty).

Also try `clinical-diabetes-followup`. On our server, with the LLM on its own GPU, lane updates took about 1.5–7.5 s (median), and the final note arrived 12–24 s after the audio ended. Report what you measure here.

## 4b. Speaker labels, the checked note and the paperwork (standard and best tiers)
decosa-api runs the rest itself, all with receipts:
- during the visit it sends rolling windows of the session audio (every 15 s of new audio) to a diarize service and labels each window's speakers as Clinician/Patient/Other;
- when the visit stops it commits the last window and writes the cited note from those turns;
- it self-checks the note, then drafts the paperwork.

The results arrive as lanes (`turns`, `final_note`, `selfcheck`, `paperwork`, `considerations`) and in the same `done` event. You only have to run the diarize service and set `DECOSA_DIARIZE_URL`.
- The service is `services/diarize` in the decosa-api source (MOSS-Transcribe-Diarize 0.9B, weights revision `704aa4a9c304e8520be88901e0d1960158ef5b15`). It has no published image yet. Install it from the source with the `uv` commands at the top of `services/diarize/requirements.txt` and run `DIARIZE_HOST=<address> DIARIZE_PORT=8092 DIARIZE_DEVICE=cuda:0 .venv-diarize/bin/python services/diarize/server.py` on this box (a systemd unit is fine). Bind it to the docker bridge address (`ip -4 addr show docker0`, usually `172.17.0.1`), which the `api` container can reach and the LAN cannot.
- Add `extra_hosts: ["host.docker.internal:host-gateway"]` to the `api` service, set `DECOSA_DIARIZE_URL=http://host.docker.internal:8092` in `.env`, and run `docker compose up -d api`. `curl -s localhost:8445/healthz` should then show `"diarize": true`.
- For standard, lower `LLM_GPU_UTIL` to 0.55 so the diarizer (about 4 GB) fits beside the live models. That's an estimate: the one-card fit has not been measured.
- Smoke test: rerun the step 4 replay with `"script_id":"clinical-visit-copilot"` (2:52, a back injury at work with a work note and FMLA). `done.summary` must now have:
  - `final_transcript.lines` with `Clinician` and `Patient` roles;
  - `final_note.status: "self-checked"`, with every kept sentence ending in `[[n]]`;
  - `selfcheck.summary`;
  - `paperwork.drafts` for `instructions`, `work_note` and `wh380e`, with `signatures_filled: 0`.

  On our server (2026-09-29, gateway route under shared load) everything arrived 34–72 s after the audio ended (median 60 s over 3 runs); on the direct route it took 20 s.
- **The M17 detail checker (optional, CPU):** it upgrades the self-check's "detail not in the visit" flags to changed-detail errors. Set it up as in the Check an AI note kit (`services/detail_checker`, weights `decosaai/decosa-note-detail-checker-modernbert-large`, `DETAIL_PORT=8495`), then give the api `DECOSA_DETAIL_URL=http://host.docker.internal:8495`. Without it, the self-check runs the sentence judge alone and says `detail_checker: "off"`.
- If the diarizer is down, the visit still finishes. `done` arrives with a note written from the live captions and marked not self-checked, and an `error` event says the speaker labels failed.
- Paperwork uses the clinic's own details. Set `DECOSA_CLINIC_PROFILE` in `.env` to one JSON line with `clinician`, `clinic`, `address`, `specialty`, `phone`, `fax` and `email`; `POST /copilot/paperwork` also takes them per draft. Without it, the drafts carry the built-in demo profile, which is labelled demo.

Alternative: a standalone worker for recordings saved to disk (not needed when the api runs pass 2 as above).
- Image `decosa-pass2` (local): `FROM vllm/vllm-openai:v0.27.1`, then `RUN pip install --no-cache-dir "git+https://github.com/OpenMOSS/MOSS-Transcribe-Diarize" soundfile openai` and `ENTRYPOINT []`. If transformers conflicts, pin what the MOSS repo asks for. Add it to compose as service `pass2`: GPU as `llm`, `hf-cache` mounted, `./visits:/visits`, `depends_on: llm`, `command: ["python","/app/after_visit.py","--watch","/visits"]`.
- Write `after_visit.py`. For each new `/visits/<id>.wav` (16 kHz mono, saved by the clinic app on this box):
  1. **Transcribe.** Load `OpenMOSS-Team/MOSS-Transcribe-Diarize` (revision `704aa4a9c304e8520be88901e0d1960158ef5b15`) with `AutoModelForCausalLM.from_pretrained(..., trust_remote_code=True, dtype=torch.bfloat16, attn_implementation="sdpa")` and `AutoProcessor`. Then call `moss_transcribe_diarize.inference_utils.build_transcription_messages(wav, prompt=DEFAULT_PROMPT)` and `generate_transcription(model, processor, messages, max_new_tokens=8192, do_sample=False, ...)`, and parse the result with `parse_transcript`. You get segments `{start, end, speaker, text}` labelled `S01`, `S02` and so on.
  2. **Role map.** Make one call to `http://llm:8000/v1` (model `qwen3.8-27b`, `max_tokens` 60, `chat_template_kwargs: {"enable_thinking": false}`) with this system prompt: "Decide which label is the CLINICIAN (asks questions, examines, advises, plans) and which is the PATIENT. Answer with one JSON object only, e.g. {"S01": "Doctor", "S02": "Patient"}; map extra labels to "Other"." Relabel the lines with the answer.
  3. **Cited note.** Number the transcript lines. Ask for a SOAP note with reasoning off and this rule: "Every sentence you write must end with a citation of the line number(s) that support it, formatted like [[12]] or [[12,15]]. A sentence with no supporting line must not be written." For the best tier, send this call to the DeepSeek endpoint instead.
  4. **Verify.** For each sentence, pass its cited lines ±1 line to the LLM as a judge. The judge first compares each fact (drug, dose, number, side, duration, negation, who it is about), then ends with the line `VERDICT: <SUPPORTED|PARTIAL|UNSUPPORTED> LEAK=<yes|no>`. Keep the verdict last.
  5. Write `/visits/<id>.note.json` and `.note.md`. Flag every sentence that isn't SUPPORTED, and every LEAK, for the clinician.
- Smoke test: `docker compose cp api:/app/demo_scripts/audio/clinical-back-pain.wav ./visits/test.wav` (a synthetic visit), wait for `test.note.md`, and check that most sentences carry `[[n]]` and every flagged one shows its verdict. For reference, a 9-minute visit took 39 s for pass 2 on our server. The role map, note and verifier have not been timed.
- Best tier only: serve DeepSeek V4 Flash with tensor parallel 2 on two other 96 GB cards. It takes about 83 GB of weights per GPU. Put it on its own OpenAI-compatible port on this box and point step 3 at it. It is not a decosa image, so ask me for the serving kit.

## 4c. Clinical considerations (all tiers, no extra service)
After each visit the api also emits a `considerations` lane (and `done.summary.considerations`): warning features stated in the visit and differential considerations from a bundled public-domain pack, each with transcript quotes, a cited source and the note's documentation gaps. For the clinician only; nothing is written into the note. Check it on its own with the synthetic chest-pain transcript in the image:
```bash
docker compose cp api:/app/rehearsal/clinical/inputs/chest-pain-transcript.txt /tmp/cp.txt
T2=$(curl -s localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"clinical"}' | jq -r .token)
jq -Rs '{transcript: .}' /tmp/cp.txt | curl -s localhost:8445/clinical/considerations -H "authorization: Bearer $T2" \
  -H 'content-type: application/json' -d @- | tee /tmp/cons.json | jq '[.items[] | {kind, topic, quotes: [.quotes[].text]}], .directive_lint.ok'
```
Pass: an item with `"topic": "rf_acs"` quoting the jaw, `true` for the lint, and every receipt `attested`. Then `jq '{statement, items, decisions: [{id: "c1", decision: "accepted"}]}' /tmp/cons.json | curl -s localhost:8445/clinical/considerations/decisions -H "authorization: Bearer $T2" -H 'content-type: application/json' -d @- | jq .record_check` must print `"ok": true`. On our server (26 Sep 2026, gateway route) this took 20-35 s and about 10 model calls.

## 5. Point the app at the local API
- Base URL: `http://localhost:8445` (for the Decosa web app: `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`).
- Get a token with `POST /demo/session {"vertical":"clinical"}`, open `ws://localhost:8445/ws/live?vertical=clinical&token=<t>`, send 16 kHz mono PCM16 little-endian binary frames of about 100 ms, then `{"type":"stop"}`. Treat the latest `lane` event per lane id as the full state (replace, don't append). Show the `support` lane in its own panel with its `label` visible at all times, to the clinician only, and never copy its items into the note or the EHR automatically. Its red flags are considerations worded as "Possible red flag: consider whether urgent evaluation is needed"; the app must not turn them into alarms, pop-ups or sounds, and they are not for emergencies. The final note and codes arrive in `done`; your EHR integration decides where they go.
- The app's origin must be in `DECOSA_CORS_ORIGINS`, or the WebSocket closes with `4403`.
- LAN access: put a TLS reverse proxy (Caddy or nginx) in front, set `DECOSA_TRUSTED_PROXIES` to the proxy's address, and restrict it to the clinic network.
- Outbound calls: the codes lane queries NLM Clinical Tables, NLM RxNav and openFDA with only a condition name or a drug name. If policy forbids any egress, block it for the `api` container. The lanes then report no match, and the note is still written.

When done, print a short report: the tier, GPU and build chosen, image sources (pulled or built), health output, pass-1 and pass-2 smoke-test results and latencies, and anything that failed.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

13 laws, rules and guidance pages cited; 10 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
The visit copilot: transcribes with speaker labels, reminds you what is still worth asking (each item with its source), writes a note checked sentence by sentence before you see it, codes it, and drafts the paperwork, on your own GPU.
Who it's for
Family physicians, NPs and PAs, and clinics that want the note, codes and forms from the visit on their own hardware.
Where it runs
Self-host (patient data stays on site)
Key numbers
  • 8.4% Medical-term miss rate, live ASR (Voxtral Mini 4B Realtime) (held out, n = 57)
  • 99.71% Speaker labels during the visit, word-level role accuracy (rolling MOSS windows) (held out, n = 11)
  • 14/15 and 13/15 Considerations: warning-feature encounter recall (synthetic, held-out) (held out, n = 15)
  • 60.0 s Median end-to-end run, hosted (QA sweep 2026-09-29)
All results, datasets and caveats
Models
Voxtral 4B · MOSS diarizer · Qwen3.8-27B · M17 detail checker
Where
Self-host (patient data stays on site)
Checks
Note checked before it is final
Industry
Healthcare
Output
Notes, reports and drafts · Structured data
Data
Patient data (PHI)
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

What does the Decosa visit copilot do that a free AI scribe does not?

Besides the note, the visit copilot shows speaker-labelled turns during the visit, lists history and checklist items still worth covering with their public sources, checks every note sentence against the transcript before showing it as final, codes the clinician's stated assessment, and drafts patient instructions, a work or school note and the FMLA WH-380-E provider sections, never signed or sent.

How accurate are the visit copilot's speaker labels?

On 11 PriMock57 mock consultations (110 minutes), the rolling speaker labels the visit copilot shows during the visit put 99.71% of words with the right speaker, against 99.97% for one pass over the whole recording afterwards. Only two-speaker visits were tested.

Does the visit copilot's guidance tell the clinician what the patient has?

No. The visit copilot's guidance lists considerations for the clinician, each with the transcript words that raised it and a cited public source, in fixed wording with no alarms, and nothing enters the note. On 38 synthetic visits it showed 1 false-alarm visit in 16. Counsel review of the in-visit use is pending.

How grounded is the visit copilot's note?

A blind reviewer checked every sentence of notes from 11 held-out PriMock57 consultations against the human transcript: 289 of 308 were supported (93.8%). Of the 19 misses, 3 came from the note writer and 16 from words the speech recogniser misheard, which a check against the visit's own transcript cannot catch. The note marks sentences that may rest on a misheard word (about half of them in a second held-out test), so check names, numbers and yes/no answers against the cited turn.

Does the visit copilot sign or send forms?

Never. The visit copilot leaves signatures, signing dates and the employer's section blank and sends nothing; 0 signature boxes were filled in any test. On held-out FMLA visits, 91% of the WH-380-E boxes it filled were right, and each box shows the turn it came from, so check every box before you sign.

Can the visit copilot run without sending patient data out?

Yes. The visit copilot is built to self-host: speech, speaker labels, the language model and the detail checker run on the clinic's own GPU. The only outbound calls are code-table lookups of a condition name or drug name; block them and the note is still written. The hosted demo takes synthetic visits only.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Visit copilot

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.