Skip to content
decosa
LiveHostedSelf-hostMacSelf-host first for real data

Check an AI note

Every note sentence checked against the transcript, with the line it came from or 'not in the transcript': invented facts, changed doses and details, exam findings nobody made, the wrong speaker, and key items left out. Plus a shorter note that only removes words, and a signed scorecard per vendor.

Held-out test47 (94%)Changed doses and details caught as errors (held-out test)
On production28 smedian on production (2026-09-28); slower when the service is busy
List price~$0.024 per visit notemeasured, at list price

Built on: Grounding, Speaker diarization, Signed record

1. Pick a sample

Sample

Real patient data: request confidential access, or run the tool on your own hardware. The demo takes samples or made-up data only.

2. Run it

On production the sample took 28 s (median, 2026-09-28; slowest 41 s). Slower when the service is busy.

Result

The answer appears here first, then what it found, the draft, and how long it took. Sample: Video visit, AI note with four planted errors (synthetic).

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the judge; the diarizer (optional, for audio) needs a second GPU slot.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the clinical ai assurance monitor API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
clinical-ai-monitor

Use the hosted API

# Decosa Clinical AI assurance monitor: try the hosted API

You are wiring Decosa's clinical AI assurance monitor into a test harness. It takes a visit transcript and the note an
ambient AI scribe wrote, checks every note sentence against the transcript, types each problem (not in the visit,
contradicts the visit, wrong speaker, invented exam finding, a changed detail), lists the medications, allergy status,
plan items and impression the note left out, and signs a per-visit report. Every verdict is one model call with its
own signed receipt. Use only what is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API is for the public mock consultations and synthetic notes only. Never send real patient data
  (PHI) to it.** Real visits run on our own hardware: use the self-host prompt. Refuse to send anything that looks like a
  real patient's visit.
- It measures documentation quality against the transcript, not care. Say so wherever you show results.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page, in `DECOSA_API_KEY`, never in code. Send
   `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "clinical-ai-monitor"}` returns `{"token", "expires_at", "budget"}`.
   a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); each session has a token allowance (its `budget`) (about four notes). Over a limit: HTTP 429 with `Retry-After`.
3. A check needs about 170 generated tokens per note sentence plus 1,500 for the checklist (402 otherwise). One check
   at a time per demo token (409).

## Endpoints
- `GET /monitor/samples` (no token): three PriMock57 mock consultations (CC BY 4.0), each with `segments` and `notes`:
  a clean synthetic scribe note, five variants with one planted error each (`fabricated`, `altered`, `omission`,
  `speaker`, `exam`, with `planted` saying what), and the clinician's own note (`human`).
- `POST /monitor/check` (token). Body:
  `{"transcript": "Clinician: ...\nPatient: ..." | "segments": [{"speaker": "Clinician", "text": "..."}], "note": "...", "vendor"?: "Scribe A", "vendor_version"?: "5.1", "specialty"?: "respiratory", "period"?: "2026-W39", "encounter_ref"?: "...", "checklist"?: [{"category": "medication|allergy|plan|impression", "item": "..."}]}`
  - Transcript: one turn per line, starting with who spoke (`Doctor:`, `Clinician:`, `Patient:`, `[Mother]`); anonymous
    labels (`S01`) cost one extra model call to name the clinician. Up to 40,000 characters and 400 turns. Note up to
    12,000 characters and 80 checked sentences.
  - JSON by default: `{decision: clean|review|errors, counts, sentences: [{i, text, section, verdict, issue, finding, lines, quotes, reason, receipt_ids}], checklist: [{category, item, lines, status, finding}], transcript, receipts, report, budget}`.
  - With `Accept: text/event-stream` (or `"stream": true`): `ready`, then per sentence a `receipt` and a `sentence`
    event as each finishes, `checklist`, `coverage`, `report`, `budget`, `done`.
- `POST /monitor/transcribe` (token): body = WAV audio (≤ 15 minutes) → `{lines, segments, ...}` via the MOSS diarizer;
  pass `segments` to `/monitor/check` with `"source": "audio"`. Counts against the 300-second audio budget.
- `POST /monitor/verify` (no token) `{"report": {...}, "transcript"?: "...", "note"?: "..."}` → `{valid_signature, signed_by_this_server, transcript_matches?, note_matches?}`.
- `POST /monitor/summary` (no token) `{"reports": [...], "group_by"?: ["vendor", "specialty"]}` → rates with 95% intervals,
  a week-by-week trend and drift test per vendor, and a signed summary. `GET /monitor/demo-dashboard` shows one.
- `GET /monitor/info`, `GET /attest/signing-key` (no token).

Offsets are Python string indices (Unicode code points) into the note as sent.

## Example: measure a sample note and roll it into a summary (Python, `pip install httpx`)
```python
import httpx, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
enc = httpx.get(f"{API}/monitor/samples", timeout=30).json()["encounters"][1]
reports = []
for variant in ("clean", "speaker"):
    note = next(n for n in enc["notes"] if n["id"] == variant)
    r = httpx.post(f"{API}/monitor/check", headers=H, timeout=300,
                   json={"segments": enc["segments"], "note": note["note"], "vendor": "Scribe A", "specialty": enc["specialty"]})
    r.raise_for_status()
    run = r.json()
    print(variant, run["decision"], {k: v for k, v in run["counts"].items() if v and k not in ("sentences", "checked", "supported")})
    for s in run["sentences"]:
        if s["finding"]:
            print("  ", s["finding"], "|", s["text"], "| lines", s["lines"])
    for it in run["checklist"]:
        if it["finding"]:
            print("  ", it["finding"], "|", it["category"], it["item"])
    reports.append(run["report"])
summary = httpx.post(f"{API}/monitor/summary", json={"reports": reports}, timeout=60).json()
print(summary["total"]["sentence_error_rate"], summary["summary"]["id"])
```

## Honest limits
- The judge is one open model; its detection and false-flag rates are on the Stack tab. Planted errors in the eval are
  clear single errors; subtler errors will be missed more often.
- It checks the note against the transcript given; transcript errors pass into the measure.
- Clinicians' own notes get flagged for things never said aloud (brand names, routine negatives). Use it on AI drafts.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Clinical AI assurance monitor: run it yourself (containers)

You are setting up the Decosa clinical AI assurance monitor inside our network, so transcripts and notes (PHI) never
leave it. It checks each scribe note against its visit transcript, lists what the note left out, and signs per-visit
reports that roll up into rates per vendor, specialty and week. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/clinical-ai-monitor.zip (8 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py clinical-ai-monitor` (the api image carries the same bundle under /app/rehearsal/clinical-ai-monitor/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py clinical-ai-monitor --bundle clinical-ai-monitor.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "all 4 clinical sentences were checked", "every checked sentence got a verdict (none not judged or errored)", "the planted exam finding (oxygen saturation 97%) is not supported"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For `api` set `DECOSA_LLM_ROUTE=direct`,
   `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_LOCAL_SIGNING=on`, and bind every port to
   127.0.0.1. For audio input, run the diarizer (`services/diarize` in the decosa-api repo) and set
   `DECOSA_DIARIZE_URL`; otherwise leave it empty and send transcripts.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/monitor/info` lists the finding types and `judge.route: "direct"`;
   `GET /attest/signing-key` shows this box's public key. Show me the key: the committee pins it to check our reports.
5. Smoke test: get a token with `POST /demo/session {"vertical":"clinical-ai-monitor"}`, take encounter 1 from
   `GET /monitor/samples`, and run `POST /monitor/check` with its `segments` and the `speaker` note. Expect
   `decision: "errors"` with the planted sentence flagged. Then `POST /monitor/verify` with the report and the note:
   `valid_signature` and `note_matches` must be true.
6. Report back: the public key, the smoke-test decision and counts, and how long the check took.

Off, and keep it off: this box holds PHI.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/clinical-ai-monitor-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (61.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (61.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (48 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (48 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"clinical-ai-monitor"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py clinical-ai-monitor

Download the mock-data bundle (8 KB, 10 checks)expected.json

A four-sentence scribe note checked sentence by sentence against the transcript of a mock respiratory consultation (PriMock57, role-played, no patients), plus a two-item coverage checklist. The note carries one planted invented exam finding (oxygen saturation and respiratory rate from a phone consultation) that must not be supported, and the signed report must verify against the same transcript and note and fail once one verdict is changed.

What the rehearsal checks
  • all 4 clinical sentences were checked
  • every checked sentence got a verdict (none not judged or errored)
  • the planted exam finding (oxygen saturation 97%) is not supported
  • the coverage check answered for both checklist items
  • no checklist item errored
  • the signed report verifies against this server's key
  • the report matches the transcript it was made from
  • the report matches the note it was made from
  • a report with the exam verdict changed no longer verifies
  • every model call has a signed receipt

Licence: Transcript: PriMock57 mock primary-care consultation day1_consultation07 (Papadopoulos Korfiatis et al., 2022, github.com/babylonhealth/primock57), CC BY 4.0; role-played by clinicians and actors, no patients. The note is synthetic, written for this demo with one planted error. See inputs/ATTRIBUTION.txt.

Prompt for your coding agent

# Decosa Clinical AI assurance monitor: run it yourself (containers)

You are setting up the Decosa clinical AI assurance monitor inside our network, so transcripts and notes (PHI) never
leave it. It checks each scribe note against its visit transcript, lists what the note left out, and signs per-visit
reports that roll up into rates per vendor, specialty and week. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/clinical-ai-monitor.zip (8 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py clinical-ai-monitor` (the api image carries the same bundle under /app/rehearsal/clinical-ai-monitor/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py clinical-ai-monitor --bundle clinical-ai-monitor.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "all 4 clinical sentences were checked", "every checked sentence got a verdict (none not judged or errored)", "the planted exam finding (oxygen saturation 97%) is not supported"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For `api` set `DECOSA_LLM_ROUTE=direct`,
   `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_LOCAL_SIGNING=on`, and bind every port to
   127.0.0.1. For audio input, run the diarizer (`services/diarize` in the decosa-api repo) and set
   `DECOSA_DIARIZE_URL`; otherwise leave it empty and send transcripts.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/monitor/info` lists the finding types and `judge.route: "direct"`;
   `GET /attest/signing-key` shows this box's public key. Show me the key: the committee pins it to check our reports.
5. Smoke test: get a token with `POST /demo/session {"vertical":"clinical-ai-monitor"}`, take encounter 1 from
   `GET /monitor/samples`, and run `POST /monitor/check` with its `segments` and the `speaker` note. Expect
   `decision: "errors"` with the planted sentence flagged. Then `POST /monitor/verify` with the report and the note:
   `valid_signature` and `note_matches` must be true.
6. Report back: the public key, the smoke-test decision and counts, and how long the check took.

Off, and keep it off: this box holds PHI.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/clinical-ai-monitor-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsClinical AI assurance monitor on GeForce RTX 5090: use the Standard · judge plus diarizer (hosted demo) tier

The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB (fits): Text transcripts only: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for prompts up to about 16k tokens (a 25-minute visit). Estimate: same judge as the grounding check; not run here for this use case.

Standard · judge plus diarizer (hosted demo): what changesuses estimates

  • Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Monitor: decosa-api clinical AI assurance monitor (decosa_api/verticals/monitor). CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Judge: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
  • Audio in: MOSS-Transcribe-Diarize 0.9B. ~4 GB, weights 1.8 GB (estimate). MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate.
  • Detail checker: decosa-note-detail-checker-modernbert-large (M17, our own model; Apache-2.0 on Hugging Face). CPU. Runs on CPU (vram_gb 0 in stack.json).

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Clinical AI assurance monitor, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Clinical AI assurance monitor on my hardware

Fetch https://decosa.ai/prompts/clinical-ai-monitor-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=clinical-ai-monitor)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · judge plus diarizer (hosted demo) (standard). Fit check: runs with changes, about 32 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Monitor: decosa-api clinical AI assurance monitor (decosa_api/verticals/monitor), CPU
- Judge: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- Audio in: MOSS-Transcribe-Diarize 0.9B (OpenMOSS-Team/MOSS-Transcribe-Diarize), 4 GB
- Detail checker: decosa-note-detail-checker-modernbert-large (M17, our own model; Apache-2.0 on Hugging Face) (decosaai/decosa-note-detail-checker-modernbert-large), CPU

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%), MOSS-Transcribe-Diarize 0.9B ~4 GB (13%); about 0 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/clinical-ai-monitor-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 48 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh --profile live

Mac prompt for your coding agent

# Decosa Clinical AI assurance monitor: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Clinical AI assurance monitor on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 48 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/clinical-ai-monitor.zip (8 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py clinical-ai-monitor` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "all 4 clinical sentences were checked", "every checked sentence got a verdict (none not judged or errored)", "the planted exam finding (oxygen saturation 97%) is not supported"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Monitor: transcript and note parsing, sentence and section offsets, finding types, the signed report and the summary with intervals and drift (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured |
| Judge: one call per note sentence, one checklist extraction, one coverage check; also the speaker role map for anonymous labels | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |
| Detail checker (M17): after the judge, each drug, dose, frequency, route, date, duration, side and number in a sentence is read against its transcript lines and labelled same / changed / absent; a judge 'detail not in the visit' flag becomes a changed-detail error when p(changed) >= 0.5 | PyTorch on CPU (8 threads) | The same service with PyTorch on CPU (about 2.5 GB of RAM); expected to run, not timed on a Mac | Runs, not measured |
| Audio in (optional): one-pass speaker-attributed transcript of the whole visit | transformers on CUDA | MLX 8-bit (vanch007/mlx-MOSS-Transcribe-Diarize-8bit) on mlx-audio, scripts/mac/diarize_server.py | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 48 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh --profile live`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model, plus about 5 GB for speech recognition and diarization), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`, `"asr": true` and `"diarize": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py clinical-ai-monitor`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --profile live --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 28 Sep 2026 · measured 28 Sep 2026: · p50 28 s · p95 41 s · ~$0.021 per run · 30 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · assemble-prompt.md on our server: fresh clone into a clean directory, api image built from docker/api/Dockerfile, compose with a named volume, pointed at the already-running Qwen3.8-27B vLLM (127.0.0.1:8114) and MOSS diarizer (127.0.0.1:8092) instead of starting new ones; then torn down

Measured cost to run: about $0.024 per visit note (hosted, 28 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Images build, the service starts, and the smoke steps pass: the wrong-speaker sample flagged (17 s), the report verifies with transcript and note hashes, a changed count fails verification, the summary signs, and 170 s of audio came back as 22 speaker turns. Model-server startup itself was not re-verified. /healthz says ok: false on this stack because it also checks the live-scribe recogniser, which the monitor does not use.

Known limits (8)
  • Measured on synthetic notes with one clear planted error each; real scribe errors are subtler.
  • Changed details: 94% caught as errors with the detail checker (47/50 and 188/200 blind plants), 72-74% by the judge alone; the checker only upgrades the judge's own 'detail not in the visit' flags.
  • On real speech-recognition transcripts the detail checker adds some false error flags (4 in 1,469 ACI-Bench sentences, e.g. 'type i' vs 'type 1'); each comes with its transcript line.
  • Wrong-speaker errors are caught but typed correctly only 62% of the time.
  • Tighten shortened the 50 test notes by a median 15%; 5 of 362 recorded key items came back 'partly recorded' after tightening (none missing).
  • Measures against the transcript it is given; speech-recognition errors pass into the measure.
  • Clinicians' own notes get flagged for things never said aloud; use it on AI drafts.
  • Audio input is API-only (POST /monitor/transcribe); the page takes text.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the judge; the diarizer (optional, for audio) needs a second GPU slot.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the clinical ai assurance monitor API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.
The open stack

Checks any AI scribe's note against the visit transcript, sentence by sentence: invented facts, changed doses and details, exam findings nobody made, the wrong speaker and what was left out, each with the transcript line. Tightens the note without adding anything, and scores vendors in bulk.

Paste the note a commercial AI scribe wrote and the visit transcript (text, captions, a scribe export, or segments; audio goes through the API's diarizer). An open judge model checks every note sentence against the numbered, speaker-labelled transcript and types each problem: not in the visit, contradicts the visit (a changed dose, duration, side or number), the wrong speaker (the doctor's guess or a relative's history written as the patient's), or an exam finding nobody made. Each flag carries the note sentence and the transcript lines, or 'not in the transcript'. The medicines, allergy status, plan items and impression are read from the transcript and checked against the note for omissions. Tighten gives a shorter note that only deletes and merges, enforced in code, with the flagged sentences kept word for word, marked and listed for the clinician to fix. Buyer mode checks many visits from one or more scribe vendors and returns a signed scorecard by error type next to the checker's own measured catch and false-flag rates. Every visit ends in a signed report of hashes, line numbers and counts, with no text. It checks documentation against the conversation; it does not judge care.

Deployment
Self-host firstHosted demo on synthetic notes and transcripts; real patients: self-host (confidential on request, after a BAA).
Regulatory
A documentation-quality measure for the health system's own oversight, not a medical device function and not software that recommends care: it does not analyse images or signals, does not diagnose, and makes no recommendation about any patient's care; it compares a note with the conversation it came from. The closest statutory exclusion is software for administrative support of a health care facility (FD&C Act 520(o)(1)(A)); FDA's clinical decision support guidance (reissued 29 Jan 2026) covers software that recommends care, which this does not. Keep it that way: do not wire its flags into real-time changes to a patient's care. HIPAA: run it on your own hardware; measuring your own notes is a quality assessment activity, part of health care operations (45 CFR 164.501). A vendor that hosts it or can reach the PHI needs a business associate agreement. The hosted demo takes only PriMock57 mock consultations and synthetic notes. Governance: the Joint Commission and CHAI guidance on the Responsible Use of AI in Healthcare (17 Sep 2025) names clinical documentation in scope and asks for local validation, ongoing quality monitoring and error reporting to leadership and vendors; the Joint Commission's voluntary certification on it was announced 1 Jun 2026. ONC HTI-1's decision support interventions criterion (45 CFR 170.315(b)(11)) binds certified EHR developers, not hospitals, and HTI-5, proposed 29 Dec 2025 and not final when checked, would remove its source-attribute and risk-management duties; neither changes what a hospital may measure itself. State laws bind the providers that use scribes, not this monitor: California AB 3030 (disclaimers on AI-generated patient communications, from 1 Jan 2025), Texas SB 1188 (practitioners review AI-created records, from 1 Sep 2025) and TRAIGA (disclose AI use in care, from 1 Jan 2026), and Illinois Public Act 104-0054 (consent before recording therapy sessions for AI, 1 Aug 2025). Its reports can document the review these laws expect. Some scribe contracts restrict benchmarking or publishing results; check yours. Not legal advice. Model licences: Apache-2.0 (Qwen3.8-27B, MOSS-Transcribe-Diarize). Checked 25 Sep 2026.
Architecture
Text description

The visit audio (optional) goes to the MOSS diarizer, which returns a speaker-labelled transcript; or the transcript comes in as text. The transcript and the vendor's note go to the monitor, which splits the note into sentences. The Qwen3.8-27B judge checks each sentence against the numbered transcript lines and names a finding type, reads a checklist of medications, allergies, plan items and the impression from the transcript, and checks which of them the note records. The monitor signs a per-encounter report of hashes and counts, and signed reports roll up into a signed summary with rates per vendor, specialty and week and a drift test. Everything inside the hospital boundary stays local when self-hosted; on the hosted demo each judge call gets a receipt that our gateway countersigns.

Architecture

At a glance

What leaves the box
Nothing when self-hosted (DECOSA_LLM_ROUTE=direct). The hosted demo takes public mock visits only.
Retention
None: transcripts, notes and audio stay in memory for one request; logs carry counts.
What you keep
A signed report per visit (hashes, line numbers, counts; no text, the ISO week not the date) and signed summaries.
Inputs
Transcript as text with speaker labels, WebVTT or SRT captions, a scribe export with timestamps, or segments (up to 40,000 characters), or WAV audio up to 15 minutes through the API; the note as text (up to 12,000 characters). Buyer mode: up to 25 visits per call hosted, no limit self-hosted.
Typical cost
A few cents or less per note at the gateway list price (measured on the demo notes), rising with note and transcript length, including tighten; the detail checker runs on CPU at no token cost. Self-hosted: your own hardware.
Hardware
One GPU with 32 GB or more for the judge; the diarizer (audio only) needs about 4 GB more; the detail checker runs on CPU (8 threads, about 2.5 GB of RAM).
Output
Per sentence: the flag type, the note sentence and the transcript lines (or 'not in the transcript'); key items left out; a tightened note; a signed per-visit report; in buyer mode, a signed scorecard per vendor and version.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    text transcripts, one 32 GB card

    The judge only: send transcripts as text (most scribe products export one). No audio input.

    Models
    • decosa-api clinical AI assurance monitor (decosa_api/verticals/monitor)
    • Qwen3.8-27B (NVFP4)
    • decosa-note-detail-checker-modernbert-large (M17, our own model; Apache-2.0 on Hugging Face)
    Hardware
    1x RTX 5090 32 GB (estimate)
    Quality evidence
    • Same judge and prompts as standard, so the same detection and false-flag ratessee standarddocs/evals/clinical-ai-monitor.md
    Latency
    estimate: as standard on a smaller card; not timed separately.
    Verification
    Proof: strongSelf-host onlySelf-hosted calls are attested with the box's key; on a network provider they carry gateway receipts.
  • In the hosted demo

    Standard

    judge plus diarizer (hosted demo)

    Adds audio input through the MOSS diarizer. This is what the hosted API runs.

    Models
    • decosa-api clinical AI assurance monitor (decosa_api/verticals/monitor)
    • Qwen3.8-27B (NVFP4)
    • decosa-note-detail-checker-modernbert-large (M17, our own model; Apache-2.0 on Hugging Face)
    • MOSS-Transcribe-Diarize 0.9B
    Hardware
    1x RTX PRO 6000 96 GB (measured), diarizer on a second slot
    Quality evidence
    • Planted errors caught at error severity (50 held-out PriMock57 visits, one error per note): invented fact / invented exam finding / wrong speaker / key item left out / changed detail49/50 / 49/50 / 47/50 / 45/50 / changed detail 47/50 with the detail checker (36/50 judge alone); every plant flagged at least as review (50/50 each)docs/evals/clinical-ai-monitor.md
    • Changed details, 200 more blind plants on the same 50 held-out visits (8 detail types)188/200 (94%) with the detail checker, 147/200 (73.5%) judge alone; false error flags on faithful notes unchanged (4/1,135)docs/evals/clinical-ai-monitor.md (M17)
    • Wrong-speaker errors typed as wrong speaker31/50 (62%); the rest flagged as contradicts or not in the visitdocs/evals/clinical-ai-monitor.md
    • Error flags on faithful notes (1,135 sentences, 368 key items)4 sentences (0.35%; 1 real error in the note, 2 judge mistakes, 1 debatable), 3 key items (0.8%; 1 real)docs/evals/clinical-ai-monitor.md
    • Clinicians' own PriMock57 notes: lines with an error finding12.1%; in a sample of 20, 16 were statements the transcript does not support (names, routine negatives never asked)docs/evals/clinical-ai-monitor.md
    • Speech recognition, if you send audio (MOSS-Transcribe-Diarize, PriMock57)10.3% WERscribe-bench RESULTS.md (the clinical scribe's pass 2)
    Latency
    measured on our server: seconds per note straight to the model server; longer through our gateway, and a few minutes for the recorded demo runs when the shared gateway was loaded.
    Verification
    Proof: strongEvery judge call is a separate gateway call with a gateway-signed receipt; the report lists them.
  • Needs more compute

    Wanted: the best setup

    a second judge from another family

    DeepSeek-V4-Flash re-judges the transcripts to check the first judge's rates, where wrong-speaker and changed-detail errors are hardest. Patient data stays on your hardware, never on community providers. Not served yet.

    Models
    • decosa-api clinical AI assurance monitor (decosa_api/verticals/monitor)
    • Qwen3.8-27B (NVFP4)
    • decosa-note-detail-checker-modernbert-large (M17, our own model; Apache-2.0 on Hugging Face)
    • MOSS-Transcribe-Diarize 0.9B
    • DeepSeek V4 Flash
    Hardware
    Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash beside the standard card, or a Mac with 192 GB or more (MXFP4 MLX build, 156 GB on our Mac; not measured). Estimate.
    Quality evidence
    • This eval, same protocolnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.

Also runs on

  • A second judge on one cardNemotron-3-Super-120B-A12B (NVFP4)not servedA cheaper cross-family check that fits beside the 27B: smaller than DeepSeek-V4-Flash, so we would host it ourselves. sm_120 support unconfirmed. Hardware: 1x RTX PRO 6000 96 GB (80 GB of weights).

We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.

Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Monitor: transcript and note parsing, sentence and section offsets, finding types, the signed report and the summary with intervals and drift (no model; CPU)decosa-api clinical AI assurance monitor (decosa_api/verticals/monitor)
0 GBProof: partial
Judge: one call per note sentence, one checklist extraction, one coverage check; also the speaker role map for anonymous labelsQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Detail checker (M17): after the judge, each drug, dose, frequency, route, date, duration, side and number in a sentence is read against its transcript lines and labelled same / changed / absent; a judge 'detail not in the visit' flag becomes a changed-detail error when p(changed) >= 0.5decosa-note-detail-checker-modernbert-large (M17, our own model; Apache-2.0 on Hugging Face)decosaai/decosa-note-detail-checker-modernbert-large on Hugging Face (opens in a new tab)
395M · 0 GBProof: partial
Audio in (optional): one-pass speaker-attributed transcript of the whole visitMOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab)
0.9BProof: partialIn the hosted demo
284B (13B active) · 166 GBNo proof yetSelf-host only
One-card second judge from another model familyNemotron-3-Super-120B-A12B (NVFP4)nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 on Hugging Face (opens in a new tab)
120B (12B active) · about 80 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

  • 57 mock primary-care consultations, role-played by clinicians and actors, with per-speaker transcripts and the clinicians' own notes: the eval set and the hosted demo's visits.

  • scripts/monitor_eval.py and docs/evals/clinical-ai-monitor.mdApache-2.0

    The eval: synthetic scribe notes for all 57 visits with five planted error types, false flags on the clean notes, and the flag rate on the clinicians' own notes.

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /monitor/info, /monitor/samples, /monitor/demo-dashboard; POST /monitor/check (SSE or JSON), /monitor/transcribe (audio), /monitor/verify, /monitor/summary. Keeps nothing: transcripts, notes and audio stay in memory for the request.

  • vLLM (judge):8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).

  • decosa-diarize (optional):8092
    built from services/diarize in decosa-api

    MOSS-Transcribe-Diarize for audio input (POST /v1/diarize with PCM16 16 kHz mono).

Hardware

  • 1x RTX 5090 32 GB Fits

    Text transcripts only: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for prompts up to about 16k tokens (a 25-minute visit). Estimate: same judge as the grounding check; not run here for this tool.

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the hosted demo and the eval ran on this card, shared with other services; the diarizer ran on a second card.

Latency per lane

  • one note with tighten, hosted demo through the gateway (18-30 model calls)28.0 s

    Measuredmeasured on our server 2026-09-28: median of 9 runs of the three note samples (10-41 s), gateway shared with other work; the tighten proposal runs alongside the check

  • the check alone, same runs23.0 s

    Measuredmeasured on our server 2026-09-28: median of the same 9 runs' report time (10-36 s)

  • one note (about 23 checked sentences) straight to the model server, 4 calls in flight12.6 s

    Measuredmeasured on our server 2026-09-25: median over 50 clean notes in the eval

  • audio to transcript (MOSS diarizer), 170 s of audio10.3 s

    Measuredmeasured on our server 2026-09-25, one PriMock57 clip

Notes

  • Findings are typed: not in the visit, contradicts the visit, wrong speaker, invented exam finding, and left out are errors; a detail not in the visit and a partly recorded item are worth a look. A sentence that could not be judged is counted as not judged, never guessed.
  • A changed detail the judge only calls 'a detail not in the visit' is settled by the detail checker (our M17 model, CPU): if the note's value conflicts with the transcript line, the flag becomes an error. It never removes or lowers a judge flag, and if the checker is down the judge's verdicts stand.
  • Tighten only removes or condenses. Code checks every proposed sentence: its words must appear in the source sentences in the same order (a few joining words aside), a dropped negation takes the rest of its sentence, a dropped hedge or condition its whole sentence, and every number, dose, frequency, side and medicine name stays. What fails goes back to the model once, then falls back to the original words.
  • The signed report holds hashes, offsets, transcript line numbers, verdicts and counts. It holds no note or transcript text, the visit week instead of the date, and the encounter reference only as a SHA-256, so it can go in an ordinary log or to a committee.
  • It measures the note against the transcript, not against the patient: errors in the transcript (speech recognition) pass into the measure. The eval used reference transcripts.
  • Clinicians' own notes record things never said aloud (brand names, standard negatives, their own examination), and the checker flags those too. It is a check for AI drafts, not a style check for people.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

clinical-ai-monitor/assemble-prompt.md149 lines
# Assemble the Decosa clinical AI assurance monitor on this machine

You are setting up a monitor that measures an ambient AI scribe's notes against the visit transcripts, inside this
hospital's network. For each visit it takes the transcript (or the audio) and the note the scribe product wrote, checks
every note sentence against the transcript, lists the medications, allergy status, plan items and impression the note
left out, and signs a per-visit report of hashes and counts. Signed reports roll up into rates per vendor, specialty and
week. Work step by step, show me each command before you run anything with `sudo`, and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/clinical-ai-monitor.zip (8 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py clinical-ai-monitor` (the api image carries the same bundle under /app/rehearsal/clinical-ai-monitor/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py clinical-ai-monitor --bundle clinical-ai-monitor.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "all 4 clinical sentences were checked", "every checked sentence got a verdict (none not judged or errored)", "the planted exam finding (oxygen saturation 97%) is not supported"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Models: Qwen3.8-27B (Apache-2.0) as the judge; MOSS-Transcribe-Diarize 0.9B (Apache-2.0) only if we send audio.
  decosa-api is AGPL-3.0-or-later.
- Transcripts and notes are PHI. Bind every port to 127.0.0.1 (or the internal interface I name), never expose it to the
  internet, and keep `DECOSA_LLM_ROUTE=direct` so no model call leaves this machine. The service keeps nothing: text
  and audio stay in memory for one request, and logs carry counts only. Keep it that way; do not add request logging.
- It is a documentation-quality measure, not software that recommends care: do not wire its flags into anything that
  changes a patient's care in real time. Check with me whether our scribe contract allows us to benchmark the product.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB for the judge (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV
   cache; a 25-minute visit is about 10k tokens of transcript, sent first in every prompt). Driver 570 or newer.
   Blackwell cards run NVFP4; older cards use `Qwen/Qwen3.8-27B-FP8`.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
   from the official Docker and NVIDIA repositories after asking me, then run
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 30 GB free for the model and images.
4. Audio (optional): the diarizer needs its own GPU memory (about 4 GB) and a Python 3.12 venv with PyTorch.

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a release that contains
  `decosa_api/verticals/monitor/`, and `docker build -f docker/api/Dockerfile -t decosa-api:local .`
- `vllm/vllm-openai:v0.29.0` for the judge; weights `nvidia/Qwen3.8-27B-NVFP4` (revision
  `482ca0f3832238542f8f5295dde86b5f22711d80`), or `Qwen/Qwen3.8-27B-FP8` on pre-Blackwell cards.
- Audio only: from the same clone, `uv venv .venv-diarize --python 3.12` and
  `uv pip install -p .venv-diarize -r services/diarize/requirements.txt --index-strategy unsafe-best-match --extra-index-url https://download.pytorch.org/whl/cu130`,
  then run `DIARIZE_HOST=127.0.0.1 DIARIZE_PORT=8092 .venv-diarize/bin/python services/diarize/server.py` (it downloads
  `OpenMOSS-Team/MOSS-Transcribe-Diarize`, 1.8 GB). `curl -s 127.0.0.1:8092/health` must say `"ok": true`.

## 3. docker-compose.yml
Write this in `~/decosa/monitor/`:

```yaml
services:
  llm:
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--revision", "482ca0f3832238542f8f5295dde86b5f22711d80",
              "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768", "--enable-prefix-caching",
              "--kv-cache-dtype", "fp8_e4m3", "--gpu-memory-utilization", "0.85"]
    ports: ["127.0.0.1:8114:8000"]
    ipc: host
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck:
      test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"]
      interval: 30s
      retries: 40
      start_period: 900s
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>        # or decosa-api:local from step 2
    ports: ["127.0.0.1:8445:8445"]
    extra_hosts: ["host.docker.internal:host-gateway"]
    environment:
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_DIARIZE_URL: http://host.docker.internal:8092   # empty string if we only send transcripts
      DECOSA_ASR_WS: ws://127.0.0.1:1/unused                 # the live scribe is not part of this stack
      DECOSA_BUDGET_LLM_TOKENS: "400000"                      # per session; a note needs about 170 tokens a sentence
      DECOSA_BUDGET_AUDIO_S: "7200"
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_LOCAL_SIGNING: "on"
      DECOSA_SIGNER_NAME: "<hospital> quality monitor"
      DECOSA_MONITOR_MAX_CONCURRENT: "3"
    volumes: ["decosa-data:/data"]
    depends_on: { llm: { condition: service_healthy } }
volumes:
  decosa-data:
```

If the diarizer runs on the host, bind it to the Docker bridge address instead of 127.0.0.1 so the container can
reach it, or leave `DECOSA_DIARIZE_URL` empty and send transcripts only. On first start the api service creates this
box's Ed25519 key in the `decosa-data` volume. Back it up and never print it: reports and summaries are signed with it
(an attestation by us, the operator, not a proof), and the committee pins its public key.

## 4. Smoke test
1. `docker compose up -d`, then wait for `llm` to be healthy (the first start downloads about 20 GB).
2. `curl -s localhost:8445/monitor/info | jq '{findings: (.findings|keys), judge, diarizer}'`. (`/healthz` reports
   `ok: false` on this stack because it also checks the live-scribe recogniser, which we do not run; `llm: true` is what
   matters here.)
3. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"clinical-ai-monitor"}' | jq -r .token)`.
4. Run a bundled sample with a planted wrong-speaker error:
   `curl -s localhost:8445/monitor/samples | jq '{segments: .encounters[1].segments, note: (.encounters[1].notes[] | select(.id=="speaker") | .note), vendor: "Scribe A", specialty: .encounters[1].specialty}' > req.json`
   `curl -s -XPOST localhost:8445/monitor/check -H "authorization: Bearer $T" -H 'content-type: application/json' -d @req.json > run.json`
   Expect `decision: "errors"` and `.sentences[] | select(.finding)` to include the planted sentence (compare
   `.encounters[1].notes[]|select(.id=="speaker")|.planted` from the samples). Run the `clean` note too: expect `clean`
   or at most one flag.
5. Verify: `jq -s '{report: .[1].report, note: .[0].note, segments: .[0].segments}' req.json run.json | curl -s -XPOST localhost:8445/monitor/verify -H 'content-type: application/json' -d @-`
   must show `valid_signature`, `transcript_matches` and `note_matches` true. Change one count in the report and verify again: it must fail.
6. Summary: `jq '{reports: [.report]}' run.json | curl -s -XPOST localhost:8445/monitor/summary -H 'content-type: application/json' -d @-`
   returns rates and a signed `summary`.
7. Audio (if the diarizer runs): `curl -s -XPOST localhost:8445/monitor/transcribe -H "authorization: Bearer $T" --data-binary @visit.wav`
   returns `lines` with `Clinician` and `Patient` turns; pass its `segments` to `/monitor/check` with `"source": "audio"`.
8. Time it and tell me: on our RTX PRO 6000, shared with other work, a 27-sentence note against a 14-minute visit took
   about 15-20 s straight to the model server, almost all of it judge calls.

## 4b. The detail checker (optional, CPU)
Our M17 model settles the judge's "a detail not in the visit" flags: when the note's drug, dose, frequency, route, date,
duration, side or number conflicts with the transcript line, the flag becomes a changed-detail error (changed details
caught: 94% with it, 72% without, on our held-out mock visits). It needs no GPU: about 2.5 GB of RAM and 8 CPU threads.
The weights are not published yet (their model card is under review); until they are, ask us for them, or skip this
step: without `DECOSA_DETAIL_URL` the tool runs the judge alone and says so in `/monitor/info` (`detail_checker.configured: false`).
1. `uv venv ./detail-venv --python 3.12` and `uv pip install -p ./detail-venv -r services/detail_checker/requirements.txt
   --index-url https://download.pytorch.org/whl/cpu --extra-index-url https://pypi.org/simple --index-strategy unsafe-best-match`.
2. Put the weights in `./detail-weights/` and start it from the decosa-api checkout:
   `CUDA_VISIBLE_DEVICES= DETAIL_MODEL_DIR=./detail-weights DETAIL_PORT=8495 ./detail-venv/bin/python services/detail_checker/server.py`.
3. Give the api service `DECOSA_DETAIL_URL=http://host.docker.internal:8495` (or the host's address) and restart it;
   `curl -s localhost:8445/monitor/info | jq .detail_checker` shows `reachable: true` and the weights sha256.

## 5. Wire it in
Keep reports, not text: for each sampled visit, call `POST /monitor/check` with the vendor, version, specialty and the
ISO week, append `report` to a JSON-lines file, and once a week or month send those reports to `POST /monitor/summary`
as `{"reports": [...], "pin": true}` (`jq -s '{reports: ., pin: true}' reports.jsonl`). The signed summary lists every report it counts; that is the committee's evidence. Use
`encounter_ref` for your own opaque visit id (the report keeps only its hash), never a medical record number. Point the
Decosa site at it with `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` if you want the console. Contract:
`API_CONTRACT.md`, section "Clinical AI assurance monitor".
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

2 laws, rules and guidance pages cited; 2 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Checks any AI scribe's note against the visit transcript, sentence by sentence: invented facts, changed doses and details, exam findings nobody made, the wrong speaker and what was left out, each with the transcript line. Tightens the note without adding anything, and scores vendors in bulk.
Who it's for
Teams in healthcare and compliance and trust.
Where it runs
Self-host (patient data stays on site)
Key numbers
  • 49 (98%) Invented fact caught as an error (test split, n = 50)
  • 47 (94%) Changed detail caught as an error (test split, n = 50)
  • 45 (90%) Key item left out caught (test split, n = 50)
  • 28.0 s Median end-to-end run, hosted (QA sweep 2026-09-28)
All results, datasets and caveats
Models
Qwen3.8-27B · MOSS diarizer (audio)
Where
Self-host (patient data stays on site)
Checks
Receipt per verdict; signed reports and summary
Output
Structured data · Signed record or verdict
Data
Patient data (PHI)
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

How does Check an AI note measure an ambient AI scribe?

Check an AI note takes the visit transcript, or audio it diarizes into one, and the note the scribe wrote. An open judge model, Qwen3.8-27B, checks every note sentence against the numbered, speaker-labelled transcript and types each problem: not in the visit, contradicts the visit, changed detail, wrong speaker, or invented exam finding. A checklist of medications, allergy status, plan items and impression finds omissions. It measures documentation quality; it does not judge care.

Where do I get the transcript to check an AI scribe's note?

Check an AI note needs what was said in the visit: the transcript your scribe product shows or exports (text, WebVTT or SRT captions, or a timestamped export all work), or the visit audio through the API, which a speaker-separating recogniser turns into a transcript. If your scribe keeps neither the audio nor the transcript, its notes cannot be checked this way.

How accurate is Check an AI note?

On 50 held-out PriMock57 mock visits with synthetic notes, Check an AI note caught 94% of changed doses and details as errors (47 of 50; 188 of 200 more blind plants), 98% of invented facts and invented exam findings, 94% of wrong-speaker errors and 90% of left-out key items, and flagged 0.35% of faithful sentences. The 94% for changed details uses our detail checker after the language model; the language model alone caught 72%. Each note had one clean planted error, so these rates are an upper bound, and real scribe notes were not measured.

Does PHI leave the hospital when the checker runs?

Self-hosted on the direct route, nothing leaves the box: transcripts, notes and audio stay in memory for one request and logs carry counts. What Check an AI note keeps is a signed per-visit report of hashes, line numbers and counts, with the ISO week instead of the date and no text. Measuring your own notes is a quality assessment activity within HIPAA health care operations (45 CFR 164.501). The hosted demo takes mock consultations only.

Is Check an AI note a medical device?

It is designed not to be one. Check an AI note compares a note with the conversation it came from; it does not diagnose, analyse images or signals, or recommend care. The closest statutory exclusion is administrative support software under FD&C Act 520(o)(1)(A), and FDA's clinical decision support guidance covers software that recommends care. Do not wire its flags into real-time care decisions. Not legal advice.

How does the checker relate to the Joint Commission's Responsible Use of AI guidance?

The Joint Commission and CHAI guidance (17 Sep 2025) names clinical documentation in scope and asks for local validation, ongoing quality monitoring and error reporting to leadership and vendors; a voluntary certification based on it was announced 1 Jun 2026. Check an AI note rolls signed per-visit reports into rates, intervals and a drift test per vendor, specialty and week for the governance committee. It supplies evidence; it does not certify anyone.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Clinical AI assurance monitor

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.