{"schema_version":"1","site":"https://decosa.ai","id":"clinical","num":"01","name":"Visit copilot","tool_name":"Run the visit with a copilot","short":"Copilot","blurb":"The visit copilot: a live transcript with speaker labels, guidance on what is still worth asking (each item with its source), a note checked sentence by sentence against the visit, codes, prior-auth prep, and the work note, FMLA sections and patient instructions drafted, never signed or sent.","status":"live","labels":{"industry":["healthcare"],"job":["transcribe","draft"],"input":["voice"],"deploy":["selfhost"],"demo":"synthetic","status":"live","output":["text","data"],"data":["phi"],"hardware":"gpu-96","licence":"permissive"},"industries":["healthcare"],"runs_in":["selfhost","crp-phi-research"],"part_of":[],"built_from":["live-asr","diarize","grounding","formfill"],"models":"Voxtral 4B · MOSS diarizer · Qwen3.8-27B · M17 detail checker","where":"Self-host (patient data stays on site)","hardware":"1× RTX PRO 6000 (96 GB), or 2× RTX 5090","final_artifact":"The self-checked note, validated codes, the paperwork drafts, and the clinical considerations to accept or dismiss.","self_host_first":true,"verification":{"receipt_coverage":"full","summary":"Note checked before it is final","manual_qa":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":60000,"p95_ms":71600,"runs":null,"receipts_per_run":80,"cost_per_run_usd":0.07},"selfhost":{"date":"2026-09-29","result":"pass","method":"fresh clone of the branch into a clean directory, api image built from docker/api/Dockerfile, compose with a named volume and DECOSA_CLINIC_PROFILE, pointed at the model servers already running on our server (Qwen3.8-27B, Voxtral, MOSS diarizer, M17) instead of starting new ones; then torn down","notes":"The image builds and starts; /healthz ok with asr, llm and diarize true. The copilot sample at 2x passed end to end: 28 speaker turns, guidance, a self-checked note (28 sentences, M17 on), codes from the assessment, three paperwork drafts with the clinic profile from the environment, 0 signatures; done 44.5 s after the audio on a GPU shared with our evals. Model-server startup itself was not re-verified."},"known_limits":["Hosted is for synthetic visits only; real visits must be self-hosted (no BAA yet).","Speaker labels were tested on two-speaker visits; a third speaker is labelled Other.","The self-check compares the note with the visit's own transcript, so a word the recogniser misheard passes it. The note marks sentences that may rest on one (where the two recognisers disagree, or a word is unknown): 8 of 17 such sentences on held-out visits, with 8% of good sentences also marked. Check names, numbers and yes/no answers against the cited turn.","WH-380-E drafts still need every box checked (91% of filled boxes right on held-out visits; vague essential-function wording and incomplete date lists are the usual misses); third-party insurer FMLA forms are not bundled yet.","No EHR write-back yet: copy the note and the drafts, or use the API.","Guidance is a 26-topic public-domain pack plus the CMS history elements; a missing item does not mean nothing is missing."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Medical-term miss rate, live ASR (Voxtral Mini 4B Realtime)","value":"8.4%","unit":null,"n":57,"split":"heldout","note":"146 of 1,741 lexicon terms; Nemotron-3.5 streaming 12.7%; WER 13.2 (Nemotron-3.5 streaming 11.8; MOSS-TD pass 2 10.3)"},{"name":"Speaker labels during the visit, word-level role accuracy (rolling MOSS windows)","value":"99.71%","unit":null,"n":11,"split":"heldout","note":"PriMock57, 110 min of audio; whole-recording pass 99.97%; worst consultation 98.1%"},{"name":"Considerations: warning-feature encounter recall (synthetic, held-out)","value":"14/15 and 13/15","unit":null,"n":15,"split":"heldout","note":"two runs; dev 7/7 both; topic recall held-out 13/16 and 12/16"},{"name":"Guidance: false-alarm encounters","value":"1/16 and 1/16","unit":null,"n":16,"split":"heldout","note":"both flags outside the guideline pack (tea-coloured urine on a statin; drowsy driving), shown as 'no bundled guideline'"},{"name":"Guidance: directive wording shown","value":"0","unit":null,"n":null,"split":"heldout","note":"all runs, the live lint"},{"name":"Guidance items judged useful at that moment (blind clinician judge)","value":"checklist 95/145 (66%); on real GP consultations 20/31 (65%)","unit":null,"n":425,"split":"heldout","note":"blind clinician judge: 76 moments in 38 synthetic visits and 24 in 12 PriMock57 consultations; warning-feature items 4/10 and 1/5 useful, the rest already covered; history elements 11% and 18% useful, so they fold away during the visit; 3 of 523 items judged harmful, all synthetic"},{"name":"Note sentences supported by the human transcript (blind judge, held-out)","value":"289/308 (93.8%)","unit":null,"n":308,"split":"heldout","note":"PriMock57, 11 new consultations; the 28 Sep writer 229/254 (90.2%) on the same visits; writer-caused misses 12 -> 3; 16 misheard by speech recognition, which a transcript self-check cannot see"},{"name":"WH-380-E boxes filled right (held-out FMLA visits written blind)","value":"91.4% (85/93)","unit":null,"n":8,"split":"heldout","note":"28 Sep code 79.6% on the same visits; 85.3% vs 75.7% by code scoring alone (free text judged blind otherwise); false fills 10 -> 2. Dev: work/school note 90%, instructions 97%"},{"name":"Signatures filled by a paperwork draft","value":"0","unit":null,"n":28,"split":"heldout","note":"signature and signing-date boxes"},{"name":"Guidance checklist items the GP went on to ask (PriMock57, real GPs)","value":"25/43 (58%)","unit":null,"n":43,"split":"heldout","note":"history elements 88/107 (82%); what the lane listed that the clinician asked later in the same visit"}],"dataset":"PriMock57 (57 recorded mock primary-care consultations, CC BY 4.0) for speech, speaker labels and note grounding; 38 synthetic US visits with planted warning-feature labels (written and labelled by separate agents) for guidance; 12 synthetic visits with paperwork truth written blind by a separate agent. Judges: Claude Code Opus 5.5 as blind sub-agents.","held_out":true,"caveats":["Synthetic visits and mock consultations, not real clinic audio.","Two-speaker visits only.","The judges are models (Claude Code Opus as blind sub-agents), not clinicians.","The guidance was tested on a 26-topic pack.","Costs are at list price; eval runs used the direct route (same weights), the e2e runs the gateway."],"date":"2026-09-26","doc_url":null},"stack":{"summary":"One screen for the whole visit. Live captions come first, then speaker-labelled turns a few seconds behind (the MOSS diarizer on rolling windows). While the patient is in the room, a guidance lane lists history and checklist items still worth covering and considerations when a warning feature comes up, each with the words that triggered it and a cited public source; nothing there is an instruction or an alarm, and none of it enters the note. After the visit the note is written from the speaker turns with a citation on every sentence and checked (Check an AI note plus our M17 detail checker) before it is shown as final; codes follow the clinician's stated assessment; and the patient instructions, a work or school note and the FMLA WH-380-E provider sections are drafted from the visit, never signed or sent. It runs on the clinic's own GPU.","tagline":"The visit copilot: transcribes with speaker labels, reminds you what is still worth asking (each item with its source), writes a note checked sentence by sentence before you see it, codes it, and drafts the paperwork, on your own GPU.","deployment":"self-host-first","regulatory_note":"HIPAA: run it on the clinic's own hardware so audio, transcripts and notes stay on-site; there is no hosted BAA yet, so the hosted demo takes synthetic visits only. Drafts for clinician review; not a diagnostic device. The guidance lane is designed to stay outside FDA's device definition under FD&C Act 520(o)(1)(E) (FDA guidance on software functions, January 2026): no image or signal analysis (the audio is only transcribed), clinician only, considerations rather than directives, the basis (transcript words and a public source) shown on every item, no alerting or time pressure, nothing written into the note. That is our reading, not a legal opinion: counsel review of the in-visit use is pending before the guidance lane is marketed. Paperwork drafts never fill signatures, signing dates or attestations and are never sent; the clinician reviews, signs and sends. California AB 3030 exempts only clinician-reviewed AI messages to patients. CPT codes are AMA-licensed and hidden in the demo.","components":[{"id":"asr-live","role":"Pass 1: live streaming transcript for the in-visit view (no speakers)","name":"Voxtral Mini 4B Realtime","hf_repo":"mistralai/Voxtral-Mini-4B-Realtime-2602","license":"Apache-2.0","params":"4.4B","quant":"BF16","vram_gb":24,"memory_gb_estimate":null,"engine":"vLLM realtime WebSocket (/v1/realtime)","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["all"],"alternative_to":null},{"id":"asr-live-alt","role":"Pass 1 alternative: streaming ASR (cache-aware, 1.1 s chunks)","name":"Nemotron-3.5-ASR-Streaming 0.6B","hf_repo":"nvidia/nemotron-speech-streaming-en-0.6b","license":"NVIDIA Open Model License","params":"0.6B","quant":null,"vram_gb":null,"memory_gb_estimate":null,"engine":"NeMo toolkit (cache-aware streaming)","receipt_coverage":"partial","in_hosted_demo":false,"tiers":["alternative"],"alternative_to":"asr-live"},{"id":"asr-pass2","role":"Speaker labels during the visit: rolling windows (every 15 s of new audio, 6 s overlap) re-transcribed with speaker labels; the committed turns become the visit transcript the note cites","name":"MOSS-Transcribe-Diarize 0.9B","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize","license":"Apache-2.0","params":"0.9B","quant":"BF16","vram_gb":null,"memory_gb_estimate":null,"engine":"transformers (trust_remote_code) + moss_transcribe_diarize package; hosted: decosa-diarize service, 127.0.0.1:8092, GPU0","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"asr-pass2-alt","role":"Pass 2 alternative, part 1: offline transcript","name":"Parakeet-TDT 0.6B v3","hf_repo":"nvidia/parakeet-tdt-0.6b-v3","license":"CC-BY-4.0","params":"0.6B","quant":null,"vram_gb":null,"memory_gb_estimate":null,"engine":"NeMo toolkit","receipt_coverage":"partial","in_hosted_demo":false,"tiers":["alternative"],"alternative_to":"asr-pass2"},{"id":"diar-pass2-alt","role":"Pass 2 alternative, part 2: speaker diarization (up to 4 speakers)","name":"Sortformer diarizer 4spk v1","hf_repo":"nvidia/diar_sortformer_4spk-v1","license":"CC-BY-NC-4.0 (non-commercial)","params":"0.12B","quant":null,"vram_gb":null,"memory_gb_estimate":null,"engine":"NeMo toolkit","receipt_coverage":"partial","in_hosted_demo":false,"tiers":["alternative"],"alternative_to":"asr-pass2"},{"id":"llm-lite","role":"Language model for the lite tier: live lanes and the final note from the live transcript","name":"Qwen3.8-27B (official FP8)","hf_repo":"Qwen/Qwen3.8-27B-FP8","license":"Apache-2.0","params":"27.8B","quant":"FP8","vram_gb":null,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 (same decosa-llm image)","receipt_coverage":"strong","in_hosted_demo":false,"tiers":["lite"],"alternative_to":null},{"id":"llm","role":"Language model: live SOAP draft, guidance report, practitioner lanes, window role map, cited note, the self-check's sentence judge, assessment codes and the paperwork field mapping","name":"Qwen3.8-27B (NVIDIA NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP) + FP8 (attention/GDN), FP8 KV cache, MTP speculative decoding k=3","vram_gb":57,"memory_gb_estimate":null,"engine":"vLLM 0.29.0","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"writer-alt","role":"Note writer alternative: the loop's best writer","name":"DeepSeek V4 Flash","hf_repo":"deepseek-ai/DeepSeek-V4-Flash","license":"MIT","params":"284B","quant":"FP8 + FP4 experts (official mixed precision)","vram_gb":166,"memory_gb_estimate":null,"engine":"vLLM, tensor parallel 2, Blackwell build with DSpark speculative decoding","receipt_coverage":"strong","in_hosted_demo":false,"tiers":["best"],"alternative_to":"llm"},{"id":"detail-checker","role":"Note self-check, detail step (M17): each drug, dose, frequency, date, side and number in a note sentence read against its transcript lines; a 'detail not in the visit' flag becomes a changed-detail error","name":"decosa-note-detail-checker-modernbert-large (M17, our own model)","hf_repo":"decosaai/decosa-note-detail-checker-modernbert-large","license":"Apache-2.0","params":"395M","quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"PyTorch CPU (services/detail_checker)","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["standard","best"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · one 48 GB card, live pass only","summary":"Live transcript, lanes and a final note written from the streaming transcript. No speaker labels, no citations, no verifier; the pipeline the benchmark calls vanilla.","components":["asr-live","llm-lite"],"hardware":"1x L40S or RTX 6000 Ada 48 GB (FP8 weights; starting-point settings, not measured)","quality_evidence":[{"metric":"ACI-Bench ROUGE-L (Qwen3.8-27B FP8, human transcript)","value":"34.2","source":"scribe-bench wiki models.md / RESULTS.md"},{"metric":"PriMock57 note composite, vanilla pipeline (streaming ASR -> Qwen3.8-27B, official weights, test 37)","value":"41.01","source":"scribe-bench wiki vanilla-vs-best"},{"metric":"Medical-term miss / WER, live ASR (Voxtral Mini 4B Realtime, PriMock57, 57 visits)","value":"8.4% / 13.2 (Nemotron-3.5 streaming: 12.7% / 11.8)","source":"scribe-bench asr_score on PriMock57 (57 visits) through the live realtime endpoint; eval results file asr-voxtral-primock57.json (2026-09-23)"}],"latency_note":"not measured yet on a 48 GB card; decode 49 tok/s single-stream for FP8 on an RTX PRO 6000 (measured, benchmark page 17)","in_hosted_demo":false,"receipt_coverage":"strong","receipt_note":null,"hosting":null},{"id":"standard","label":"Standard · one 96 GB Blackwell card, two passes","summary":"The hosted demo: live captions, speaker labels, guidance and lanes during the visit; then the cited note, its self-check (with M17 on CPU), assessment codes and paperwork, all on one Qwen3.8-27B.","components":["asr-live","llm","asr-pass2","detail-checker"],"hardware":"1x RTX PRO 6000 Blackwell 96 GB","quality_evidence":[{"metric":"Medical-term miss / WER, pass 2 MOSS-TD vs the Voxtral live pass (PriMock57, 57 visits)","value":"8.4% / 10.3 vs 8.4% / 13.2","source":"scribe-bench RESULTS.md (MOSS-TD); eval results file asr-voxtral-primock57.json (Voxtral, 2026-09-23)"},{"metric":"DER / word speaker misattribution, MOSS-TD","value":"11.4 / 1.0%","source":"scribe-bench RESULTS.md, wiki decoder-finding"},{"metric":"Composite, two-pass + role map + Qwen3.8-27B minus vanilla (official weights, test 37)","value":"+2.0 [-1.0, +5.7], not resolved; misattributions -0.05","source":"scribe-bench wiki vanilla-vs-best (measured with Sortformer + Parakeet as pass 2)"},{"metric":"Verifier recall on injected errors / flags on clean, Qwen3.8-27B judge","value":"99.1% / 6.5%","source":"scribe-bench wiki verifier"}],"latency_note":"measured on our server: live lanes arrive seconds apart; after audio ends, pass 2 (diarize, role map, cited note, verifier; the verifier is slowest) finishes and the done event arrives about a minute later (decosa-api smoke runs on our server; contract Changes: after-visit pass 2)","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":null,"hosting":null},{"id":"best","label":"Best · adds DeepSeek V4 Flash as the note writer on 2x 96 GB","summary":"Standard stack plus the loop's best writer: higher term precision and, with citations, the most grounded notes; costs two extra 96 GB cards and some plan recall.","components":["asr-live","llm","asr-pass2","writer-alt"],"hardware":"2x RTX PRO 6000 96 GB for the writer (about 83 GB weights per GPU) plus the standard card for the live pass","quality_evidence":[{"metric":"ACI-Bench base ROUGE-L / term precision / plan recall","value":"35.8 / 70.3 / 93 (Qwen3.8-27B ROUGE-L 34.2)","source":"scribe-bench wiki models.md"},{"metric":"PriMock57 cited note, official weights (test 37): grounded / term precision","value":"91.33% / 31.06 (vanilla 88.62% / 26.41)","source":"scribe-bench wiki vanilla-vs-best, citations-and-verifiability"},{"metric":"Composite vs vanilla, official weights","value":"+1.25 [-3.76, +6.23], not resolved; plan recall -8.1 [-13.8, -1.7]","source":"scribe-bench wiki vanilla-vs-best"}],"latency_note":"measured on our server 2026-09-12: about 220 tok/s decode with the DSpark draft, TP2 (wiki live-synthesis); full-stack latency not measured yet","in_hosted_demo":false,"receipt_coverage":"strong","receipt_note":null,"hosting":null}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:0.1.0","purpose":"Lane engine and HTTP/WS API (/ws/live, /demo/replay, /healthz). No GPU. Binds 127.0.0.1 by default."},{"name":"decosa-llm","port":8000,"image":"${DECOSA_REGISTRY}/decosa-llm:0.1.0","purpose":"vLLM OpenAI endpoint for Qwen3.8-27B, served as qwen3.8-27b. Internal to the compose network."},{"name":"decosa-asr","port":8000,"image":"${DECOSA_REGISTRY}/decosa-asr:0.1.0","purpose":"vLLM realtime endpoint for Voxtral Mini 4B Realtime (served as voxtral-realtime). Internal to the compose network."},{"name":"decosa-diarize","port":8092,"image":null,"purpose":"MOSS-Transcribe-Diarize 0.9B pass-2 service (decosa-api services/diarize, GPU0, loopback only). decosa-api runs pass 2 on stop: diarize, role map, cited note, verifier. No published image yet; self-host builds it (see the assemble prompt)."}],"tools":[{"name":"NLM Clinical Tables (ICD-10-CM)","url":"https://clinicaltables.nlm.nih.gov/api/icd10cm/v3/search","license":"Free public API from the US National Library of Medicine","purpose":"Validates every ICD-10-CM code against the official code table; only the condition name is sent."},{"name":"NLM RxNav (RxNorm API)","url":"https://rxnav.nlm.nih.gov/REST","license":"Free public API; NLM says no license is needed for the RxNorm API","purpose":"Maps each drug heard to an RxCUI and ingredient; only the drug name is sent."},{"name":"openFDA drug label","url":"https://api.fda.gov/drug/label.json","license":"Public domain, CC0 1.0 (openFDA terms)","purpose":"Boxed warnings, contraindications, interactions and an example NDC for each drug."},{"name":"Prior-auth rules and CMS Part D flags (bundled)","url":"https://data.cms.gov/provider-summary-by-type-of-service/medicare-part-d-prescribers/monthly-prescription-drug-plan-formulary-and-pharmacy-network-information","license":null,"purpose":"Local JSON: criteria summarised from Medicare LCD/NCD and public payer policies for common orders (e.g. lumbar MRI), plus the share of Part D formularies that require prior auth per drug. No network call."},{"name":"E/M office-visit level table (AMA 2021 MDM)","url":"https://www.ama-assn.org/system/files/2023-e-m-descriptors-guidelines.pdf","license":"AMA copyright; CPT codes are AMA-licensed","purpose":"Built-in table: problems, data and risk, two of three, gives the office-visit E/M code with the evidence for each element."},{"name":"Clinical considerations guideline pack (bundled)","url":"https://decosa.ai/evals/clinical-considerations-sources.json","license":"Excerpts from US federal public-domain pages (CDC, NIH institutes, MedlinePlus, HHS Office on Women's Health) and openFDA drug labels (CC0); topics and checklists written for Decosa (Apache-2.0)","purpose":"26 topics (10 warning-feature, 16 differential) with 39 sources, each quoted verbatim with its URL and access date (26 Sep 2026). The considerations lane cites only these. No USPSTF, NICE or other licensed text. No network call."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB (or B200/GB200)","fits":true,"notes":"Default compose split: LLM 0.60 of the card (~57 GB), ASR 0.25 (~24 GB), about 4 simultaneous visits. Measured on our server with the two models on separate cards (LLM 0.92 of GPU1, ASR 0.40 of GPU0 using 34 GB); the one-card split itself is not yet measured. The after-visit pass (0.9B, BF16 weights 1.8 GB) is meant to share the card after the visit; that fit is not measured."},{"tier":"2x RTX PRO 6000 96 GB, DeepSeek V4 Flash as writer","fits":null,"notes":"DeepSeek V4 Flash ran on our server across both cards (TP2, weights about 83 GB per GPU, 32k context) with little room left, so the live models would need another card. Not measured as a full stack."},{"tier":"1x H100 80 GB / H200 141 GB (Hopper, no NVFP4)","fits":null,"notes":"Not measured. Use Qwen/Qwen3.8-27B-FP8 (Apache-2.0) with LLM_GPU_UTIL=0.62, ASR_GPU_UTIL=0.25 as a starting point."},{"tier":"1x L40S / RTX 6000 Ada 48 GB","fits":null,"notes":"Not measured. FP8 checkpoint, LLM_MAX_LEN=16384, LLM_GPU_UTIL=0.70, ASR_GPU_UTIL=0.22, DECOSA_LIVE_CAP=2 as a starting point."},{"tier":"Cards under 48 GB","fits":false,"notes":"Both models do not fit with useful KV cache (decosa-api self-host guide)."}],"latency":[{"lane":"transcript","typical_ms":1400,"source":"measured on our server 2026-09-23: first transcript event 1.4 s after audio start on both clinical replays (ops/record-all.json)"},{"lane":"note","typical_ms":7354,"source":"measured on our server 2026-09-23: per-replay median 7354 ms (clinical-back-pain, n=10) and 4724 ms (clinical-diabetes-followup, n=9); gateway route, LLM alone on its own GPU, one session"},{"lane":"codes","typical_ms":1449,"source":"measured on our server 2026-09-23: median 1449 ms (clinical-diabetes-followup, n=8); back-pain median was 4 ms because repeat lookups hit the in-memory cache, max 5302 ms"},{"lane":"priorauth","typical_ms":5606,"source":"measured on our server 2026-09-23: per-replay median 5606 ms (clinical-back-pain, n=4) and 2036 ms (clinical-diabetes-followup, n=1)"},{"lane":"visit_level","typical_ms":6819,"source":"measured on our server 2026-09-23: per-replay median 6819 ms (clinical-back-pain, n=2) and 2684 ms (clinical-diabetes-followup, n=2)"},{"lane":"final note after stop, single pass (before pass 2 was added)","typical_ms":24100,"source":"measured on our server 2026-09-23: done event 24.1 s (clinical-back-pain) and 12.1 s (clinical-diabetes-followup) after the audio ended"},{"lane":"note, 4 live sessions in parallel","typical_ms":13698,"source":"measured on our server 2026-09-23: median 13698 ms, max 26594 ms (clinical-back-pain beside three other verticals, ops/smoke-parallel-1.json)"},{"lane":"pass 2 transcript (9-minute visit)","typical_ms":39000,"source":"measured on our server 2026-09-12: MOSS-Transcribe-Diarize, GPU1 (wiki live-synthesis)"},{"lane":"final_transcript (diarize, after audio ends)","typical_ms":9000,"source":"measured, median of 3 clinical runs, range 6.3-11.3 s over 3 clinical runs; decosa-api smoke runs 2026-09-23/24 on our server (ops/record-pass2.json, ops/smoke-final.json; contract Changes: after-visit pass 2)"},{"lane":"role map","typical_ms":2100,"source":"measured, median of 3 clinical runs, range 1.3-3.3 s; decosa-api smoke runs 2026-09-23/24 on our server (ops/record-pass2.json, ops/smoke-final.json; contract Changes: after-visit pass 2)"},{"lane":"final_note (cited)","typical_ms":4100,"source":"measured, median of 3 clinical runs, range 3.3-4.4 s; decosa-api smoke runs 2026-09-23/24 on our server (ops/record-pass2.json, ops/smoke-final.json; contract Changes: after-visit pass 2)"},{"lane":"verifier","typical_ms":30000,"source":"measured, median of 3 clinical runs, range 18.6-51.3 s, one judge call per claim (14-21 claims); decosa-api smoke runs 2026-09-23/24 on our server (ops/record-pass2.json, ops/smoke-final.json; contract Changes: after-visit pass 2)"},{"lane":"considerations (text route, one encounter)","typical_ms":18050,"source":"measured on our server 2026-09-26: median over 26 held-out encounters, 3 in parallel on the shared gateway, max 25.4 s (decosa-api docs/evals/clinical-considerations.md); in the live scribe it runs beside pass 2 after the visit"},{"lane":"all lanes, one-card self-host (direct route)","typical_ms":null,"source":"estimate: not measured; the numbers above used the gateway route with the LLM on its own card"},{"lane":"done after audio ends (note, self-check, codes, paperwork)","typical_ms":60000,"source":"measured on our server 2026-09-29: 3 gateway runs of the 2:52 synthetic visit under shared load, 33.7-71.6 s; the direct route took 20 s (28 Sep)"},{"lane":"turns (speaker labels behind the captions)","typical_ms":15000,"source":"measured 2026-09-28/29: a window every 15 s of new audio, 2.3 s median diarizer time per window on a shared GPU"},{"lane":"guidance","typical_ms":4300,"source":"measured 2026-09-28: the guidance call beside the live view on each tick (every 8-15 s of speech), median model latency on the gateway 4.3 s"}],"benchmark":null,"notes":[]},"buyer_facts":[{"label":"Measured latency","value":"The note, its self-check, the assessment codes and the paperwork drafts arrive about a minute after the audio ends on the shared gateway (measured), sooner on the direct route. Speaker labels follow the captions after a short delay; guidance updates every few seconds of speech."},{"label":"Typical run cost","value":"Several cents per visit for the synthetic sample visit: model calls at the gateway list price plus a little speech-recognition GPU time (measured). The \"in model calls\" figure after a run is the first part only. Cost grows with visit length (the guidance and the live view reread the transcript every tick)."},{"label":"Data retention","value":"Audio and transcript live in server memory for the session and are dropped when it ends; paperwork PDFs are rendered on request and not stored. The receipt store keeps hashes of each model call, not the text."},{"label":"What leaves the box (self-host)","value":"Only the codes lane's lookups: a condition name to NLM Clinical Tables and a drug name to NLM RxNav and openFDA. Block egress for the api container and the lanes report no match; the note is still written."},{"label":"Input","value":"16 kHz mono PCM16 over WebSocket (about 100 ms frames), or a canned sample through /demo/replay. US English visits."},{"label":"Clinical considerations","value":"After the visit: up to 6 items (warning features first, then differentials) from a 26-topic public-domain pack, each with the patient's words, the source excerpt and URL, and what the note does not record. Clinician only; nothing enters the note; each accept or dismiss is sealed in a signed record. About 10 model calls and USD 0.006 for the chest-pain sample transcript (text route, 2026-09-26). Not a diagnosis and not for urgent decisions."},{"label":"Guidance during the visit","value":"Warning-feature considerations, checklist items of a triggered topic not yet asked, and history elements not yet covered (CMS 1995 Documentation Guidelines). Every item shows the transcript words and a public source, in fixed wording, and turns 'covered' when asked. No sound, no alarm colours, no deadlines. A blind clinician judge found 66% of checklist items useful at that moment but only 11% of history items, so history folds away until you ask ('Only when I ask' mode)."},{"label":"Paperwork","value":"Patient instructions (the clinician's own words), a work or school note (dates said in the visit are worked out in code: 'through Friday, October 2nd' becomes 10/02/2026, citing the turn), and the Department of Labor's WH-380-E provider sections. Signatures, signing dates and the employer's section are never filled; nothing is sent. WH-380-E drafts need every box checked: 91.4% of the boxes it filled were right on 8 held-out visits."}],"data_handling":{"page":"/data#clinical","self_host":{"level":"confidential","leaves":"identifiers","summary":"Runs on your machine; by default only short identifiers or a digest go to the public services listed in external_calls."},"hosted":{"level":"operator-processed","demo_only":true,"summary":"Hosted demo on sample or public data only; self-host for real data.","gpus":"operator-contracted","third_parties":[],"retention":"Audio and transcript live in server memory for the session and are dropped when it ends; paperwork PDFs are rendered on request and not stored. The receipt store keeps hashes of each model call, not the text.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[{"to":"NLM Clinical Tables, NLM RxNav and openFDA","route":"both","sends":"identifiers","what":"The codes lane looks up a condition name and drug names (not the transcript or the note).","default":"on","off":"Block egress for the api container: the codes lane reports no match and the note is still written."}]},"console":{"href":"/clinics/visit-copilot","input":"mic","lanes":[{"id":"support","title":"Considerations (live)","kind":"support"},{"id":"note","title":"Visit note (SOAP)","kind":"markdown"},{"id":"codes","title":"Codes (ICD-10-CM, RxNorm)","kind":"codes"},{"id":"priorauth","title":"Prior authorization","kind":"list"},{"id":"visit_level","title":"Visit level (E/M)","kind":"json"},{"id":"considerations","title":"Clinical considerations","kind":"considerations"}],"samples":[{"n":1,"id":"clinical-back-pain","title":"Back pain visit","deep_link":"/clinics/visit-copilot?sample=1&autorun=0"},{"n":2,"id":"clinical-diabetes-followup","title":"Diabetes follow-up","deep_link":"/clinics/visit-copilot?sample=2&autorun=0"},{"n":3,"id":"clinical-chest-pain","title":"Chest burning on the stairs","deep_link":"/clinics/visit-copilot?sample=3&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/clinical-hosted.md","selfhost":"/prompts/clinical-selfhost.md","assemble":"/prompts/clinical-assemble.md","mac":"/prompts/clinical-mac.md"},"rehearsal":{"bundle":"/samples/clinical.zip","bundle_url":"https://decosa.ai/samples/clinical.zip","folder":"/samples/clinical/","expected":"/samples/clinical/expected.json","files":["/samples/clinical/expected.json","/samples/clinical/inputs/back-pain-visit-29s.wav","/samples/clinical/inputs/back-pain-visit-script.json","/samples/clinical/inputs/chest-pain-transcript.txt"],"bytes":676298,"checks":["the session ends with a done event","the session reports no errors","the captions carry the complaint and the MRI order","the final SOAP note names the radiculopathy and the MRI","the note is self-checked before it is final: at least 3 sentences checked and supported by the transcript","the transcript comes back with speaker labels: the clinician","and the patient","the final note is marked self-checked","the paperwork drafts never fill a signature","a validated ICD-10 code for radiculopathy is suggested","a visit level (E/M) is proposed from the problems addressed","the live session also returns the clinical considerations lane with its label","the chest-pain visit shows the cardiac warning feature as a consideration","the cardiac item quotes the patient's words about the jaw","the cardiac item cites a CDC or NHLBI page","everything the lane wrote passes the directive-language check","the output is for the clinician only and is never inserted into the note","the accept decision seals a record that verifies","the record verifier accepts it","the clinical AI monitor's audit finds the output conforming (quotes, citations, wording, labels, signature)","every speech and model call has a signed receipt"],"licence":"Synthetic: a script written for Decosa (no real people, patients or companies) read by Decosa house voices (Kokoro-82M stock voicepacks, Apache-2.0), each allowed by the consent ledger for project decosa-clinical-demo. Part of decosa-api, AGPL-3.0-or-later.","about":"Two cuts from a synthetic primary-care visit (the patient's complaint of eight weeks of back pain shooting down the right leg, and the doctor's assessment: S1 radiculopathy, MRI of the lumbar spine ordered), streamed over the live WebSocket as if from a microphone. The session must produce captions and a SOAP note whose claims are checked against the transcript, with validated ICD-10 codes for lumbar radiculopathy and a visit level. Then, without the speech model, the clinical considerations lane runs on the typed transcript of a second synthetic visit (chest burning after meals, and tightness on the stairs that goes up into the jaw, which the clinician puts down to reflux): it must show the cardiac warning feature with the patient's own words and a cited CDC source, carry the 'not a diagnosis' label with no directive wording, and one accept decision must seal a signed record that verifies and that the clinical AI monitor's audit finds conforming.","run":{"containers":"docker compose exec api python scripts/rehearse.py clinical","checkout":"python scripts/rehearse.py clinical --bundle clinical.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py clinical"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=clinical","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":85.6,"basis":"estimate","unknown":[]},{"id":"best","gpu_gb":277.6,"basis":"estimate","unknown":[]}],"mac":{"fit":"full","memory_gb":48}},"links":{"page":"/clinics/visit-copilot","json":"/use-cases/clinical.json","metrics":"/metrics/clinical","console":"/clinics/visit-copilot","console_sample":"/clinics/visit-copilot?sample=1&autorun=0","stack":"/clinics/visit-copilot#stack","try_live":"/clinics/visit-copilot","watch":"/clinics/visit-copilot","build":"/clinics/visit-copilot#build","self_host":"/clinics/visit-copilot#self-host","prompts":{"hosted":"/prompts/clinical-hosted.md","selfhost":"/prompts/clinical-selfhost.md","assemble":"/prompts/clinical-assemble.md","mac":"/prompts/clinical-mac.md"}}}