{"schema_version":"1","site":"https://decosa.ai","id":"clinical-ai-monitor","num":"29","name":"Clinical AI assurance monitor","tool_name":"Check an AI note","short":"AI monitor","blurb":"Checks any commercial AI scribe's note against the visit transcript: every sentence typed as not in the visit, a changed detail, the wrong speaker or an invented exam finding, with the transcript line, plus the medicines, allergies and plan items the note left out. Tightens the note without adding a word. Buyer mode scores vendors into a signed scorecard.","status":"live","labels":{"industry":["healthcare","compliance-trust"],"job":["review","attest"],"input":["text","voice"],"deploy":["selfhost"],"demo":"synthetic","status":"live","output":["data","record"],"data":["phi"],"hardware":"gpu-96","licence":"permissive"},"industries":["healthcare","compliance-trust"],"runs_in":["selfhost","crp-phi-research"],"part_of":[],"built_from":["grounding","diarize","signed-record"],"models":"Qwen3.8-27B · MOSS diarizer (audio)","where":"Self-host (patient data stays on site)","hardware":"1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the judge; the diarizer (optional, for audio) needs a second GPU slot","final_artifact":"A signed per-encounter report and a signed summary of rates per vendor, specialty and week.","self_host_first":true,"verification":{"receipt_coverage":"full","summary":"Receipt per verdict; signed reports and summary","manual_qa":{"hosted":{"date":"2026-09-28","result":"pass","p50_ms":28000,"p95_ms":41400,"runs":null,"receipts_per_run":30,"cost_per_run_usd":0.021},"selfhost":{"date":"2026-09-25","result":"pass","method":"assemble-prompt.md on our server: fresh clone into a clean directory, api image built from docker/api/Dockerfile, compose with a named volume, pointed at the already-running Qwen3.8-27B vLLM (127.0.0.1:8114) and MOSS diarizer (127.0.0.1:8092) instead of starting new ones; then torn down","notes":"Images build, the service starts, and the smoke steps pass: the wrong-speaker sample flagged (17 s), the report verifies with transcript and note hashes, a changed count fails verification, the summary signs, and 170 s of audio came back as 22 speaker turns. Model-server startup itself was not re-verified. /healthz says ok: false on this stack because it also checks the live-scribe recogniser, which the monitor does not use."},"known_limits":["Measured on synthetic notes with one clear planted error each; real scribe errors are subtler.","Changed details: 94% caught as errors with the detail checker (47/50 and 188/200 blind plants), 72-74% by the judge alone; the checker only upgrades the judge's own 'detail not in the visit' flags.","On real speech-recognition transcripts the detail checker adds some false error flags (4 in 1,469 ACI-Bench sentences, e.g. 'type i' vs 'type 1'); each comes with its transcript line.","Wrong-speaker errors are caught but typed correctly only 62% of the time.","Tighten shortened the 50 test notes by a median 15%; 5 of 362 recorded key items came back 'partly recorded' after tightening (none missing).","Measures against the transcript it is given; speech-recognition errors pass into the measure.","Clinicians' own notes get flagged for things never said aloud; use it on AI drafts.","Audio input is API-only (POST /monitor/transcribe); the page takes text."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Invented fact caught as an error","value":"49 (98%)","unit":null,"n":50,"split":"test","note":null},{"name":"Changed detail caught as an error","value":"47 (94%)","unit":null,"n":50,"split":"test","note":"With the detail checker (M17). The judge alone: 36 (72%). False error flags on faithful sentences unchanged: 4 of 1,135"},{"name":"Key item left out caught","value":"45 (90%)","unit":null,"n":50,"split":"test","note":"48/50 after a guard was removed post-test; that number is not held out"},{"name":"Wrong speaker caught / typed correctly","value":"47 (94%) / 31 (62%)","unit":null,"n":50,"split":"test","note":null},{"name":"Invented exam finding caught as an error","value":"49 (98%)","unit":null,"n":50,"split":"test","note":null},{"name":"False error flags on faithful notes (sentences)","value":"4 (0.35%)","unit":null,"n":1135,"split":"test","note":"Key items called left out: 3 of 368 (0.8%); notes with at least one error finding: 6 of 50 (12%)"},{"name":"Tighten: words removed (median per note)","value":"15%","unit":null,"n":50,"split":"test","note":"Range 0-31%; 15,835 to 13,499 words over 50 clean notes; 0 words added (checked in code on every output)"},{"name":"Tighten: key items that stopped being fully recorded","value":"5 of 362 (1.4%)","unit":null,"n":362,"split":"test","note":"All 5 came back 'partly recorded', none missing; e.g. 'over the counter' dropped from a medicine"},{"name":"Changed detail caught as an error, 200 more blind plants","value":"188 (94.0%)","unit":null,"n":200,"split":"test","note":"95% CI 89.8-96.5%; judge alone 147 (73.5%); written by a blind sub-agent on the 50 held-out visits, never trained on"},{"name":"Detail checker alone on CPU: changed details caught / faithful sentences flagged","value":"33 of 50 (66%) / 3 of 1,135 (0.26%)","unit":null,"n":50,"split":"test","note":"No language model, windows by word overlap, p(changed) >= 0.9; median 7.1 s per note on 8 CPU threads"}],"dataset":"All 57 PriMock57 mock primary-care consultations (CC BY 4.0) with reference transcripts; one synthetic cited scribe note per visit plus five copies each with one planted error; 7 visits dev, 50 held out.","held_out":true,"caveats":["The notes are synthetic and the plants are clean single errors; real scribe errors are subtler, so these detection rates are an upper bound.","Reference transcripts: no speech-recognition error enters this eval.","One judge model family; a second judge is not run. Plants were checked by a validator, not reviewed by a clinician.","Primary care, English, remote visits, 50 test visits; no specialty or inpatient data and no clinician agreement study.","The detail checker's thresholds were fixed on the dev split before the test run. On ACI-Bench's human notes (real speech-recognition transcripts) it turned 9 of 1,469 sentences into errors: 2 real conflicts, 3 details never said aloud, 4 not errors. On ACI-Bench plants the judge alone caught 71/78 and with the checker 72/78.","A blind frontier model (Claude Opus via Claude Code) proofreading 40 of the same notes caught 20/20 changed details with 1 flag on 20 faithful notes; this tool caught 19/20 with 2 flags, on open weights you can run yourself."],"date":"2026-09-28","doc_url":"https://decosa.ai/metrics/evals/clinical-ai-monitor"},"stack":{"summary":"Paste the note a commercial AI scribe wrote and the visit transcript (text, captions, a scribe export, or segments; audio goes through the API's diarizer). An open judge model checks every note sentence against the numbered, speaker-labelled transcript and types each problem: not in the visit, contradicts the visit (a changed dose, duration, side or number), the wrong speaker (the doctor's guess or a relative's history written as the patient's), or an exam finding nobody made. Each flag carries the note sentence and the transcript lines, or 'not in the transcript'. The medicines, allergy status, plan items and impression are read from the transcript and checked against the note for omissions. Tighten gives a shorter note that only deletes and merges, enforced in code, with the flagged sentences kept word for word, marked and listed for the clinician to fix. Buyer mode checks many visits from one or more scribe vendors and returns a signed scorecard by error type next to the checker's own measured catch and false-flag rates. Every visit ends in a signed report of hashes, line numbers and counts, with no text. It checks documentation against the conversation; it does not judge care.","tagline":"Checks any AI scribe's note against the visit transcript, sentence by sentence: invented facts, changed doses and details, exam findings nobody made, the wrong speaker and what was left out, each with the transcript line. Tightens the note without adding anything, and scores vendors in bulk.","deployment":"self-host-first","regulatory_note":"A documentation-quality measure for the health system's own oversight, not a medical device function and not software that recommends care: it does not analyse images or signals, does not diagnose, and makes no recommendation about any patient's care; it compares a note with the conversation it came from. The closest statutory exclusion is software for administrative support of a health care facility (FD&C Act 520(o)(1)(A)); FDA's clinical decision support guidance (reissued 29 Jan 2026) covers software that recommends care, which this does not. Keep it that way: do not wire its flags into real-time changes to a patient's care. HIPAA: run it on your own hardware; measuring your own notes is a quality assessment activity, part of health care operations (45 CFR 164.501). A vendor that hosts it or can reach the PHI needs a business associate agreement. The hosted demo takes only PriMock57 mock consultations and synthetic notes. Governance: the Joint Commission and CHAI guidance on the Responsible Use of AI in Healthcare (17 Sep 2025) names clinical documentation in scope and asks for local validation, ongoing quality monitoring and error reporting to leadership and vendors; the Joint Commission's voluntary certification on it was announced 1 Jun 2026. ONC HTI-1's decision support interventions criterion (45 CFR 170.315(b)(11)) binds certified EHR developers, not hospitals, and HTI-5, proposed 29 Dec 2025 and not final when checked, would remove its source-attribute and risk-management duties; neither changes what a hospital may measure itself. State laws bind the providers that use scribes, not this monitor: California AB 3030 (disclaimers on AI-generated patient communications, from 1 Jan 2025), Texas SB 1188 (practitioners review AI-created records, from 1 Sep 2025) and TRAIGA (disclose AI use in care, from 1 Jan 2026), and Illinois Public Act 104-0054 (consent before recording therapy sessions for AI, 1 Aug 2025). Its reports can document the review these laws expect. Some scribe contracts restrict benchmarking or publishing results; check yours. Not legal advice. Model licences: Apache-2.0 (Qwen3.8-27B, MOSS-Transcribe-Diarize). Checked 25 Sep 2026.","components":[{"id":"monitor","role":"Monitor: transcript and note parsing, sentence and section offsets, finding types, the signed report and the summary with intervals and drift (no model; CPU)","name":"decosa-api clinical AI assurance monitor (decosa_api/verticals/monitor)","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12; the grounding module's sentence splitter; Ed25519 signing (cryptography)","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard","wanted"],"alternative_to":null},{"id":"judge","role":"Judge: one call per note sentence, one checklist extraction, one coverage check; also the speaker role map for anonymous labels","name":"Qwen3.8-27B (NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP NVFP4, GDN/attention FP8) + FP8 KV cache; MTP head, 3 draft tokens","vram_gb":20,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, temperature 0, thinking off, prefix caching","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["lite","standard"],"alternative_to":null},{"id":"detail-checker","role":"Detail checker (M17): after the judge, each drug, dose, frequency, route, date, duration, side and number in a sentence is read against its transcript lines and labelled same / changed / absent; a judge 'detail not in the visit' flag becomes a changed-detail error when p(changed) >= 0.5","name":"decosa-note-detail-checker-modernbert-large (M17, our own model; Apache-2.0 on Hugging Face)","hf_repo":"decosaai/decosa-note-detail-checker-modernbert-large","license":"Apache-2.0","params":"395M","quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"PyTorch CPU, 8 threads (services/detail_checker, 127.0.0.1:8495)","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard","wanted"],"alternative_to":null},{"id":"diarizer","role":"Audio in (optional): one-pass speaker-attributed transcript of the whole visit","name":"MOSS-Transcribe-Diarize 0.9B","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize","license":"Apache-2.0","params":"0.9B","quant":"BF16","vram_gb":null,"memory_gb_estimate":null,"engine":"transformers (trust_remote_code) + moss_transcribe_diarize package; hosted: decosa-diarize service, 127.0.0.1:8092","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["standard","wanted"],"alternative_to":null},{"id":"judge-wanted","role":"Stronger judge (wanted)","name":"DeepSeek V4 Flash","hf_repo":"deepseek-ai/DeepSeek-V4-Flash","license":"MIT","params":"284B","quant":"FP8 + FP4 experts (official mixed precision)","vram_gb":166,"memory_gb_estimate":null,"engine":"vLLM, tensor parallel 2","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null},{"id":"super-judge","role":"One-card second judge from another model family","name":"Nemotron-3-Super-120B-A12B (NVFP4)","hf_repo":"nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4","license":"NVIDIA Nemotron Open Model License","params":"120B","quant":"NVFP4 (80.3 GB of weights)","vram_gb":null,"memory_gb_estimate":80,"engine":"vLLM (sm_120 support unconfirmed: smoke-test first)","receipt_coverage":"none","in_hosted_demo":false,"tiers":["alternates"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · text transcripts, one 32 GB card","summary":"The judge only: send transcripts as text (most scribe products export one). No audio input.","components":["monitor","judge","detail-checker"],"hardware":"1x RTX 5090 32 GB (estimate)","quality_evidence":[{"metric":"Same judge and prompts as standard, so the same detection and false-flag rates","value":"see standard","source":"docs/evals/clinical-ai-monitor.md"}],"latency_note":"estimate: as standard on a smaller card; not timed separately.","in_hosted_demo":false,"receipt_coverage":"strong","receipt_note":"Self-hosted calls are attested with the box's key; on a network provider they carry gateway receipts.","hosting":null},{"id":"standard","label":"Standard · judge plus diarizer (hosted demo)","summary":"Adds audio input through the MOSS diarizer. This is what the hosted API runs.","components":["monitor","judge","diarizer","detail-checker"],"hardware":"1x RTX PRO 6000 96 GB (measured), diarizer on a second slot","quality_evidence":[{"metric":"Planted errors caught at error severity (50 held-out PriMock57 visits, one error per note): invented fact / invented exam finding / wrong speaker / key item left out / changed detail","value":"49/50 / 49/50 / 47/50 / 45/50 / changed detail 47/50 with the detail checker (36/50 judge alone); every plant flagged at least as review (50/50 each)","source":"docs/evals/clinical-ai-monitor.md"},{"metric":"Changed details, 200 more blind plants on the same 50 held-out visits (8 detail types)","value":"188/200 (94%) with the detail checker, 147/200 (73.5%) judge alone; false error flags on faithful notes unchanged (4/1,135)","source":"docs/evals/clinical-ai-monitor.md (M17)"},{"metric":"Wrong-speaker errors typed as wrong speaker","value":"31/50 (62%); the rest flagged as contradicts or not in the visit","source":"docs/evals/clinical-ai-monitor.md"},{"metric":"Error flags on faithful notes (1,135 sentences, 368 key items)","value":"4 sentences (0.35%; 1 real error in the note, 2 judge mistakes, 1 debatable), 3 key items (0.8%; 1 real)","source":"docs/evals/clinical-ai-monitor.md"},{"metric":"Clinicians' own PriMock57 notes: lines with an error finding","value":"12.1%; in a sample of 20, 16 were statements the transcript does not support (names, routine negatives never asked)","source":"docs/evals/clinical-ai-monitor.md"},{"metric":"Speech recognition, if you send audio (MOSS-Transcribe-Diarize, PriMock57)","value":"10.3% WER","source":"scribe-bench RESULTS.md (the clinical scribe's pass 2)"}],"latency_note":"measured on our server: seconds per note straight to the model server; longer through our gateway, and a few minutes for the recorded demo runs when the shared gateway was loaded.","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Every judge call is a separate gateway call with a gateway-signed receipt; the report lists them.","hosting":null},{"id":"wanted","label":"Wanted · a second judge from another family","summary":"DeepSeek-V4-Flash re-judges the transcripts to check the first judge's rates, where wrong-speaker and changed-detail errors are hardest. Patient data stays on your hardware, never on community providers. Not served yet.","components":["monitor","judge","judge-wanted","diarizer","detail-checker"],"hardware":"Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash beside the standard card, or a Mac with 192 GB or more (MXFP4 MLX build, 156 GB on our Mac; not measured). Estimate.","quality_evidence":[{"metric":"This eval, same protocol","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"On your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.","hosting":"own-hardware"}],"alternates":[{"id":"one-card-judge","label":"A second judge on one card","components":["super-judge"],"hardware":"1x RTX PRO 6000 96 GB (80 GB of weights)","use":"A cheaper cross-family check that fits beside the 27B: smaller than DeepSeek-V4-Flash, so we would host it ourselves. sm_120 support unconfirmed.","status":"not served"}],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"GET /monitor/info, /monitor/samples, /monitor/demo-dashboard; POST /monitor/check (SSE or JSON), /monitor/transcribe (audio), /monitor/verify, /monitor/summary. Keeps nothing: transcripts, notes and audio stay in memory for the request."},{"name":"vLLM (judge)","port":8114,"image":"vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1","purpose":"Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host)."},{"name":"decosa-diarize (optional)","port":8092,"image":"built from services/diarize in decosa-api","purpose":"MOSS-Transcribe-Diarize for audio input (POST /v1/diarize with PCM16 16 kHz mono)."}],"tools":[{"name":"PriMock57 (Papadopoulos Korfiatis et al., 2022)","url":"https://github.com/babylonhealth/primock57","license":"CC BY 4.0","purpose":"57 mock primary-care consultations, role-played by clinicians and actors, with per-speaker transcripts and the clinicians' own notes: the eval set and the hosted demo's visits."},{"name":"scripts/monitor_eval.py and docs/evals/clinical-ai-monitor.md","url":null,"license":"Apache-2.0","purpose":"The eval: synthetic scribe notes for all 57 visits with five planted error types, false flags on the clean notes, and the flag rate on the clinicians' own notes."}],"hardware":[{"tier":"1x RTX 5090 32 GB","fits":true,"notes":"Text transcripts only: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for prompts up to about 16k tokens (a 25-minute visit). Estimate: same judge as the grounding check; not run here for this tool."},{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured on our server: the hosted demo and the eval ran on this card, shared with other services; the diarizer ran on a second card."}],"latency":[{"lane":"one note with tighten, hosted demo through the gateway (18-30 model calls)","typical_ms":28000,"source":"measured on our server 2026-09-28: median of 9 runs of the three note samples (10-41 s), gateway shared with other work; the tighten proposal runs alongside the check"},{"lane":"the check alone, same runs","typical_ms":23000,"source":"measured on our server 2026-09-28: median of the same 9 runs' report time (10-36 s)"},{"lane":"one note (about 23 checked sentences) straight to the model server, 4 calls in flight","typical_ms":12600,"source":"measured on our server 2026-09-25: median over 50 clean notes in the eval"},{"lane":"audio to transcript (MOSS diarizer), 170 s of audio","typical_ms":10300,"source":"measured on our server 2026-09-25, one PriMock57 clip"}],"benchmark":null,"notes":["Findings are typed: not in the visit, contradicts the visit, wrong speaker, invented exam finding, and left out are errors; a detail not in the visit and a partly recorded item are worth a look. A sentence that could not be judged is counted as not judged, never guessed.","A changed detail the judge only calls 'a detail not in the visit' is settled by the detail checker (our M17 model, CPU): if the note's value conflicts with the transcript line, the flag becomes an error. It never removes or lowers a judge flag, and if the checker is down the judge's verdicts stand.","Tighten only removes or condenses. Code checks every proposed sentence: its words must appear in the source sentences in the same order (a few joining words aside), a dropped negation takes the rest of its sentence, a dropped hedge or condition its whole sentence, and every number, dose, frequency, side and medicine name stays. What fails goes back to the model once, then falls back to the original words.","The signed report holds hashes, offsets, transcript line numbers, verdicts and counts. It holds no note or transcript text, the visit week instead of the date, and the encounter reference only as a SHA-256, so it can go in an ordinary log or to a committee.","It measures the note against the transcript, not against the patient: errors in the transcript (speech recognition) pass into the measure. The eval used reference transcripts.","Clinicians' own notes record things never said aloud (brand names, standard negatives, their own examination), and the checker flags those too. It is a check for AI drafts, not a style check for people."]},"buyer_facts":[{"label":"What leaves the box","value":"Nothing when self-hosted (DECOSA_LLM_ROUTE=direct). The hosted demo takes public mock visits only."},{"label":"Retention","value":"None: transcripts, notes and audio stay in memory for one request; logs carry counts."},{"label":"What you keep","value":"A signed report per visit (hashes, line numbers, counts; no text, the ISO week not the date) and signed summaries."},{"label":"Inputs","value":"Transcript as text with speaker labels, WebVTT or SRT captions, a scribe export with timestamps, or segments (up to 40,000 characters), or WAV audio up to 15 minutes through the API; the note as text (up to 12,000 characters). Buyer mode: up to 25 visits per call hosted, no limit self-hosted."},{"label":"Typical cost","value":"A few cents or less per note at the gateway list price (measured on the demo notes), rising with note and transcript length, including tighten; the detail checker runs on CPU at no token cost. Self-hosted: your own hardware."},{"label":"Hardware","value":"One GPU with 32 GB or more for the judge; the diarizer (audio only) needs about 4 GB more; the detail checker runs on CPU (8 threads, about 2.5 GB of RAM)."},{"label":"Output","value":"Per sentence: the flag type, the note sentence and the transcript lines (or 'not in the transcript'); key items left out; a tightened note; a signed per-visit report; in buyer mode, a signed scorecard per vendor and version."}],"hosted_now":{"needs":["qwen3.8-27b"],"off":[],"live_by_default":true,"live_status":"https://api.decosa.ai/status"},"data_handling":{"page":"/data#clinical-ai-monitor","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":true,"summary":"Hosted demo on sample or public data only; self-host for real data.","gpus":"operator-contracted","third_parties":[],"retention":"None: transcripts, notes and audio stay in memory for one request; logs carry counts.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/clinics/clinical-ai-monitor","input":"monitor","lanes":[{"id":"sentences","title":"Note sentences","kind":"list"},{"id":"checklist","title":"Key items","kind":"list"},{"id":"report","title":"Signed report","kind":"json"},{"id":"dashboard","title":"Committee dashboard","kind":"json"}],"samples":[{"n":1,"id":"video-errors","title":"Video errors","deep_link":"/clinics/clinical-ai-monitor?sample=1&autorun=0"},{"n":2,"id":"video-clean","title":"Video clean","deep_link":"/clinics/clinical-ai-monitor?sample=2&autorun=0"},{"n":3,"id":"gp-changed","title":"Gp changed","deep_link":"/clinics/clinical-ai-monitor?sample=3&autorun=0"},{"n":4,"id":"gp-frequency","title":"Gp frequency","deep_link":"/clinics/clinical-ai-monitor?sample=4&autorun=0"},{"n":5,"id":"scorecard","title":"Scorecard","deep_link":"/clinics/clinical-ai-monitor?sample=5&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/clinical-ai-monitor-hosted.md","selfhost":"/prompts/clinical-ai-monitor-selfhost.md","assemble":"/prompts/clinical-ai-monitor-assemble.md","mac":"/prompts/clinical-ai-monitor-mac.md"},"rehearsal":{"bundle":"/samples/clinical-ai-monitor.zip","bundle_url":"https://decosa.ai/samples/clinical-ai-monitor.zip","folder":"/samples/clinical-ai-monitor/","expected":"/samples/clinical-ai-monitor/expected.json","files":["/samples/clinical-ai-monitor/expected.json","/samples/clinical-ai-monitor/inputs/ATTRIBUTION.txt","/samples/clinical-ai-monitor/inputs/checklist.json","/samples/clinical-ai-monitor/inputs/note.txt","/samples/clinical-ai-monitor/inputs/transcript.txt"],"bytes":8074,"checks":["all 4 clinical sentences were checked","every checked sentence got a verdict (none not judged or errored)","the planted exam finding (oxygen saturation 97%) is not supported","the coverage check answered for both checklist items","no checklist item errored","the signed report verifies against this server's key","the report matches the transcript it was made from","the report matches the note it was made from","a report with the exam verdict changed no longer verifies","every model call has a signed receipt"],"licence":"Transcript: PriMock57 mock primary-care consultation day1_consultation07 (Papadopoulos Korfiatis et al., 2022, github.com/babylonhealth/primock57), CC BY 4.0; role-played by clinicians and actors, no patients. The note is synthetic, written for this demo with one planted error. See inputs/ATTRIBUTION.txt.","about":"A four-sentence scribe note checked sentence by sentence against the transcript of a mock respiratory consultation (PriMock57, role-played, no patients), plus a two-item coverage checklist. The note carries one planted invented exam finding (oxygen saturation and respiratory rate from a phone consultation) that must not be supported, and the signed report must verify against the same transcript and note and fail once one verdict is changed.","run":{"containers":"docker compose exec api python scripts/rehearse.py clinical-ai-monitor","checkout":"python scripts/rehearse.py clinical-ai-monitor --bundle clinical-ai-monitor.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py clinical-ai-monitor"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=clinical-ai-monitor","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":61.6,"basis":"estimate","unknown":[]},{"id":"wanted","gpu_gb":253.6,"basis":"estimate","unknown":[]},{"id":"alternate-one-card-judge","gpu_gb":81.6,"basis":"estimate","unknown":[]}],"mac":{"fit":"full","memory_gb":48}},"links":{"page":"/clinics/clinical-ai-monitor","json":"/use-cases/clinical-ai-monitor.json","metrics":"/metrics/clinical-ai-monitor","console":"/clinics/clinical-ai-monitor","console_sample":"/clinics/clinical-ai-monitor?sample=1&autorun=0","stack":"/clinics/clinical-ai-monitor#stack","try_live":"/clinics/clinical-ai-monitor","watch":"/clinics/clinical-ai-monitor","build":"/clinics/clinical-ai-monitor#build","self_host":"/clinics/clinical-ai-monitor#self-host","prompts":{"hosted":"/prompts/clinical-ai-monitor-hosted.md","selfhost":"/prompts/clinical-ai-monitor-selfhost.md","assemble":"/prompts/clinical-ai-monitor-assemble.md","mac":"/prompts/clinical-ai-monitor-mac.md"}}}