59 · Healthcare · Sales and marketing · live
Medicare sales-call record
Eval results
Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)
- Per-rule accuracy, test B run 1 (never used to change anything)40/40held outn = 40Run 2: 40/40
- Per-rule accuracy, test run 4101/101test splitn = 101Run 1 (before the one code change): 101/101
- Planted typed answers found, test run 49/9test splitn = 9Test B: 2/2; audio: 9/9 (also on the re-voiced audio)
- Call-sheet items, test run 4 / test B run 179/79 / 31/32test splitn = 111Test run 1, before the needs change: 77/79
- Benefit claims graded right (good ok, bad flagged), test run 425/25 / 4/4test splitn = 29Test B: 8/8 / 2/2; audio (re-voiced 26 Sep 2026 with Decosa house voices) runs 1 and 2: 18/19 / 2/3, both misses on t9 where the diarizer split a line and the scorer paired claims with the wrong gold lines (first build: 19/19 / 3/3 and 18/19 / 3/3)
- False alarms on clean calls, test run 4 / test B run 13 of 87 / 0 of 42test splitn = 129All 3 are claims the plan facts do not mention
- Cited time inside the spoken gold line, audio run 128/29test splitn = 29Audio re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M) and re-run: 21/29 within 3 s of the line's start; run 2: 29/29. First build (macOS voices): 29/29, 16/29 within 3 s
Dataset
17 synthetic scripted Medicare sales calls with fictional agency, plans and people, written from 42 CFR 422/423 subpart V: 3 dev, 10 test, 4 test B, plus 8 of them as TTS audio (Decosa house voices, Kokoro-82M, each allowed by the consent ledger) run through the diarizer.
Caveats
- The same author wrote the scripts, the labels, the prompts and the code; the labels are one reading of the rule text.
- 17 short scripted calls with clean TTS audio; real sales calls (long, accents, transfers, Spanish) are not measured.
- One code change (the premiums needs topic) was made after test run 1, so test is not fully held out for that item; test B was never used to change anything. The scorer's audio matching was refined after the audio runs, before this write-up.
- Benefit claims are judged against the plan facts given, not the plan's filed benefits.
- Answer probabilities and claim confidences come from tables fit on other domains and are not calibrated for sales calls.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 26 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 7.5 s
- Receipts
- 15
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.006
Self-host verification
Verified on 26 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after
The assembly prompt's smoke tests ran against the already-running local Qwen3.8-27B vLLM (127.0.0.1:8114, network_mode host instead of the compose llm service): tb1-showcase flagged the disclaimer timing, "free", the final-expense pitch and the $3,000 dental claim (contradicted), 15 attested receipts, record verified, retention 2029-11-05 / 2032-11-05, 3.8 s; the pasted cold call flagged Medicare, free, unsolicited, no SOA and the dental claim; no consent gave 400.
Rehearsal bundle: medicare-call-record.zip (4 KB, 12 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.
- Measured on 17 short synthetic role-plays written by the building agent, with clean TTS audio. Not measured on real sales calls (20-60 minutes, accents, transfers, Spanish) or with an independent reviewer's labels.
- Claims the plan facts do not mention come back flagged even when true (2-3 on a clean drug-plan call); give the full Summary of Benefits.
- ASR errors become claim errors: "eyewear" heard as "in-store" was flagged in one audio run.
- Audio intake (/medicare/transcribe) is self-host only. The hosted demo's audio samples were transcribed on our server and are bundled.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Two extractions (products, first benefit discussion, enrollment steps; the agent's benefit claims), one typed yes/no/unclear check per question, and one grounding judgment per benefit claim against the plan factsQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
- Recording to a timed, speaker-labelled transcript (POST /medicare/transcribe, self-host; the demo's audio samples were transcribed with it)MOSS-Transcribe-Diarize 0.9BApache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 48 GB card (1)
- typed checks correct / claims graded right: not measured yet
Standard · the hosted demo, one 96 GB card (4)
- held-out test B, 4 synthetic calls: typed checks / call-sheet items / claims graded right / false alarms on the 2 clean calls: 40/40 / 31/32 / 10/10 / 0 of 42decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route, run 1 (run 2 the same); prompts frozen on a 3-call dev split; data, labels and prompts written by the building agent
- test set, 10 calls: typed checks / planted typed answers / call-sheet items / good claims ok / bad claims flagged: 101/101 / 9/9 / 79/79 / 25/25 / 4/4decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route, run 4, after one code change made on test run 1 (77/79 call-sheet items before it)
- false alarms on the 4 clean test calls (flag or review): 3 of 87 (claims the plan facts do not mention)decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route, run 4
- on diarized TTS audio, 8 calls: typed checks / call-sheet items / claims graded right / cited time inside the spoken line: 81/81 / 63/63 / 20/22 / 28/29 (run 2: 20/22 claims, 29/29)decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26 on audio re-voiced 26 Sep 2026 from macOS voices to Decosa house voices (Kokoro-82M) (the first build: 22/22 and 21/22 claims, 29/29), gateway route, asr runs 1 and 2
Best · two 96 GB cards (1)
- typed checks correct / claims graded right: not measured yet
Wanted · two large judges from different families (1)
- typed checks correct / claims graded right: not measured yet