Skip to content
decosa

59 · Healthcare · Sales and marketing · live

Medicare sales-call record

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)

  • Per-rule accuracy, test B run 1 (never used to change anything)40/40held outn = 40Run 2: 40/40
  • Per-rule accuracy, test run 4101/101test splitn = 101Run 1 (before the one code change): 101/101
  • Planted typed answers found, test run 49/9test splitn = 9Test B: 2/2; audio: 9/9 (also on the re-voiced audio)
  • Call-sheet items, test run 4 / test B run 179/79 / 31/32test splitn = 111Test run 1, before the needs change: 77/79
  • Benefit claims graded right (good ok, bad flagged), test run 425/25 / 4/4test splitn = 29Test B: 8/8 / 2/2; audio (re-voiced 26 Sep 2026 with Decosa house voices) runs 1 and 2: 18/19 / 2/3, both misses on t9 where the diarizer split a line and the scorer paired claims with the wrong gold lines (first build: 19/19 / 3/3 and 18/19 / 3/3)
  • False alarms on clean calls, test run 4 / test B run 13 of 87 / 0 of 42test splitn = 129All 3 are claims the plan facts do not mention
  • Cited time inside the spoken gold line, audio run 128/29test splitn = 29Audio re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M) and re-run: 21/29 within 3 s of the line's start; run 2: 29/29. First build (macOS voices): 29/29, 16/29 within 3 s

Dataset

17 synthetic scripted Medicare sales calls with fictional agency, plans and people, written from 42 CFR 422/423 subpart V: 3 dev, 10 test, 4 test B, plus 8 of them as TTS audio (Decosa house voices, Kokoro-82M, each allowed by the consent ledger) run through the diarizer.

Caveats

  • The same author wrote the scripts, the labels, the prompts and the code; the labels are one reading of the rule text.
  • 17 short scripted calls with clean TTS audio; real sales calls (long, accents, transfers, Spanish) are not measured.
  • One code change (the premiums needs topic) was made after test run 1, so test is not fully held out for that item; test B was never used to change anything. The scorer's audio matching was refined after the audio runs, before this write-up.
  • Benefit claims are judged against the plan facts given, not the plan's filed benefits.
  • Answer probabilities and claim confidences come from tables fit on other domains and are not calibrated for sales calls.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
26 Sep 2026
Latency, this run
n/a
p50 over passed runs
7.5 s
Receipts
15
Model calls
n/a
Tokens
n/a
Cost per run
$0.006

Self-host verification

Verified on 26 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

The assembly prompt's smoke tests ran against the already-running local Qwen3.8-27B vLLM (127.0.0.1:8114, network_mode host instead of the compose llm service): tb1-showcase flagged the disclaimer timing, "free", the final-expense pitch and the $3,000 dental claim (contradicted), 15 attested receipts, record verified, retention 2029-11-05 / 2032-11-05, 3.8 s; the pasted cold call flagged Medicare, free, unsolicited, no SOA and the dental claim; no consent gave 400.

Rehearsal bundle: medicare-call-record.zip (4 KB, 12 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.
  • Measured on 17 short synthetic role-plays written by the building agent, with clean TTS audio. Not measured on real sales calls (20-60 minutes, accents, transfers, Spanish) or with an independent reviewer's labels.
  • Claims the plan facts do not mention come back flagged even when true (2-3 on a clean drug-plan call); give the full Summary of Benefits.
  • ASR errors become claim errors: "eyewear" heard as "in-store" was flagged in one audio run.
  • Audio intake (/medicare/transcribe) is self-host only. The hosted demo's audio samples were transcribed on our server and are bundled.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Two extractions (products, first benefit discussion, enrollment steps; the agent's benefit claims), one typed yes/no/unclear check per question, and one grounding judgment per benefit claim against the plan factsQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
  • Recording to a timed, speaker-labelled transcript (POST /medicare/transcribe, self-host; the demo's audio samples were transcribed with it)MOSS-Transcribe-Diarize 0.9BApache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 48 GB card (1)
  • typed checks correct / claims graded right: not measured yet
Standard · the hosted demo, one 96 GB card (4)
  • held-out test B, 4 synthetic calls: typed checks / call-sheet items / claims graded right / false alarms on the 2 clean calls: 40/40 / 31/32 / 10/10 / 0 of 42decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route, run 1 (run 2 the same); prompts frozen on a 3-call dev split; data, labels and prompts written by the building agent
  • test set, 10 calls: typed checks / planted typed answers / call-sheet items / good claims ok / bad claims flagged: 101/101 / 9/9 / 79/79 / 25/25 / 4/4decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route, run 4, after one code change made on test run 1 (77/79 call-sheet items before it)
  • false alarms on the 4 clean test calls (flag or review): 3 of 87 (claims the plan facts do not mention)decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route, run 4
  • on diarized TTS audio, 8 calls: typed checks / call-sheet items / claims graded right / cited time inside the spoken line: 81/81 / 63/63 / 20/22 / 28/29 (run 2: 20/22 claims, 29/29)decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26 on audio re-voiced 26 Sep 2026 from macOS voices to Decosa house voices (Kokoro-82M) (the first build: 22/22 and 21/22 claims, 29/29), gateway route, asr runs 1 and 2
Best · two 96 GB cards (1)
  • typed checks correct / claims graded right: not measured yet
Wanted · two large judges from different families (1)
  • typed checks correct / claims graded right: not measured yet

How we measure · All tools