Eval: Medicare sales-call record (59)
Run on 26 Sep 2026 against the pre-release server (127.0.0.1:8459; the audio set again on 127.0.0.1:8479 after the audio was re-voiced) through the model gateway, with Qwen3.8-27B at temperature 0. Every model call had a signed gateway receipt. The gateway was shared with other workloads' runs (and was down for about 25 minutes in the middle; runs from that window are not counted), so latency is as measured under load.
What is measured
Each call goes through POST /medicare/check with its call sheet, the plan facts (a fictional Summary of Benefits) and a
consent statement. Expected answers come from tags on the script lines (scripts/medicare_eval/convs.py), not from a
model. The expected call-sheet items were written by hand from the rule text before the code ran.
- Per-rule accuracy. The typed answer (yes, no, unclear or n/a) equals the expected answer. Two answers accept more than one value: a vague "I guess, if you think it's best" / "Well, I suppose, if I have to" accepts no or unclear for the clear-yes check; and "supplement presented?" accepts n/a or no on a drug-plan-only call.
- Planted recall. Of the expected typed answers that mark a problem ("free" used, claims to be Medicare, no enrollment notice, and so on), how many came back that way.
- Call-sheet items. Disclaimer timing and numbers, SOA on file, products outside the SOA, non-health products, cold calls, needs topics and the pre-enrollment checklist: the status must equal the hand label.
- Benefit-claim grounding. Every benefit or cost the agent states in a script is tagged good (the plan facts support
it) or bad (the plan facts contradict it or lack it). A good claim must come back
ok(supported), a bad oneflag(unsupported or contradicted). Claims the extraction adds beyond the tagged ones are listed; on clean calls a flagged one counts as a false alarm. - False alarms. On the clean calls, every typed check, call-sheet item and claim that came back
flagorreview. - Timestamp accuracy. For each correct "yes", and for the first discussion of benefits: on text, the cited line is a gold line; on audio, the cited time falls inside a gold line as spoken (from 3 s before its start to the next line's start). The diarizer splits a long script line into several segments and the quote is often on a later one, so the collections metric (within 3 s of the line's start) is also reported; it under-counts here.
Data
Everything is synthetic, written from 42 CFR 422/423 subpart V as read on eCFR on 26 Sep 2026 (with the CY2027 amendments
in force since 1 Jun 2026) and the MCMG's telephonic SOA elements. The agency (Prairie Lantern Benefits), the MA
organizations (Northfield Harbor Health Plan, Juniper Ridge Health), their three plans and Summaries of Benefits
(scripts/medicare_eval/plans.py) and every person are fictional; no Medicare numbers are spoken.
- Dev (3 calls). Used to write the prompts. Run 1 scored 30/31 typed answers: "the Medicare enrollment center" was not read as claiming to be Medicare. The endorsement question now names that pattern (dev run 2: 31/31). No prompt changed after that.
- Test (10 calls: 4 clean, 6 planted) and Test B (4 calls: 2 clean, 2 planted): both written before any model run.
- One code change after test run 1: the needs topic "premiums" now also counts when the extraction found the plan's premium and copays gone over (the pre-enrollment checklist's costs item). Test run 1 had two needs items at review on calls where the premium was plainly stated. Test is therefore not fully held out for the needs item; test B was run only after the change and never used to change anything.
- Audio (8 calls from test and test B). Each line spoken by a Decosa house voice (Kokoro-82M stock voicepacks,
Apache-2.0, on CPU: af_heart for the agent, am_michael for the beneficiary), then transcribed and diarized by
MOSS-Transcribe-Diarize through
/medicare/transcribe(the diarizer the other verticals use); the transcripts keep the ASR errors. Gold times come from the audio assembly.- The audio changed on 26 Sep 2026. The first build used macOS system voices, which Apple licenses for personal
use, so it was re-voiced. Each line keeps the start time it had (
scripts/medicare_eval/timeline.json; some lines were spoken up to 1.22x faster to fit), the eight calls were transcribed again and the audio set was re-run twice with the same scorer. The text sets do not use audio and were not re-run. - Consent.
make_audio.pyasks the consent ledger's gate for each voice before it writes a call (projectdecosa-medicare-demo, purposecharacter_dialogue). The entries are operator-owned and scoped to this demo only: ce_e45bc1f19f0c (af_heart), ce_ba7fd6db3efe (am_michael), ce_4c36fbd57f30 (af_nicole, for a caregiver; no audio call has one). The signed decisions per call are indocs/evals/medicare-call-record/voices.json, and one speaker check per voice indemo_voices/speaker-checks.json.
- The audio changed on 26 Sep 2026. The first build used macOS system voices, which Apple licenses for personal
use, so it was re-voiced. Each line keeps the start time it had (
- Scorer changes after the audio runs, before this write-up: the timestamp rule above (inside the spoken line) and the
claim matcher (two claims on one line are matched by key words before position). Neither changes what the API returned;
both only change how a returned item is matched to a gold line. All runs were re-scored the same way
(
run_eval.py --rescore).
Planted problems:
- typed: "free" for a $0 premium (3 calls), claiming to be Medicare (2), an Advantage plan called a supplement, a $50 gift card, switch-today pressure and a false last-day deadline, no recording notice, an enrollment with no enrollment notice, a vague yes, and no Summary of Benefits location;
- call sheet: benefits before the disclaimer (3 calls), wrong disclaimer numbers, a cold call, no SOA, an SOA over 12 months old, Medigap outside an MA-only SOA, final-expense life insurance pitched (2), needs topics and the checklist skipped before an enrollment, and a July call where the disclaimer came before benefits but after the first minute;
- claims: a $3,000 dental allowance (the plan: $1,500), $100 a month OTC (the plan: $50 a quarter), hearing aids "at no cost" (the plan: $699 or $999 per aid), a $150 Part B giveback on a plan with none, "no copays at all for specialists" (the plan: $35).
The clean calls include hard cases: "free" used correctly for $0 preventive care and a no-cost gym; "Medicare-approved"; real enrollment-period dates; an agency that sells every plan (the second disclaimer wording); an SOA taken on the call for a drug-plan enrollment (Part D needs list); a caregiver on the call with the beneficiary giving the yes; and a March 2026 call under the old first-minute rule with the SHIP wording.
Results
| Set | Per-rule accuracy | Planted typed answers found | Call-sheet items | Benefit claims (good ok / bad flagged) | False alarms on clean calls | Timestamps |
|---|---|---|---|---|---|---|
| Dev run 2 (after the prompt change) | 31/31 | 2/2 | 23/23 | 8/9 / 1/1 | 1 of 43 | 12/12 |
| Test, run 1 (before the code change) | 101/101 | 9/9 | 77/79 | 25/25 / 4/4 | 3 of 87 | 32/33 on the gold line |
| Test, run 4 (after the change) | 101/101 | 9/9 | 79/79 | 25/25 / 4/4 | 3 of 87 | 33/33 |
| Test B, run 1 (held out) | 40/40 | 2/2 | 31/32 | 8/8 / 2/2 | 0 of 42 | 15/15 |
| Test B, run 2 | 40/40 | 2/2 | 31/32 | 8/8 / 2/2 | 0 of 42 | 15/15 |
| Audio (ASR transcripts), run 1 | 81/81 | 9/9 | 63/63 | 18/19 / 2/3 | 4 of 65 | 28/29 inside the spoken line (21/29 within 3 s of its start; mean distance from the line start 2.5 s, max 14.3 s) |
| Audio, run 2 | 81/81 | 9/9 | 63/63 | 18/19 / 2/3 | 3 of 65 | 29/29 (20/29) |
| Audio, first build (macOS voices), run 1 | 81/81 | 9/9 | 63/63 | 19/19 / 3/3 | 3 of 65 | 29/29 (16/29; mean 3.6 s, max 13.7 s) |
| Audio, first build, run 2 | 81/81 | 9/9 | 63/63 | 18/19 / 3/3 | 4 of 65 | 29/29 (16/29) |
The audio rows are for the house-voice audio (see Data); the first build's rows are kept for comparison. The typed checks, the call-sheet items and the planted answers are unchanged. The claims moved on one call, t9, on both runs: see Errors. Run 1's one timestamp outside its line was t1's "intent confirmed", cited 8.6 s from the gold line.
(Test runs 2 and 3 fell in the gateway outage: most model calls failed. They are kept in the folder but not counted. The 503-with-no-record behaviour added after run 2 is what a caller now gets in that case.)
Per rule, test run 4: recording notice 10/10, TPMO disclaimer 10/10, SOA read on the call 1/1, "free" 10/10, claimed Medicare 10/10, supplement 10/10, inducement 10/10, pressure 10/10, enrollment notice 10/10, clear yes 10/10, SB location 10/10. Call-sheet items: disclaimer timing 10/10, disclaimer numbers 10/10, SOA on file 9/9, SOA scope 10/10, non-health 10/10, cold call 10/10, needs 10/10, pre-enrollment checklist 10/10.
- First discussion of benefits (which the timing item rests on): on the gold line in 10/10 test, 4/4 test B and 8/8 audio calls (first build and re-voiced audio). Enrollment taken or not: 10/10, 4/4, 8/8. Products discussed (the scope and non-health items): exact on every call.
- Latency per call, under shared load: test B medians 31 s (run 1, a busy gateway) and 4.4 s (run 2); audio medians 6.2 s and 6.0 s on the first build, 41.2 s and 13.1 s on the re-voiced audio (a much busier gateway); the slowest single call 71.4 s. A call is 13 to 19 receipted model calls.
- Cost: the smoke sample used 15 model calls and 15,120 tokens, about $0.0058 at the gateway list price.
Errors
- Claims the plan facts do not mention are flagged, even when true. On the clean drug-plan call (t3) the claim extraction also listed "the plan requires network pharmacies" and "emergency out-of-network fills are covered"; the fictional Summary of Benefits does not say either, so both came back "not in the plan facts" (flag). These are 2 of the 3 false alarms on the clean test calls in run 4 (the third: "the plan only adds drug coverage"). It is the intended behaviour for a claim the plan facts lack, and the product says so, but it means the plan facts must be the full Summary of Benefits, not an excerpt.
- ASR errors become claim errors. On the first build's audio, "a two hundred dollar eyewear allowance" was heard as "in-store allowance" in run 2 and flagged unsupported; "Medicare-covered preventive care is free, a zero copay for every member" became "the plan has a $0 copay" (unsupported, t4, both runs). On the re-voiced audio the eyewear line is heard right; the t4 claim is extracted as "the plan has a zero copay for every member" and now comes back contradicted (both runs).
- t9 on the re-voiced audio: a scoring miss after a diarizer split. The diarizer split line 7 across two segments ("...back on your part" / "be premium every month"). The check itself still flagged the planted $150 giveback as contradicted and the $50 gift card as unsupported. But the scorer, which pairs claims with gold lines by time and key words, paired the gold giveback with "the plan has a zero monthly premium" (supported) and the gold zero-premium claim with the gift card (flag), so the flagged giveback landed among the extra claims. Scored the same way as before, that is one bad claim missed and one good claim flagged on both runs: the 2/3 and 18/19 in the audio rows. The scorer was not changed.
- tb4 (test B), needs topics: expected review (2 of 5 topics missing), got flag (3 missing). The premium was stated ("the premium is zero dollars") but the extraction marked neither the needs "costs" topic nor the checklist costs item, so the code change above did not apply. Same on both runs.
- Test run 1: the two needs items described under Data, and one timestamp: a stitched quote ("I'm now taking your enrollment request... Your Summary of Benefits is there too.") cited its longest piece, the enrollment line, not the SB line (the engine's rule from collections).
- Dev: "The out-of-pocket maximum is $4,900" came back partial (the plan facts say "$4,900 for in-network services"), so review rather than ok, in one of two runs.
- Outside the plan's claims: a Medigap Plan G price ("about $140 a month") is judged against the MA plan's facts and comes back contradicted or unsupported. It is on a planted call (t7, where Medigap is already flagged as outside the SOA), so it is not counted as a clean-call false alarm, but it is noise.
Honest limits
- 17 short scripted calls (under two minutes) with clean TTS audio (two stock synthetic voices). Real sales calls run 20-60 minutes, with accents, crosstalk, hold music, transfers from lead generators and Spanish; none of that is measured. Long calls also mean many more benefit claims than the 12 the extraction keeps.
- The same author wrote the scripts, the labels, the prompts and the code. The labels are one reading of the rule text, not an independent label set. "Pressure tactics" is not a defined term in the rule; it is checked as misleading conduct.
- Benefit claims are checked against the plan facts given, not against the plan's filed benefits or its Evidence of Coverage. A partial Summary of Benefits produces false flags.
- The call sheet is taken as given (SOA dates, who asked for the call, plan counts). The timing item depends on the model finding the first discussion of benefits; CMS's line between "mentioning" and "discussing" a benefit (91 FR 17448-17449) is a judgement.
- Answer probabilities come from typed-judgment's stated-confidence table, fit on another domain; claim confidences from grounding's calibration table. Neither is calibrated for sales calls.
Expected properties of the sample run (for the rehearsal kit)
For POST /medicare/check {"sample_id": "tb1-showcase"} (the diarized audio transcript) or the bundle's text version:
tpmo_disclaimer_timingisflag(on the audio sample: benefits discussed at 00:20, the disclaimer at 00:42), andtpmo_disclaimerisyes.free_misuseisyes/flagwith a quote containing "free premiums".non_healthisflag, naming life insurance (the final-expense pitch).- The dental claim ("$3,000 dental allowance") is
flag, verdictcontradicted, with the span "Comprehensive dental allowance of $1,500 per year." retention.audio_untilis2029-11-05andretain_untilis2032-11-05;POST /record/verifyonrecordreturnsok: true.- About 15 receipts, every one signed.
Files
docs/evals/medicare-call-record/*.json: every run, with the full checklists, claims, extraction and scores;asr-transcripts.jsonholds the diarizer output of the eight audio calls.scripts/medicare_eval/: the calls (convs.py), the plan facts (plans.py), audio (make_audio.py: house voices on our server, gated by the consent ledger;timeline.json), transcription, the runner and scorer (run_eval.py), the demo samples and the site's replay fixture.