Eval: Collections and servicing call QA (55)
Run on 26 Sep 2026 against the pre-release server (127.0.0.1:8455; the audio set again on 127.0.0.1:8479 after the audio was re-voiced) through the model gateway, with Qwen3.8-27B at temperature 0. Every model call had a signed gateway receipt. The gateway was shared with other workloads' runs, so latency is as measured under load.
What is measured
Each call goes through POST /collections/check with its call log and a consent statement. The expected answers come from
tags on the script lines (scripts/collections_eval/convs.py), not from a model. The expected call-log findings were
written by hand from the rule text before the log code was run.
- Per-rule accuracy. The typed answer (yes, no, unclear or n/a) equals the expected answer. Five answers accept more
than one value:
- the mini-Miranda and identity checks on the two third-party calls accept n/a or no;
- t5's partial disclosure accepts no or unclear;
- tb2's identity check accepts any answer, because the spouse confirmed the address but not that she is the consumer;
- tb3's loss-mitigation check accepts n/a or yes.
- Planted recall. Of the expected answers that mark a problem (for example "false or misleading? yes" or "debt-collector disclosure? no"), how many came back that way.
- Extraction. Requests found in what the called person said: stop calling, a lawyer, call back, a dispute, a hardship. Also: requests invented where none was made, and whether the cited line is right.
- Call-log findings. For each hand-labelled item, both the status and the exact set of calls behind it must match. Any item without a label must come back ok or n/a. These findings are computed in code from the log, plus the events the extraction places on it (a stop request, a lawyer, consent to call back).
- False alarms. On the clean calls, how many call checks or log items came back as
flagorreview. - Timestamp accuracy. Scored for every correct "yes": does the quote's line fall on a gold line? On the audio set, a hit means the cited time is within 3 s of the start of a gold line in the audio.
Data
Everything is synthetic, written from 12 CFR part 1006 and 12 CFR 1024.39-1024.41 as read on eCFR on 26 Sep 2026. The agency (Harbor Ridge Recovery), the servicer (Cedar Gate Mortgage Servicing), the card issuer (Brightwater Card Company) and every person are fictional. Phone numbers are in the 555-01xx range, which is reserved for fiction. The CFPB complaint database suggested in the wiki was not used: the scripts were written directly from the rule text.
Dev (3 calls). Used to write the prompts. The first run scored 25/25. Two things changed after it, and no prompt changed:
- one label was fixed ("my hours got cut" is a hardship);
- quote citing: a quote stitched from pieces now cites its longest piece (it had cited the first, often a greeting).
The gold lines for "identity confirmed" were also widened to include the name question and the answer to it. Both changes were made before any test run.
Test (10 calls: 4 clean, 6 planted). Written before any model run.
Test B (4 calls: 2 clean, 2 planted). Written after test run 1, before any run of these. Writing tb4 ("don't call me at work") led to one change before any run. The extraction now has a "work" scope that bars the numbers marked
kind: "work". Before, it mapped to the number of the recorded call. Test was re-run after that change (run 2).Audio (6 calls from test and test B). Each line is spoken by a Decosa house voice (Kokoro-82M stock voicepacks, Apache-2.0, on CPU: af_heart for the agent, am_michael for the consumer, am_fenrir for the borrower), then transcribed and diarized by MOSS-Transcribe-Diarize, through
/collections/transcribeon the same diarizer the other verticals use. The transcripts keep the ASR errors. Gold times come from the audio assembly.- The audio changed on 26 Sep 2026. The first build used macOS system voices, which Apple licenses for personal
use, so it was re-voiced the same day. Each line keeps the start time it had (
scripts/collections_eval/timeline.json; three lines were spoken up to 1.06x faster to fit), so the times quoted below and in the rehearsal bundle still hold. The six calls were then transcribed again and the audio set was re-run twice. The text sets do not use audio and were not re-run. - Consent. Before a call is spoken,
make_audio.pyasks the consent ledger's gate for each voice (projectdecosa-collections-demo, purposecharacter_dialogue), with that voice's lines as a sample for the speaker check. The ledger has an operator-owned entry per voice, scoped to this demo only and enrolled through its API withscripts/demo_voices.py enroll(ce_7d7887b6365c af_heart, ce_bc39c2ffd91e am_michael, ce_e2b3818b07d5 am_fenrir). The signed decisions for every call are indocs/evals/collections-call-qa/voices.json. All 12 were allowed, and every voice matched its enrolled consent clip (scores 0.82 to 0.93, threshold 0.585).
- The audio changed on 26 Sep 2026. The first build used macOS system voices, which Apple licenses for personal
use, so it was re-voiced the same day. Each line keeps the start time it had (
Planted problems cover:
- debt mentioned before identity was confirmed;
- a missing or partial debt-collector disclosure;
- the debt disclosed to a roommate and to a coworker;
- threats of the sheriff, arrest and garnishment;
- a paid-off dispute dismissed;
- "stop calling" refused;
- a lawyer named, then more calls;
- no recording notice;
- a servicer that demands the arrears before any modification and mentions no options;
- in the logs: 8 calls in 7 days, calls within 7 days after a conversation, a 7:10 a.m. call to a Denver mobile, a 9:20 p.m. call, calls after a stop request (including a work number only), calls after a lawyer was named, a coworker called twice, and no live contact by day 36.
The clean calls include three hard cases:
- a "call me back Thursday" that makes a call 2 days after a conversation allowed (1006.14(b)(3)(i));
- a spouse who is treated as the consumer (1006.6(a)(1));
- a borrower who says they forgot and will pay in full, so offering loss-mitigation options is not needed (comment 39(a)-4.i.B).
Results
| Set | Per-rule accuracy | Planted answers found | Call-log findings | False alarms on clean calls | Timestamps on the gold line |
|---|---|---|---|---|---|
| Dev (after the two changes) | 25/25 | 4/4 | 8/8 | 0 of 21 | 9/9 |
| Test, run 1 | 81/83 | 11/12 | 33/33 | 0 of 48 | 31/31 |
| Test, run 2 (after the "work" scope) | 81/83 | 11/12 | 33/33, plus 1 unlabelled flag | 0 of 48 | 31/31 |
| Test B, run 1 (held out) | 32/33 | 3/3 | 11/11 | 2 of 21 | 10/10 |
| Test B, run 2 | 32/33 | 3/3 | 11/11 | 1 of 21 | 10/10 |
| Audio (ASR transcripts), run 1 | 48/49 | 8/9 | 17/17 | 0 of 11 | 19/21 (mean error 0.53 s, max 4.2 s) |
| Audio, run 2 | 48/49 | 8/9 | 17/17 | 0 of 11 | 20/21 (mean error 0.33 s, max 3.4 s) |
| Audio, first build (macOS voices), runs 1 and 2 | 48/49 | 8/9 | 17/17 | 0 of 11 | 20/21 (mean error 0.56 s, max 3.4 s) |
The audio rows are for the house-voice audio (see Data). Against the first build, the only change is one timestamp on run 1. The answers, the planted finds, the log findings and the extraction are the same on every run.
Per rule, test run 1:
| Rule | Correct |
|---|---|
| recording disclosed | 10/10 |
| debt-collector disclosure | 9/9 |
| identity confirmed | 10/10 |
| debt before identity | 9/10 |
| third-party disclosure | 10/10 |
| false or misleading | 10/10 |
| stop request refused | 9/10 |
| dispute dismissed | 10/10 |
| loss-mitigation options mentioned | 2/2 |
| loss mitigation misstated | 2/2 |
- Extraction (test): 8/8 requests found, every one on the right line; 0 of 42 invented in run 1, 1 of 42 in run 2.
- Call-log maths: 61/61 hand-labelled items right across test, test B and audio (first runs). The code is deterministic, so repeat runs only change when the extraction changes the events it adds. That happened once: see t10 under Errors. It is also covered by unit tests: DST, two time zones, prior-consent windows, unconnected calls, the one allowed reply after the person calls in, and the day-36 deadline.
- Latency per call (median, under shared load): 4 to 5 s. The slowest single call took 16.3 s.
- Cost of a check: 9 or 10 receipted model calls. The smoke sample used 9 calls and 6,321 tokens, about $0.0026 at the gateway list price.
Errors
- t6, "debt mentioned before identity was confirmed". Expected yes, got no on every run (text and audio). The agent's first line is "This is Derek with Harbor Ridge Recovery, a debt collector. ... Is this Tom Kowalski?". The label says that naming yourself a debt collector before confirming who answered reveals the collection purpose. The model reads it as not mentioning "the debt". It is a strict label, but it is a real miss for a QA team that treats it as a fail.
- t10, "stop request refused". Expected n/a, got no on both runs. "Please talk to her [my lawyer] from now on" was read as a stop request that the agent accepted. The status is ok either way. The lawyer event itself was found, and the calls after it were flagged for review.
- t10, test run 2: an unlabelled log flag. After the "work" scope change, the extraction read "Please talk to her from now on" as a stop request. So "calls after a stop request" flagged the two later calls, which were already under review as calls after a lawyer was named. This is a request the label does not contain. It is on a planted call, so it is not counted as a clean-call false alarm, but it is an extra flag.
- tb2 (clean), debt-collector disclosure with the spouse. Expected yes, got no on run 1 (a false flag) and n/a on run
- The model reasoned that the spouse is not the consumer. For 1006.18 that reading is arguable: the spouse counts as the consumer only for 1006.6 (1006.2(e)). The disclosure was plainly given, so "no" is wrong.
- tb2 identity check came back "no" (review) on both runs. The label accepts it, but it still counts as a clean-call review item.
- Hardship was extracted where none was said ("I can't pay it right now") once per run on tb1. It only shows in the "said" panel and changes no check.
- Audio timestamps. On both runs, t6's recording notice was cited 3.4 s late. The diarizer split the greeting, and the quote landed on the later piece, which holds the words "calls are recorded" (the same miss as on the first build). On run 1, tb1's debt-collector disclosure was also cited 4.2 s late. The diarizer split that line into three pieces, and the quote landed on the third ("and any information obtained will be used for that purpose").
Honest limits
- 17 short scripted calls with clean TTS audio (two stock synthetic voices per call). The following are not measured:
- real collection calls: accents, crosstalk, Spanish (1006.18(e)(4) needs the disclosure in the call's language), long calls, and agents reading scripts at speed;
- voicemails and limited-content messages.
- The same author wrote the scripts, the labels, the prompts and the log code. The labels are one reading of the rule text, not an independent label set. The log labels were written before the code ran, but by the same person.
- The call log is taken as given. The presumptions depend on facts outside it: letters, texts and emails, consent given elsewhere, whether an attorney's name and address were known, and debts split per account. Counting days in the consumer's first time zone is a choice the rule does not spell out.
- Answer probabilities come from typed-judgment's stated-confidence table, fit on public dev sets from another domain. They are not calibrated for collection calls.
Expected properties of the sample run (for the rehearsal kit)
For POST /collections/check {"sample_id": "tb1-showcase"}:
false_misleadingisyes/flag, with a quote containing "sheriff" at about 00:27 (the sample is the diarized audio transcript).freq_7in7isflag,max_in_7_daysis 8, andcallsis["c8"].calling_hoursisflag, and its first call isc1withoutside[0] = {"zone": "America/Denver", "local": "07:10"}.mini_mirandaisyes/ok(initial communication:call.initial_communicationis true).retain_untilis2029-03-08(three years after the call date), andPOST /record/verifyonrecordreturnsok: true.- The number of receipts is 9, and every receipt is signed.
Files
docs/evals/collections-call-qa/*.json: every run, with the full checklists, extraction, log items and scores.docs/evals/collections-call-qa/asr-transcripts.json: the diarizer output, the speech receipts and the gold line starts for the audio set.scripts/collections_eval/:convs.py: the role-plays, tags and log labels;run_eval.py;make_audio.py: the house voices, gated by the consent ledger (timeline.jsonholds the line start times it keeps);transcribe_audio.py;build_samples.py.
docs/evals/collections-call-qa/voices.json: the voice, speed and signed consent decision for every line of the audio.scripts/demo_voices.py: the consent entries (enrol) and the gate call. The consent clips are indemo_voices/consent-clips/decosa-collections-demo/.