Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: Collections and servicing call QA (55)

Run on 26 Sep 2026 against the pre-release server (127.0.0.1:8455; the audio set again on 127.0.0.1:8479 after the audio was re-voiced) through the model gateway, with Qwen3.8-27B at temperature 0. Every model call had a signed gateway receipt. The gateway was shared with other workloads' runs, so latency is as measured under load.

What is measured

Each call goes through POST /collections/check with its call log and a consent statement. The expected answers come from tags on the script lines (scripts/collections_eval/convs.py), not from a model. The expected call-log findings were written by hand from the rule text before the log code was run.

  • Per-rule accuracy. The typed answer (yes, no, unclear or n/a) equals the expected answer. Five answers accept more than one value:
    • the mini-Miranda and identity checks on the two third-party calls accept n/a or no;
    • t5's partial disclosure accepts no or unclear;
    • tb2's identity check accepts any answer, because the spouse confirmed the address but not that she is the consumer;
    • tb3's loss-mitigation check accepts n/a or yes.
  • Planted recall. Of the expected answers that mark a problem (for example "false or misleading? yes" or "debt-collector disclosure? no"), how many came back that way.
  • Extraction. Requests found in what the called person said: stop calling, a lawyer, call back, a dispute, a hardship. Also: requests invented where none was made, and whether the cited line is right.
  • Call-log findings. For each hand-labelled item, both the status and the exact set of calls behind it must match. Any item without a label must come back ok or n/a. These findings are computed in code from the log, plus the events the extraction places on it (a stop request, a lawyer, consent to call back).
  • False alarms. On the clean calls, how many call checks or log items came back as flag or review.
  • Timestamp accuracy. Scored for every correct "yes": does the quote's line fall on a gold line? On the audio set, a hit means the cited time is within 3 s of the start of a gold line in the audio.

Data

Everything is synthetic, written from 12 CFR part 1006 and 12 CFR 1024.39-1024.41 as read on eCFR on 26 Sep 2026. The agency (Harbor Ridge Recovery), the servicer (Cedar Gate Mortgage Servicing), the card issuer (Brightwater Card Company) and every person are fictional. Phone numbers are in the 555-01xx range, which is reserved for fiction. The CFPB complaint database suggested in the wiki was not used: the scripts were written directly from the rule text.

  • Dev (3 calls). Used to write the prompts. The first run scored 25/25. Two things changed after it, and no prompt changed:

    • one label was fixed ("my hours got cut" is a hardship);
    • quote citing: a quote stitched from pieces now cites its longest piece (it had cited the first, often a greeting).

    The gold lines for "identity confirmed" were also widened to include the name question and the answer to it. Both changes were made before any test run.

  • Test (10 calls: 4 clean, 6 planted). Written before any model run.

  • Test B (4 calls: 2 clean, 2 planted). Written after test run 1, before any run of these. Writing tb4 ("don't call me at work") led to one change before any run. The extraction now has a "work" scope that bars the numbers marked kind: "work". Before, it mapped to the number of the recorded call. Test was re-run after that change (run 2).

  • Audio (6 calls from test and test B). Each line is spoken by a Decosa house voice (Kokoro-82M stock voicepacks, Apache-2.0, on CPU: af_heart for the agent, am_michael for the consumer, am_fenrir for the borrower), then transcribed and diarized by MOSS-Transcribe-Diarize, through /collections/transcribe on the same diarizer the other verticals use. The transcripts keep the ASR errors. Gold times come from the audio assembly.

    • The audio changed on 26 Sep 2026. The first build used macOS system voices, which Apple licenses for personal use, so it was re-voiced the same day. Each line keeps the start time it had (scripts/collections_eval/timeline.json; three lines were spoken up to 1.06x faster to fit), so the times quoted below and in the rehearsal bundle still hold. The six calls were then transcribed again and the audio set was re-run twice. The text sets do not use audio and were not re-run.
    • Consent. Before a call is spoken, make_audio.py asks the consent ledger's gate for each voice (project decosa-collections-demo, purpose character_dialogue), with that voice's lines as a sample for the speaker check. The ledger has an operator-owned entry per voice, scoped to this demo only and enrolled through its API with scripts/demo_voices.py enroll (ce_7d7887b6365c af_heart, ce_bc39c2ffd91e am_michael, ce_e2b3818b07d5 am_fenrir). The signed decisions for every call are in docs/evals/collections-call-qa/voices.json. All 12 were allowed, and every voice matched its enrolled consent clip (scores 0.82 to 0.93, threshold 0.585).

Planted problems cover:

  • debt mentioned before identity was confirmed;
  • a missing or partial debt-collector disclosure;
  • the debt disclosed to a roommate and to a coworker;
  • threats of the sheriff, arrest and garnishment;
  • a paid-off dispute dismissed;
  • "stop calling" refused;
  • a lawyer named, then more calls;
  • no recording notice;
  • a servicer that demands the arrears before any modification and mentions no options;
  • in the logs: 8 calls in 7 days, calls within 7 days after a conversation, a 7:10 a.m. call to a Denver mobile, a 9:20 p.m. call, calls after a stop request (including a work number only), calls after a lawyer was named, a coworker called twice, and no live contact by day 36.

The clean calls include three hard cases:

  • a "call me back Thursday" that makes a call 2 days after a conversation allowed (1006.14(b)(3)(i));
  • a spouse who is treated as the consumer (1006.6(a)(1));
  • a borrower who says they forgot and will pay in full, so offering loss-mitigation options is not needed (comment 39(a)-4.i.B).

Results

Set Per-rule accuracy Planted answers found Call-log findings False alarms on clean calls Timestamps on the gold line
Dev (after the two changes) 25/25 4/4 8/8 0 of 21 9/9
Test, run 1 81/83 11/12 33/33 0 of 48 31/31
Test, run 2 (after the "work" scope) 81/83 11/12 33/33, plus 1 unlabelled flag 0 of 48 31/31
Test B, run 1 (held out) 32/33 3/3 11/11 2 of 21 10/10
Test B, run 2 32/33 3/3 11/11 1 of 21 10/10
Audio (ASR transcripts), run 1 48/49 8/9 17/17 0 of 11 19/21 (mean error 0.53 s, max 4.2 s)
Audio, run 2 48/49 8/9 17/17 0 of 11 20/21 (mean error 0.33 s, max 3.4 s)
Audio, first build (macOS voices), runs 1 and 2 48/49 8/9 17/17 0 of 11 20/21 (mean error 0.56 s, max 3.4 s)

The audio rows are for the house-voice audio (see Data). Against the first build, the only change is one timestamp on run 1. The answers, the planted finds, the log findings and the extraction are the same on every run.

Per rule, test run 1:

Rule Correct
recording disclosed 10/10
debt-collector disclosure 9/9
identity confirmed 10/10
debt before identity 9/10
third-party disclosure 10/10
false or misleading 10/10
stop request refused 9/10
dispute dismissed 10/10
loss-mitigation options mentioned 2/2
loss mitigation misstated 2/2
  • Extraction (test): 8/8 requests found, every one on the right line; 0 of 42 invented in run 1, 1 of 42 in run 2.
  • Call-log maths: 61/61 hand-labelled items right across test, test B and audio (first runs). The code is deterministic, so repeat runs only change when the extraction changes the events it adds. That happened once: see t10 under Errors. It is also covered by unit tests: DST, two time zones, prior-consent windows, unconnected calls, the one allowed reply after the person calls in, and the day-36 deadline.
  • Latency per call (median, under shared load): 4 to 5 s. The slowest single call took 16.3 s.
  • Cost of a check: 9 or 10 receipted model calls. The smoke sample used 9 calls and 6,321 tokens, about $0.0026 at the gateway list price.

Errors

  • t6, "debt mentioned before identity was confirmed". Expected yes, got no on every run (text and audio). The agent's first line is "This is Derek with Harbor Ridge Recovery, a debt collector. ... Is this Tom Kowalski?". The label says that naming yourself a debt collector before confirming who answered reveals the collection purpose. The model reads it as not mentioning "the debt". It is a strict label, but it is a real miss for a QA team that treats it as a fail.
  • t10, "stop request refused". Expected n/a, got no on both runs. "Please talk to her [my lawyer] from now on" was read as a stop request that the agent accepted. The status is ok either way. The lawyer event itself was found, and the calls after it were flagged for review.
  • t10, test run 2: an unlabelled log flag. After the "work" scope change, the extraction read "Please talk to her from now on" as a stop request. So "calls after a stop request" flagged the two later calls, which were already under review as calls after a lawyer was named. This is a request the label does not contain. It is on a planted call, so it is not counted as a clean-call false alarm, but it is an extra flag.
  • tb2 (clean), debt-collector disclosure with the spouse. Expected yes, got no on run 1 (a false flag) and n/a on run
    1. The model reasoned that the spouse is not the consumer. For 1006.18 that reading is arguable: the spouse counts as the consumer only for 1006.6 (1006.2(e)). The disclosure was plainly given, so "no" is wrong.
  • tb2 identity check came back "no" (review) on both runs. The label accepts it, but it still counts as a clean-call review item.
  • Hardship was extracted where none was said ("I can't pay it right now") once per run on tb1. It only shows in the "said" panel and changes no check.
  • Audio timestamps. On both runs, t6's recording notice was cited 3.4 s late. The diarizer split the greeting, and the quote landed on the later piece, which holds the words "calls are recorded" (the same miss as on the first build). On run 1, tb1's debt-collector disclosure was also cited 4.2 s late. The diarizer split that line into three pieces, and the quote landed on the third ("and any information obtained will be used for that purpose").

Honest limits

  • 17 short scripted calls with clean TTS audio (two stock synthetic voices per call). The following are not measured:
    • real collection calls: accents, crosstalk, Spanish (1006.18(e)(4) needs the disclosure in the call's language), long calls, and agents reading scripts at speed;
    • voicemails and limited-content messages.
  • The same author wrote the scripts, the labels, the prompts and the log code. The labels are one reading of the rule text, not an independent label set. The log labels were written before the code ran, but by the same person.
  • The call log is taken as given. The presumptions depend on facts outside it: letters, texts and emails, consent given elsewhere, whether an attorney's name and address were known, and debts split per account. Counting days in the consumer's first time zone is a choice the rule does not spell out.
  • Answer probabilities come from typed-judgment's stated-confidence table, fit on public dev sets from another domain. They are not calibrated for collection calls.

Expected properties of the sample run (for the rehearsal kit)

For POST /collections/check {"sample_id": "tb1-showcase"}:

  1. false_misleading is yes / flag, with a quote containing "sheriff" at about 00:27 (the sample is the diarized audio transcript).
  2. freq_7in7 is flag, max_in_7_days is 8, and calls is ["c8"].
  3. calling_hours is flag, and its first call is c1 with outside[0] = {"zone": "America/Denver", "local": "07:10"}.
  4. mini_miranda is yes / ok (initial communication: call.initial_communication is true).
  5. retain_until is 2029-03-08 (three years after the call date), and POST /record/verify on record returns ok: true.
  6. The number of receipts is 9, and every receipt is signed.

Files

  • docs/evals/collections-call-qa/*.json: every run, with the full checklists, extraction, log items and scores.
  • docs/evals/collections-call-qa/asr-transcripts.json: the diarizer output, the speech receipts and the gold line starts for the audio set.
  • scripts/collections_eval/:
    • convs.py: the role-plays, tags and log labels;
    • run_eval.py;
    • make_audio.py: the house voices, gated by the consent ledger (timeline.json holds the line start times it keeps);
    • transcribe_audio.py;
    • build_samples.py.
  • docs/evals/collections-call-qa/voices.json: the voice, speed and signed consent decision for every line of the audio.
  • scripts/demo_voices.py: the consent entries (enrol) and the gate call. The consent clips are in demo_voices/consent-clips/decosa-collections-demo/.