Skip to content
decosa

55 · Finance and insurance · Compliance and trust · live

Collections and servicing call QA

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)

  • Per-rule accuracy, test run 181/83test splitn = 83Run 2: 81/83
  • Per-rule accuracy, test B run 1 (written after test run 1)32/33held outn = 33
  • Planted answers found, test run 111/12test splitn = 12Test B: 3/3; audio: 8/9
  • False alarms on clean calls: test / test B run 10 of 48 / 2 of 21test splitn = 69Test B run 2: 1 of 21
  • Call-log findings (hand-labelled, first runs)61/61test splitn = 61Test, test B and audio
  • Per-rule accuracy on ASR transcripts of TTS audio, run 148/49test splitn = 49Audio re-voiced with Decosa house voices (Kokoro-82M) and re-run: accuracy unchanged from the earlier macOS-voice audio. Timestamps on the gold line 19/21 (mean error 0.53 s, max 4.2 s); run 2: 20/21

Dataset

17 synthetic scripted calls written from 12 CFR part 1006 and 1024.39-41 with fictional companies and people: 3 dev, 10 test, 4 test B, plus 6 of them as TTS audio (Decosa house voices, Kokoro-82M, each allowed by the consent ledger) run through the diarizer.

Caveats

  • The same author wrote the scripts, the labels, the prompts and the log code; the labels are one reading of the rule text.
  • 17 short scripted calls with clean TTS audio; real calls (accents, crosstalk, Spanish, long calls), voicemails and limited-content messages are not measured.
  • The call log is taken as given; presumptions depend on facts outside it.
  • Answer probabilities come from a confidence table fit on another domain and are not calibrated for collection calls.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
26 Sep 2026
Latency, this run
n/a
p50 over passed runs
4.2 s
Receipts
9
Model calls
n/a
Tokens
n/a
Cost per run
$0.003

Self-host verification

Verified on 26 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

The assembly prompt's smoke tests ran against the already-running local Qwen3.8-27B vLLM (127.0.0.1:8114, network_mode host instead of the compose llm service): tb1-showcase flagged the sheriff threat at 00:27, 7-in-7 (c8) and calling hours (c1), 9 attested receipts, record verified, 3.4 s; the pasted 21:30 New York call flagged the threat and the hours; no consent gave 400. One prompt bug found and fixed: the log-only step fed back the normalised call log, which the API does not accept as input. Model-server startup and the diarizer path were not re-run.

Rehearsal bundle: collections-call-qa.zip (3 KB, 12 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.
  • Measured on 17 synthetic role-plays written by the building agent, with clean TTS audio. Not measured on real collection calls (accents, Spanish, long calls, voicemails) or with an independent reviewer's labels.
  • One systematic miss: an agent who names themselves a debt collector before confirming who answered is not flagged as revealing the debt before identity (t6 on every run).
  • The call-log findings are the rule's presumptions computed from the log given. Letters, texts, emails, consent given elsewhere and an attorney's response are outside it unless added as events. Days are counted in the consumer's first time zone.
  • Audio intake (/collections/transcribe) is self-host only. The hosted demo's audio samples were transcribed on our server and are bundled.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Extraction (what the called person asked for or said, with the line) and one typed yes/no/unclear check per QA questionQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
  • Recording to a timed, speaker-labelled transcript (POST /collections/transcribe, self-host; the demo's audio samples were transcribed with it)MOSS-Transcribe-Diarize 0.9BApache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 48 GB card (1)
  • typed checks correct / planted problems found: not measured yet
Standard · the hosted demo, one 96 GB card (5)
  • held-out test B, 4 synthetic calls: typed checks correct / planted problems found / call-log findings: 32/33 / 3/3 / 11/11decosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route, run 1; data, labels and prompts written by the building agent, prompts frozen on a separate 3-call dev split
  • test set, 10 calls: typed checks correct / planted found / call-log findings: 81/83 / 11/12 / 33/33 (repeat run the same, plus one unlabelled log flag)decosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route
  • false alarms on clean calls (flag or review): test 0/48; test B 2/21 (repeat 1/21)decosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route
  • 6 calls on real speech-recognition transcripts of synthetic audio: checks / planted / log findings / cited time within 3 s: 48/49 / 8/9 / 17/17 / 20/21decosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route (asr runs)
  • real collection calls reviewed by a compliance QA lead: not measured yet
Best · DeepSeek-V4-Flash on two more cards (1)
  • typed checks correct / planted problems found: not measured yet
Wanted · two large judges from different families (1)
  • typed checks correct / planted problems found: not measured yet

How we measure · All tools