Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: Auto F&I disclosure record (46)

Run on 25 Sep 2026 against the pre-release server (127.0.0.1:8466; the audio set again on 26 Sep on 127.0.0.1:8479 after the audio was re-voiced) through the model gateway, with Qwen3.8-27B at temperature 0. Every model call had a signed gateway receipt. The gateway was shared with other workloads' runs, so latency is as measured under load.

What is measured

Each conversation goes through POST /fi/check with its deal jacket and a consent statement. The expected answers come from tags on the script lines (scripts/fi_eval/convs.py), not from a model.

  • Per-check accuracy. The typed answer (yes, no, unclear or na) equals the expected answer. Three checks accept more than one answer: in t6, "cancel right said" accepts no or unclear, and t8's two "presented as required" checks are ambiguous and unscored.
  • Planted recall. Of the expected answers that mark a planted problem (for example "presented as required? yes" or "said optional? no"), how many came back that way.
  • Mismatch recall. Planted jacket-versus-conversation mismatches found, matched by kind and add-on.
  • False alarms. On the clean conversations, how many conversation checks or mismatches came back as flag or review. Jacket-only items are deterministic and not counted.
  • Timestamp accuracy. For every correct "yes", whether the quote's line is a gold line. For consent, the gold lines are the customer's agreement and the staff statement one or two lines above it. On the audio set, a hit means the cited time is within 3 s of a gold line's start on the audio.

Data

All synthetic: a fictional dealer (Larkspur Point Motors) and invented customers. The role-plays were written from the text of SB 766 and Penal Code 632.

  • Dev (3 conversations). Used to write the prompts. The first run scored 25/25, and no prompt was changed after it.
  • Test (10 conversations: 4 clean, 6 with planted problems). Written before any model run.
  • Test B (4 conversations: 2 clean, 2 planted). Written after test run 1, before any run of these. Never used to change anything.
  • Audio (5 conversations from test and test B). Each line is spoken by a Decosa house voice (Kokoro-82M stock voicepacks, Apache-2.0, on CPU: af_heart for the finance manager, am_michael for the customer), then transcribed and diarized by MOSS-Transcribe-Diarize (the same service as the hosted diarizer). The transcripts keep the ASR errors. Gold times come from the audio assembly.
    • The audio changed on 26 Sep 2026. The first build used macOS system voices, which Apple licenses for personal use, so it was re-voiced. Each line keeps the start time it had (scripts/fi_eval/timeline.json; some lines were spoken up to 1.1x faster to fit), the five recordings were transcribed again and the audio set was re-run twice. The first build's transcripts had errors such as "GAAP" for GAP and "twenty-one-nine"; the new ones do not have those two, and have about the same number of errors overall. The text sets do not use audio and were not re-run.
    • Consent. make_audio.py asks the consent ledger's gate for each voice before it writes a recording (project decosa-fi-demo, purpose character_dialogue). The entries are operator-owned and scoped to this demo only: ce_949d9af2e2e5 (af_heart), ce_e029f82f6385 (am_michael), ce_c96885097703 (af_nicole, for a co-buyer; no audio call has one). The signed decisions per recording are in docs/evals/fi-disclosure-record/voices.json, and one speaker check per voice (the voice's lines against its enrolled consent clip) in demo_voices/speaker-checks.json.

Results

Set Per-check accuracy Planted answers found Mismatches found False alarms on clean calls Timestamps on the gold line
Test, run 1 (before the two fixes) 76/78 12/13 6/6 2 of 30 items (1 check, 1 mismatch) 38/39
Test, run 2 (after the fixes) 77/78 12/13 6/6 0 of 30 40/40
Test, repeat 77/78 12/13 6/6 0 of 30 40/40
Test B, run 1 (held out) 31/32 8/8 3/3 1 of 14 15/15
Test B, repeat 32/32 8/8 3/3 0 of 14 15/16
Audio (ASR transcripts), run 1 45/45 10/10 6/6 0 of 9 21/21 (mean error 0.04 s)
Audio, repeat 45/45 10/10 6/6 0 of 9 20/21 (mean 0.22 s, max 3.7 s)
Audio, first build (macOS voices), run 1 45/45 10/10 6/6 0 of 9 21/21 (mean error 0.19 s)
Audio, first build, repeat 44/45 9/10 6/6 0 of 9 19/21 (mean 0.53 s, max 3.7 s)

The audio rows are for the house-voice audio (see Data); the first build's rows are kept for comparison. The repeat run on the new audio did not repeat the first build's t6 miss.

Per check family on test run 2: add-on said optional 19/19, add-on presented as required 17/17, consent 10/10, total of payments 10/10, lower-payment warning 10/10, cancel right misstated 6/6, cancel right explained 5/6.

Latency per conversation (median, under shared load): 3.5 to 4.2 s for text transcripts, and 7.5 s and 16.6 s on the two audio runs of the first build; 19.0 s and 25.4 s on the re-voiced audio, when the gateway was busier. A check uses about 10 to 14 receipted calls. The smoke sample used 10 calls and 6,382 tokens, about $0.003 at the gateway list price.

Errors, and what changed

Test run 1 led to two fixes. The test set was then re-run, and test B was written as a fresh held-out set:

  1. Vehicle prices said in thousands. The extraction read "thirty-eight nine" as $389, which gave false vehicle-price mismatches on 3 calls. The fix added the shorthand to the extraction rules (with examples) and a guard: a spoken vehicle price below half or above twice the jacket price is not compared, and the result says so.
  2. Quotes that run across lines. A consent quote spanning the staff's question and the customer's "yes" could not be matched, so a correct "yes" became "unclear". Quote matching now accepts two or three consecutive lines, and pieces joined with an ellipsis. The ellipsis part was added after test B run 1, which had one such miss.
  3. Scoring. The consent gold lines were widened to the staff statement two lines up. In run 1, one correct cite counted as wrong because of this. This is a scoring change, not a model change.

Remaining errors:

  • t6, "cancel right explained". Expected no. It came back yes on 4 of 5 runs with the first build (3 text, 1 of 2 audio), and no on both runs on the re-voiced audio, quoting "if you did cancel there's a thousand dollar restocking fee". The model reads a hint at cancelling as an explanation of the right. The same call is still flagged by "cancel right misstated" (6/6) and by the restocking-fee cap check.
  • Stitched quotes land on the neighbouring line. When the model quotes the monthly payment and the total together, the cited time is the first piece's line. In two audio cases that was the line before the total, 3.5 to 3.7 s early (one case, 3.7 s, on the re-voiced audio).
  • t9 on test run 2. The extraction marked the customer's "Oh, okay" after "it comes with the car" as no clear yes, and flagged it for review. That is arguably right.

Honest limits

  • 17 short scripted conversations with clean TTS audio (two stock synthetic voices). Real F&I rooms have crosstalk, menus read at speed, Spanish and other languages (1784.41 requires the written disclosures in the negotiation language), and 30 to 60 minutes of talk. None of that is measured here.
  • The expected answers are one author's reading of the scripts. The same author wrote the prompts, so this is not an independent label set.
  • Answer probabilities come from typed-judgment's stated-confidence table, fit on public dev sets from another domain. They are not calibrated for F&I talk. Read the answer and its quote, not the probability.
  • The jacket-only rules are deterministic. They are tested in tests/test_fi.py, not in this eval.
  • It checks what was said and what the jacket says was given. It cannot tell whether a written disclosure was clear and conspicuous, in the right language, or given before the representation.

Files

  • docs/evals/fi-disclosure-record/*.json: every run, with the full checklists, mismatches and scores.
  • docs/evals/fi-disclosure-record/asr-transcripts.json: the diarizer output and gold line starts for the audio set.
  • scripts/fi_eval/: convs.py (the role-plays and tags), run_eval.py, make_audio.py (house voices, gated by the consent ledger; timeline.json holds the line starts it keeps), transcribe_audio.py, build_samples.py.
  • docs/evals/fi-disclosure-record/voices.json: the voice, speed and signed consent decision for every recording.