Skip to content
decosa

67 · Finance and insurance · Compliance and trust · live

Reg E dispute investigation file

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)

  • Clocks agreeing with a second implementation2,763 / 2,763syntheticn = 2,763600 random disputes, 2026 to 2030; the oracle was written by the same author
  • Missing notice elements flagged (recall)0.956 (43/45)test splitn = 4540 synthetic notices, run once after the prompts were frozen
  • Missing-element flags that were right (precision)0.977 (43/44)test splitn = 44
  • Notices with all five elements right37 / 40test splitn = 40name, account, why, date, amount
  • Letter explanation specific or not20 / 20test splitn = 20templated letters
  • Right-to-documents sentence found or missing20 / 20test splitn = 20includes a paraphrase no keyword rule catches
  • Debit notice complete or not20 / 20test splitn = 20
  • Automated decisions0 / 9syntheticn = 95 samples and 4 injection attempts on the real model

Dataset

56 synthetic notices (16 dev, 40 test) from 12 dispute scenarios in four channels; 27 templated results letters (7 dev, 20 test); 600 random disputes for the clocks; 5 demo files and 4 injection attempts.

Caveats

  • The same author wrote the cases, the gold labels and the prompts.
  • Synthetic and templated only; no real bank notices or letters.
  • Letters come from a small set of variants: 20/20 shows the check works on clear cases, not on your templates.
  • Who made the transfer is not in the headline: raw agreement 24/40, with label errors on our side; it only picks a coverage note.
  • The clock oracle checks the code, not the reading of the rule.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
26 Sep 2026
Latency, this run
n/a
p50 over passed runs
38 s
Receipts
8
Model calls
n/a
Tokens
n/a
Cost per run
$0.006

Self-host verification

Verified on 26 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): clocks as expected, cnp-generic-denial gaps with cannot_close in 18.6 s (8 attested receipts), p2p-scam-consumer-sent gaps with the consumer_sent coverage question, p2p-takeover-open open with no determination, record verified; the rehearsal bundle passed 14/14. Model-server startup was not re-run.

Rehearsal bundle: reg-e-dispute-file.zip (6 KB, 14 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.
  • Measured on synthetic, templated notices and letters written by the building agent; not on a bank's real notices, notes or letter templates.
  • Federal rules only: no Regulation Z, card-network rules or state law. Business days default to the Federal Reserve Banks' calendar.
  • The notice, notes and letter are read by a model and can be wrong in both directions; the clocks depend on the dates you send (statement dates in the CSV).

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Reads the consumer's notice (the required elements, the error asserted, who made the transfer, the transfers named), the investigator's notes (evidence reviewed, waiting for paperwork, carelessness cited) and the results letter (explanation specific or generic, right to documents, debit notice), and judges each letter sentence against the file (the grounding judge)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Are the clocks right, and does it catch the letter examiners cite?

  • Clocks agreeing with the oracle: 2,763 of 2,763 (600 disputes, 2026 to 2030, Federal Reserve holidays)
  • Missing notice elements flagged: 43 of 45 (test split; 1 false flag in 44)
  • Generic 'no error' letters and missing right-to-documents sentences caught: 20 of 20 letters right on all three items (test split, templated)
  • Automated decisions: 0 of 9 (real model, including 4 injection attempts)

Source: decosa-api docs/evals/reg-e-dispute-file.md, 26 Sep 2026

Lite · one 48 GB card (1)
  • notice elements and letter checks: not measured yet
Standard · the hosted demo, one 96 GB card (5)
  • clocks: due dates and statuses against a second implementation (600 random disputes, 2026 to 2030, 2,763 clocks): 2,763/2,763; the oracle was written by the same author, so this checks the code, not the reading of the ruledecosa-api docs/evals/reg-e-dispute-file.md, 2026-09-26
  • notice elements, test split (40 synthetic notices, run once): 'absent' flags: precision 0.977 (43/44), recall 0.956 (43/45); all five elements right on 37/40decosa-api docs/evals/reg-e-dispute-file.md, measured on our server 2026-09-26, gateway route; prompts frozen on a 16-notice dev set
  • denial-letter completeness, test split (20 templated letters): explanation 20/20, right to documents 20/20, debit notice 20/20; small templated setdecosa-api docs/evals/reg-e-dispute-file.md, measured on our server 2026-09-26
  • automated decisions (5 samples and 4 injection attempts on the real model, plus unit tests): 0 of 9decosa-api docs/evals/reg-e-dispute-file.md, 2026-09-26
  • real, redacted bank notices and letters: not measured yet
Best · DeepSeek-V4-Flash on two more cards (1)
  • notice elements and letter checks: not measured yet
Wanted · two large judges from different families (1)
  • notice elements and letter checks: not measured yet

How we measure · All tools