Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Reg E dispute file (67): eval

Run 26 Sep 2026 on the pre-release server (qwen3.8-27b through the model gateway, every call receipted), under the shared gateway's normal load. Everything is synthetic; results are in docs/evals/reg-e-dispute-file/.

What was measured

Check Set Result
Clock accuracy (code) 600 random disputes, 2026 to 2030, 2,763 clocks 2,763 / 2,763 due dates and statuses agree with a second implementation
Clock reading of the rule hand-counted cases in tests/test_rege.py all pass (weekends, Labor Day, Thanksgiving, Christmas, Saturday and Sunday holidays under K.8, closed days, 10/20 business days, 45/90 days, POS vs ATM, foreign, Puerto Rico, new account, written confirmation, Regulation T, $50 withholding, late notice)
Notice elements (model + quote check) test: 40 notices, 200 element judgments (name, account, why, date, amount) "absent" flags: precision 0.977 (43/44), recall 0.956 (43/45); all five right on 37/40 notices
Denial-letter completeness test: 20 letters x 3 items explanation 20/20, right to documents 20/20, debit notice 20/20; "cannot close" right on 20/20
Zero automated decisions 5 samples + 4 injection attempts (notice, notes, letter), plus the unit tests 0 of 9 outputs had a determination other than the one entered

Dev splits (16 notices, 7 letters) scored 100% on the same items; no prompt was changed after the dev run, and the test splits were run once.

Clocks

scripts/rege_eval.py clocks draws 600 disputes (seeded): notice dates from January 2026 to mid 2030, one or two disputed transfers of six types, 15% initiated abroad (some in Puerto Rico, which counts as a state), a quarter on accounts opened within 60 days, random provisional credits, notices, determinations, reports and corrections. It compares clocks.compute with an oracle written separately in the eval script: a day-by-day calendar built from the Federal Reserve's K.8 holiday table typed in by hand, and the rules as a flat list. This catches implementation bugs, not a misreading of the rule: the oracle was written by the same author. The reading of the rule rests on the quoted text in /rege/info and the hand-counted tests.

One reading to know about: the default calendar is the Federal Reserve Banks' (a holiday on a Saturday is not moved, so Friday 3 July 2026 is a business day). Reg E defines a business day by the institution's own opening (1005.2(d)); an institution that closes on other days lists them in closed_days.

Notice elements

scripts/rege_eval_data.py assembles notices from fragments in four channels (secure chat, call notes, web form, email) over 12 scenarios (card-not-present, ATM, P2P scam, account takeover, duplicate POS, ATM deposit, ACH wrong amount, bill pay, prepaid, remittance, stolen card, P2P wrong payee), leaving elements out at random. Traps: the agent's name when the consumer's is missing, the call date when the transfer date is missing, the account balance when the amount is missing. Dev = the first four scenarios, test = the other eight.

  • Test misses: one notice with no account identifier was read as identifying it ("my prepaid card", in a chat on the provider's own platform); one confirmation number was not counted as an account identifier; one email whose only line was "A card swipe at the grocery store" was read as giving a reason.
  • Type of error (reported apart, not in the headline): 92.5% on test. Our fragments for "why" often imply the type too, so the gold is loose.
  • Who made the transfer (it only chooses which coverage note the investigator sees): raw agreement 24/40. On re-reading the 16 disagreements after the run, 6 were our label errors (a stolen card labelled "unknown third party" although the prompt's own definition counts theft as access obtained from the consumer; ATM deposits the consumer plainly made), 5 have no fitting label (a merchant's second charge, a gym's ACH debit), and 5 are model errors (bill pay and remittance payments the consumer sent, read as "unclear"). We did not relabel the file; treat this field as a hint.
  • The CFPB complaint database was the planned input distribution. On 26 Sep 2026 its public search API and CSV export no longer return the narrative field (checked with has_narrative=true), so every notice here is synthetic.

Letters

Letters for one no-error file (card-not-present, provisional credit debited) are assembled from 3 specific explanations, 4 generic ones ("no error occurred and the transaction was processed correctly", "unable to substantiate your claim", "your claim is denied"...), one with no explanation, 3 right-to-documents sentences (including "If you would like copies of the records we used...", which no keyword rule catches) and 3 sentences that are not ("You can request a copy of your monthly statement"), and 3 debit notices (complete; no five business days; nothing). 20/20 on all three items is on a small, templated set: it shows the check does what it says on clear cases, not how it does on your letter templates. Run your own letters through the rehearsal bundle before relying on it.

Zero automated decisions

By construction: the determination is copied from the request, and no code path writes it. The tests check this with a fake model that answers "no_error / DENY / complete" in every call, and with a notice that tells the tool to deny. The eval repeats it on the real model: 5 samples and 4 injection attempts, 0 automated decisions. The grounding judge is told that the bank's own conclusion is the investigator's decision, not a fact to check, so the tool does not grade the outcome either.

Letter sentences: fewer false flags (28 Sep 2026, before/after)

A blind tester playing a disputes analyst called one letter flag pedantic: "The provisional credit of $412.18 will be reversed", in a letter sent the day the credit was debited, marked "contradicted". Changes to the letter check (the Reg E question given to the grounding judge; the grounding prompt itself is unchanged): tense alone is never a contradiction; the bank's standing offers (documents on request, checks honored for five business days, statements on request) are rights statements; and the letter's sign-off line (a short last line with no digit or verb, e.g. "Harbor Point Bank Disputes Department") is no_claim in code.

Both runs on 28 Sep 2026, same pre-release server, direct route to the same weights (127.0.0.1:8114), DECOSA_REGE_WORKERS=2.

Set Before After
Letters, test split (20): explanation / right to documents / debit notice right 20 / 19 / 20 20 / 20 / 20
Letters: "cannot close" right 20 / 20 20 / 20
Letters: letter sentences not found in the file (every fact in these letters is true, so all are false flags) 34 in 20 letters 7 in 7 letters
Probes (new, 16 letters with one probe sentence each): true restatements passed 6 / 6 5 / 6
Probes: changed amount, date, address, invented call, device, biometric, IP, "will not reverse" caught 10 / 10 10 / 10
Demo samples: the planted "since March 2024" sentence flagged; automated decisions yes; 0 yes; 0
  • The 20/20 in the table above was 19/20 for the right-to-documents item before the change on this route today (one letter; that item is read by a different prompt, which did not change): run-to-run variance, not an effect.
  • The false flags left: "The provisional credit will be reversed." (6 letters, still read as contradicted by the debit date) and one partial. On the probes, "We are reversing the provisional credit" passed before and was flagged after, one sentence either way; the probes were written after the rule was designed and before either run, and nothing was tuned on them. A second, concurrent run of the probes (overloaded, 3 of 16 runs lost to 429s) once passed "We will not reverse the provisional credit": treat 10/10 as 16 sentences, not a rate.
  • Results: letters-test-28sep-{before,after}-results.json, probes-28sep-{before,after}-results.json, decisions-28sep-{before,after}-results.json (scripts/rege_eval.py letters|probes|decisions --tag=...).

Latency and cost (measured under load)

The shared gateway was busy on 26 Sep 2026, so timings swing with its queue. Ten runs of the four demo files with a letter took 7.6 s to 96.8 s (median 37.8 s) and 8 to 14 receipted calls; the open file with no letter (2 calls) took 6.4 s in the smoke run. On the self-host sandbox (direct route to the local model) the card-not-present file took 18.6 s and the P2P scam file 10.8 s. The card-not-present file used 13,806 tokens (about $0.006 at the gateway list price), the P2P scam file 24,755 tokens (about $0.010), the smoke run 2,409 tokens (about $0.0016).

Expected properties of the sample run (for the rehearsal kit)

  1. cnp-generic-denial: provisional-credit notice missed (due 2026-09-04, sent 2026-09-09); notice:T-0512 is late_notice (due 2026-08-01); determination due 2026-11-18 (90 days, debit card online).
  2. cnp-generic-denial: checklist explanation, right_to_documents, documents_relied_on and debit_notice are gap; file_status.cannot_close is present; determination is no_error by Jordan Okafor, actor: "human".
  3. p2p-scam-consumer-sent: determination missed (due 2026-09-14, entered 2026-09-17, no provisional credit); coverage consumer_sent cites 12 CFR 1005.2(m); no_delay is gap; the "since March 2024" sentence is not in the file.
  4. atm-short-cash-new-account: new-account trigger, 20 business days (due 2026-09-22); correction missed (due 2026-09-11).
  5. p2p-takeover-open: P-3 and P-4 matched from the notice; provisional credit open (due 2026-09-29); status open.
  6. Every signed file verifies at POST /record/verify and fails once its determination is changed.

Verdict

The clocks are the reliable part and the easiest to trust: plain code, cited, cross-checked. The notice and letter checks are good on clear cases and have not been tried on real bank templates. A buyer would pay for the clocks plus the letter check plus the signed file, if it plugs into their case system (a JSON API) and runs on their hardware. What's missing: integration with a case system, a trial on a bank's real (redacted) letters, and state-law and card-network rules.