Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: bank-change-check (Check a vendor bank-detail change before you pay)

29 Sep 2026, build-insurance-finance-opus. Runner: scripts/insfin_eval/bankchange_eval.py. Data and raw results: docs/evals/bank-change-check/.

Data

  • 64 synthetic emails (raw RFC 5322 with headers, some HTML parts and attachments) and a 26-vendor file, written blind by a separate author (a sub-agent that never saw the tool), with a generator (make_items.py) and a label checker (check_labels.py: every evidence quote is in the decoded message). Fictional company, vendors on reserved .example/.test domains, 555-01xx numbers, routing numbers starting 99, IBANs with TEST bank codes.
  • Classes: 30 fraud (7 obvious, 9 moderate, 7 subtle, 7 from the vendor's real, authenticated mailbox), 26 genuine (bank mergers, new controllers, a new direct line, German/Spanish/Portuguese/French), 8 with no bank change (2 look-alike phishing for a login). 6 emails are not in English.
  • Split: stratified by class and subtlety, every third email by id held out: dev 45, test 19. Code and prompt changes were made on dev only; the test set was run once, on the gateway route, after they were frozen.
  • Labels: 16 sign types with evidence. The author's conventions differ from the tool in two places: auth_fail includes dkim=none/dmarc=none (the tool counts only fail/softfail/errors), and a Reply-To on another of the vendor's own domains is a reply_to_mismatch (the tool doesn't count it).

What was measured

  • Change detection: is there a bank-detail change request at all?
  • Flag at a fixed false-flag rate: the weighted sign score at or above 3 ("warning signs"). The threshold was chosen on dev as the lowest with at most 10% false flags on genuine changes, then fixed.
  • Per sign: recall and precision against the labels.
  • M30 baseline (page 87 §4): the content signs (urgency, secrecy, pressure from a boss, "don't call", redirecting an approved payment) read by keyword rules (no model) against the open model with verbatim quotes.
  • Rails: no verdict calls an email safe; the call-back always names the number on file.

Results

dev 45 (direct route) test 19 (gateway, held out, run once)
Change requests detected (all emails) 45 / 45 19 / 19
Fraud flagged at score ≥ 3 21 / 21 9 / 9 (Wilson 95% CI 0.70-1.00)
Genuine changes flagged (false flags) 0 / 18 0 / 8 (CI 0.00-0.32)
Fraud flagged, from the vendor's real mailbox 5 / 5 2 / 2
Verdicts calling an email safe 0 0
Call-back names the number on file (matched vendor) 45 / 45 18 / 18
Cost per email, gateway list price $0.00053 mean (one model call)
Time per email, shared gateway under load 10.4 s p50 8.1 s p50, 18.9 s p95

Signs found in code (test): look-alike domain 5/5, display name over an unknown address 7/8, sender not on file 7/8, bank in another country 2/2, account holder not the vendor 4/5, Reply-To 1/1, free-mail sender 2/2, sender checks failed 2/4 (the two misses are none results the tool doesn't count), new number 1/1; no false positives on any code sign.

M30: rules vs the open model on the content signs (dev + test, 64 emails)

keyword rules (no model) open model, quotes checked
Recall 23 / 47 = 0.49 43 / 47 = 0.91
Precision 23 / 26 = 0.88 43 / 52 = 0.83
Fraud flagged, test (score ≥ 3) 9 / 9 9 / 9
Fraud from a real mailbox flagged, dev 3 / 5 5 / 5
  • The flag itself leans on the code signs: rules alone flagged every test fraud, but on dev they missed 2 of 5 real-mailbox frauds, where only the content (a new bank abroad aside) gives it away. The model's quoted reading is the default; the rules run when the model's answer can't be used, and as a no-GPU mode (mode: "rules").
  • M30 verdict: worth training later, for volume and privacy (reading every AP email on the customer's own box on CPU), not for accuracy on this set. It needs a much larger labelled set (thousands of emails in the language pack's languages, written blind) than the 64 here; 64 is an eval set, not a training set. Not trained (no GPU used).

Checkable properties of the sample runs (rehearsal)

  1. lookalike-urgent: change requested; signs include lookalike_domain, reply_to_mismatch, auth_fail, bank_country_mismatch, account_name_mismatch, avoid_callback; the call-back names +1-312-555-0142 and lists +1-312-555-0199 under "don't use".
  2. genuine-bank-acquisition: change requested, no warning sign, and the verdict still says to hold and call.
  3. no-change-phishing: no change requested, warning signs about the sender.
  4. Every record verifies at POST /record/verify; POST /bank-change/signoff refuses a number that isn't on file and a caller who is also the approver.

Caveats

  • Synthetic emails: real BEC mail is messier (forwarded chains, images, lookalike Unicode, thread hijacks months long). Small test set (19 emails, 9 frauds): the confidence intervals are wide.
  • One author wrote all 64 emails; the tool's author tuned code cues on the dev 45 (a hyphen-joined vendor-name rule, a stricter phone pattern after IBAN digits were read as phone numbers, and a narrower "redirect" definition).
  • A compromised real mailbox that writes calmly, keeps the same bank country and the vendor's name passes every check in the email; that is why the call-back is always required and the tool never says safe.
  • Not measured: domain age or ownership (no look-ups by design), account-validation services, real inboxes.