Eval: bank-change-check (Check a vendor bank-detail change before you pay)
29 Sep 2026, build-insurance-finance-opus. Runner: scripts/insfin_eval/bankchange_eval.py. Data and raw results:
docs/evals/bank-change-check/.
Data
- 64 synthetic emails (raw RFC 5322 with headers, some HTML parts and attachments) and a 26-vendor file, written
blind by a separate author (a sub-agent that never saw the tool), with a generator (
make_items.py) and a label checker (check_labels.py: every evidence quote is in the decoded message). Fictional company, vendors on reserved.example/.testdomains, 555-01xx numbers, routing numbers starting 99, IBANs with TEST bank codes. - Classes: 30 fraud (7 obvious, 9 moderate, 7 subtle, 7 from the vendor's real, authenticated mailbox), 26 genuine (bank mergers, new controllers, a new direct line, German/Spanish/Portuguese/French), 8 with no bank change (2 look-alike phishing for a login). 6 emails are not in English.
- Split: stratified by class and subtlety, every third email by id held out: dev 45, test 19. Code and prompt changes were made on dev only; the test set was run once, on the gateway route, after they were frozen.
- Labels: 16 sign types with evidence. The author's conventions differ from the tool in two places:
auth_failincludesdkim=none/dmarc=none(the tool counts only fail/softfail/errors), and a Reply-To on another of the vendor's own domains is areply_to_mismatch(the tool doesn't count it).
What was measured
- Change detection: is there a bank-detail change request at all?
- Flag at a fixed false-flag rate: the weighted sign score at or above 3 ("warning signs"). The threshold was chosen on dev as the lowest with at most 10% false flags on genuine changes, then fixed.
- Per sign: recall and precision against the labels.
- M30 baseline (page 87 §4): the content signs (urgency, secrecy, pressure from a boss, "don't call", redirecting an approved payment) read by keyword rules (no model) against the open model with verbatim quotes.
- Rails: no verdict calls an email safe; the call-back always names the number on file.
Results
| dev 45 (direct route) | test 19 (gateway, held out, run once) | |
|---|---|---|
| Change requests detected (all emails) | 45 / 45 | 19 / 19 |
| Fraud flagged at score ≥ 3 | 21 / 21 | 9 / 9 (Wilson 95% CI 0.70-1.00) |
| Genuine changes flagged (false flags) | 0 / 18 | 0 / 8 (CI 0.00-0.32) |
| Fraud flagged, from the vendor's real mailbox | 5 / 5 | 2 / 2 |
| Verdicts calling an email safe | 0 | 0 |
| Call-back names the number on file (matched vendor) | 45 / 45 | 18 / 18 |
| Cost per email, gateway list price | $0.00053 mean (one model call) | |
| Time per email, shared gateway under load | 10.4 s p50 | 8.1 s p50, 18.9 s p95 |
Signs found in code (test): look-alike domain 5/5, display name over an unknown address 7/8, sender not on file 7/8, bank
in another country 2/2, account holder not the vendor 4/5, Reply-To 1/1, free-mail sender 2/2, sender checks failed 2/4
(the two misses are none results the tool doesn't count), new number 1/1; no false positives on any code sign.
M30: rules vs the open model on the content signs (dev + test, 64 emails)
| keyword rules (no model) | open model, quotes checked | |
|---|---|---|
| Recall | 23 / 47 = 0.49 | 43 / 47 = 0.91 |
| Precision | 23 / 26 = 0.88 | 43 / 52 = 0.83 |
| Fraud flagged, test (score ≥ 3) | 9 / 9 | 9 / 9 |
| Fraud from a real mailbox flagged, dev | 3 / 5 | 5 / 5 |
- The flag itself leans on the code signs: rules alone flagged every test fraud, but on dev they missed 2 of 5
real-mailbox frauds, where only the content (a new bank abroad aside) gives it away. The model's quoted reading is the
default; the rules run when the model's answer can't be used, and as a no-GPU mode (
mode: "rules"). - M30 verdict: worth training later, for volume and privacy (reading every AP email on the customer's own box on CPU), not for accuracy on this set. It needs a much larger labelled set (thousands of emails in the language pack's languages, written blind) than the 64 here; 64 is an eval set, not a training set. Not trained (no GPU used).
Checkable properties of the sample runs (rehearsal)
lookalike-urgent: change requested; signs includelookalike_domain,reply_to_mismatch,auth_fail,bank_country_mismatch,account_name_mismatch,avoid_callback; the call-back names +1-312-555-0142 and lists +1-312-555-0199 under "don't use".genuine-bank-acquisition: change requested, no warning sign, and the verdict still says to hold and call.no-change-phishing: no change requested, warning signs about the sender.- Every record verifies at
POST /record/verify;POST /bank-change/signoffrefuses a number that isn't on file and a caller who is also the approver.
Caveats
- Synthetic emails: real BEC mail is messier (forwarded chains, images, lookalike Unicode, thread hijacks months long). Small test set (19 emails, 9 frauds): the confidence intervals are wide.
- One author wrote all 64 emails; the tool's author tuned code cues on the dev 45 (a hyphen-joined vendor-name rule, a stricter phone pattern after IBAN digits were read as phone numbers, and a narrower "redirect" definition).
- A compromised real mailbox that writes calmly, keeps the same bank country and the vendor's name passes every check in the email; that is why the call-back is always required and the tool never says safe.
- Not measured: domain age or ownership (no look-ups by design), account-validation services, real inboxes.