Eval: sanctions alert disposition record (58)
Run on 26 Sep 2026 on our server, against the list snapshot 3a3407c2… (OFAC SDN published 2026-09-23, OFAC Consolidated
2026-09-14, EU generated 2026-09-22, UK Sanctions List generated 2026-09-21; 32,452 entries). Runner:
scripts/eval_sanctions.py. Result files are in docs/evals/sanctions-disposition-record/: code.json,
model.json (with the 6 outage rows re-run; the first run is in model-first-run.json) and checker.json.
Data
The pairs are synthetic customers set against real list entries, labelled by construction
(decosa_api/verticals/sanctions/synth.py).
- A same-party pair is a fictional customer record built from the entry's own public data, with the variants real
records show. There are 13 cases:
- exact;
- a romanisation variant, with the name order sometimes swapped;
- a strong alias the entry was then stripped of ("held-out alias", so the matcher cannot find it verbatim);
- the Cyrillic name from the UK list, with that alias stripped;
- a weak alias;
- a one-letter typo;
- a second nationality;
- day and month swapped;
- a year-only listing;
- the listed passport;
- a renewed passport (a different number, same country);
- name only;
- an entity with its registration number or listed city;
- a vessel with its IMO number.
- A different-party pair is a fictional person or company who shares the name but not the identity. There are 11
cases:
- the same name born 5 to 30 years apart;
- the same year but a different date;
- a variant name with a different date of birth;
- a different national ID or registration number from the same country;
- a person or company against an entity or vessel entry;
- the same vessel name with a different IMO number;
- only the first name shared;
- name only;
- one year off;
- the same name, year and nationality against a year-only listing.
The split is by list-entry uid (a hash): dev and test never share an entry. The rules and the matcher were written and adjusted on dev; the test split was run once at the end. The generator's romanisation variants come from a list of real spellings. That list was written separately from the matcher's folding rules, but by the same agent, so it tests the same families of variation.
A. Match analysis and proposal (code, no model)
| dev | test | |
|---|---|---|
| Pairs | 1,990 | 1,967 |
| Same-party pairs proposed as false positive (the risky direction) | 0 of 1,102 | 0 of 1,094 |
| False-positive clearance precision (different party / all false-positive proposals) | 634 / 634 (100%) | 635 / 635 (100%) |
| Same-party pairs proposed as true match | 870 (78.9%) | 863 (78.9%) |
| Same-party pairs proposed as needs more information | 232 | 231 |
| True-match precision | 92.3% | 93.8% |
| Different-party pairs cleared (proposed false positive) | 71.4% | 72.7% |
| Proposal in the case's acceptable set | 99.95% | 99.80% |
| Per-field status accuracy (labelled fields) | 99.91% of 2,241 | 99.77% of 2,199 |
| Time per pair | 0.35 ms | 0.40 ms |
How to read the table:
No same-party pair was proposed as a false positive. This follows from the design: a false positive is proposed only on a strong disqualifier, meaning:
- a different full date of birth, or one more than a year off;
- a different national ID, registration or tax number from the same country;
- a different IMO number;
- a different kind of party;
- or, with nothing else in agreement, a name under 0.70. A different passport is not a strong disqualifier, because passports are renewed. Neither are nationality, address or gender.
The number is only as good as the generator. A true match whose customer record carries a wrong date of birth, or a wrong ID, would be proposed as a false positive. That case is not in the set, and the tool cannot catch it. See "Not measured".
True-match recall (79%) is low on purpose. Name-only, day/month-swapped and renewed-passport pairs cannot be confirmed from the data given, so they are proposed as needs more information, with what to ask for.
True-match precision is 94%. The misses are the "same name, year and nationality against a year-only listing" case. By OFAC FAQ 5, step 3, that case is a likely match, and the proposal says so. That is a safe over-escalation.
Different parties not cleared (27%) are the first-name-only, name-only, one-year-off and "entity elsewhere" cases. They are held as needs more information, which is the safe direction.
Test misses (4):
- Two "same year, different date" pairs where the entry also lists year-only dates that the customer's date fits or comes within a year of. They were held as needs more information, not cleared.
- Two year-only EU listings whose primary name is only in Hangul or Persian script. Those names are not transliterated, so they are "not compared" and held.
B. Model: the typed second reading and the rationale (hosted gateway, receipted)
There were 81 reviews: 3 per case over 27 cases, from the test split, through POST /sanctions/review on the branch
server. Qwen3.8-27B ran through the shared model gateway.
| Reviews with both model calls answered | 81 of 81 (6 first lost both calls when the gateway went down at about 20:20 UTC; the code analysis stood and the review said so; those 6 were re-run after the fix, as the coordinator asked, and the outage is not counted) |
| Receipts | 162, all gateway-signed |
| Typed reading agrees with the code proposal | 71 of 81 (88%) |
| Typed reading would clear a same-party pair (answer false_positive) | 0 of 45 |
| Typed reading on different-party pairs | 22 false positive, 11 needs more info, 3 true match (all 3 the same-year-and-nationality case against a year-only listing; the code says true match on 2 of them) |
| Rationale sentences | 278 |
| Rationale sentences passing the code check | 278 of 278 (100%) |
| Latency per review (both calls in parallel) | first 75: p50 8.4 s, p90 12.9 s, max 15.4 s under shared load; re-run 6: 6.9-11.5 s |
The 10 disagreements:
- 3 first-name-only cases, where the code clears on a low name and the model asks for more;
- 3 entity-city cases, where the code says true match on name, city and country, and the model asks for more;
- 1 same-year date and 1 registration conflict, where the model asks for more;
- 1 year-only listing, where the model says true match and the code says needs more;
- 1 weak-alias match (in the re-run), where the code says true match and the model asks for more.
The model was never less cautious than the code about clearing. The console shows a disagreement to the analyst. It never changes the proposal.
Manual read. The building agent read 30 randomly sampled rationale sentences against their field tables. It found 0 invented facts, 0 wrong values and 0 sentences calling a conflict a match. One sentence mixes the strength label with the score ("a medium match with a similarity score of 1.00"). The words come from the table and are accurate, but they read oddly.
C. The rationale checker against planted errors (code, no model)
Correct sentences were built from each field's own note and then mutated, on 300 test pairs. A mutation added a wrong year, a country the fields don't name, an invented passport number, flipped polarity ("matches" on a conflict), a cite to a field that doesn't exist, or no cite.
| Correct sentences passed | 531 of 531 (100%; 516 of 531 before a fix: "name order differs" read as a negative claim, fixed on seeing this set) |
| Planted errors held | 2,776 of 2,776 (100%): wrong year 121/121, wrong country 531/531, invented ID 531/531, flipped polarity 531/531, missing field 531/531, no cite 531/531 |
The mutations are templates written by the same agent that wrote the checker. They show that the checks fire. They do not show what share of real model errors would be caught, and the model made no error the checker caught in part B.
Rehearsal: checkable properties of the sample run
For rehearsal/sanctions-disposition-record (dob-mismatch sample: fictional Ali Zeaiter, born 2 Nov 1977, Lebanese,
against ofac-sdn:17037, born 24 Feb 1977, Lebanon):
analysis.proposal.decisionisfalse_positiveandruleisR-DISQ.- F.dob is
conflict(same year, different date) and F.nationality ismatch. - At least one rationale sentence is
grounded, and both model calls have signed receipts. - Deciding
true_matchwithoutoverride_reasonanswers 422. - The sealed record verifies (
POST /sanctions/verify→ok: true), and withstatement.decisionchanged it does not. - The audit sample of that one record has
manifest.size1 and a signature.
Not measured
- Real alerts. Everything here is synthetic customers against real list entries, built by the agent that wrote the rules. It shows that the rules do what they say on list-derived data. It is not accuracy on a bank's alert queue, where the customer data is messy and sometimes wrong.
- A true match whose customer data is wrong (a mistyped date of birth or ID). The tool proposes a false positive on a strong disqualifier and cannot tell a typo from a different person, except for a swapped day and month, or the same day and month one year apart. The analyst has to check the source document.
- Arabic, Persian, Chinese, Korean and other non-Latin, non-Cyrillic names. They are not transliterated; a name in those scripts is "not compared" and held.
- The screening step (finding the alerts), ownership and control (the 50 Percent Rule), licences, and list changes after the snapshot.
- A second reader for the manual read.