58 · Finance and insurance · Compliance and trust · live
Sanctions alert disposition record
Eval results
Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)
- Same-party pairs proposed as false positive (the risky direction)0 of 1,094test splitn = 1,094Synthetic customers built from real list entries; dev 0 of 1,102.
- False-positive clearance precision635 / 635 (100%)test splitn = 635Share of false-positive proposals that were different parties.
- Same-party pairs proposed as true match863 (78.9%)test splitn = 1,094The other 231 were held as needs more information (name only, day/month swapped, renewed passport).
- Per-field status accuracy (labelled fields)99.77% of 2,199test splitn = 2,199
- Model reading agrees with the code proposal71 of 81 (88%)test splitn = 816 reviews that first lost both calls to a gateway outage were re-run after the fix; the outage is not counted.
- Rationale sentences passing the code check278 of 278 (100%)test splitn = 278Manual read of 30: 0 invented facts.
- Planted rationale errors held by the checker2,776 of 2,776 (100%)test splitn = 2,776531 of 531 correct sentences passed after one fix made on seeing this set (516 before).
Dataset
Synthetic customers against real OFAC, EU and UK list entries (snapshot of 26 Sep 2026), labelled by construction: 13 same-party and 11 different-party cases, split by list-entry uid into dev (1,990 pairs) and test (1,967 pairs); 81 test reviews through the hosted model (6 re-run after a gateway outage).
Caveats
- The customers, the variants and the rules were all written by the same agent: this shows the rules do what they say on list-derived data, not accuracy on a real alert queue.
- A true match with wrong customer data (a mistyped date of birth or ID) is not in the set and would be proposed as a false positive.
- The checker's planted errors are templates; the model made no error the checker caught, so its recall on real model errors is unknown.
- Names in scripts other than Latin and Cyrillic are not compared.
- The manual read had one reader, the building agent.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 26 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 8.4 s
- Receipts
- 2
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- <$0.001
Self-host verification
Verified on 26 Sep 2026: fresh clone, compose up, sample against local model servers
A fresh clone of a decosa-api pre-release build (not yet merged to main), the api image built from docker/api/Dockerfile with DECOSA_SANCTIONS_SYNTHETIC_ONLY=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The dob-mismatch review took 0.86 s (false_positive by R-DISQ, both receipts attested, 3 grounded sentences); the record verified and an edited one failed; an override without a reason and a risky clear without a second reviewer both got 422. The documented list refresh ran in the container in 41 s. Model-server startup itself not re-verified.
Rehearsal bundle: sanctions-disposition-record.zip (2 KB, 13 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Synthetic customers only on the hosted demo, and every number here comes from our own generator on real list entries; not yet run on a real alert queue.
- A true match whose customer data is wrong (a mistyped date of birth or ID) can be proposed as a false positive; the analyst checks the source document.
- Names in scripts other than Latin and Cyrillic are not compared, so those alerts are held.
- It reviews one alert against one list entry. It does not screen, and it does not cover ownership (the 50 Percent Rule), licences or list changes after the snapshot.
- Records are not stored: you keep them for 10 years.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- List snapshot, field comparison, proposal rules, rationale check, signed record, audit sample (no model; CPU)decosa-api sanctions desk (decosa_api/verticals/sanctions), with the typed-judgment core (vertical 24) and the session hash chain (record, vertical 07)AGPL-3.0-or-later
- Independent typed reading and the rationaleQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · the comparison and the proposal, no GPU (3)
- Same-party pairs proposed as false positive (held-out test split): 0 of 1,094docs/evals/sanctions-disposition-record.md, part A, 26 Sep 2026
- False-positive clearance precision (test): 635 of 635 (100%)docs/evals/sanctions-disposition-record.md, part A
- Field status accuracy on labelled fields (test): 99.77% of 2,199docs/evals/sanctions-disposition-record.md, part A
Standard · one GPU for the model (hosted demo) (4)
- Typed reading agrees with the code proposal (81 test reviews): 71 of 81 (88%); it never proposed clearing a same-party pairdocs/evals/sanctions-disposition-record.md, part B
- Rationale sentences passing the code check: 278 of 278docs/evals/sanctions-disposition-record.md, part B
- Planted rationale errors held by the checker (wrong year, country, ID, polarity, citation): 2,776 of 2,776; 531 of 531 correct sentences passeddocs/evals/sanctions-disposition-record.md, part C
- Invented facts in a manual read of 30 sampled rationale sentences: 0docs/evals/sanctions-disposition-record.md, part B