53 · Finance and insurance · Compliance and trust · live
SAR narrative desk
Eval results
Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)
- Planted wrong numbers caught in reference narratives, first run (no model)1,598 of 1,600 (99.9%)test splitn = 1,6001,600 of 1,600 after a fix made on seeing the two misses; 0 of 1,920 correct numbers flagged.
- Wrong numbers planted in the model's own sentences caught (v3)340 of 356 (95.5%)test splitn = 356Amounts 133/133, dates 165/174, counts 34/40.
- Invented sentences held in investigator drafts18 of 20test splitn = 20The other two were flagged (partial), so all 20 were caught. Re-run 28 Sep 2026 on the same weights (direct route) after the judge was given the institution-conclusions and place-name context; the same day before that change: 19 held, 1 flagged; 26 Sep (gateway): 19 held, 1 flagged.
- True reference sentences held / flagged0 of 172 held; 19 of 172 (11%) flaggedtest splitn = 17228 Sep 2026 re-run (direct route, same weights). Before the change, same day: 0 held, 27 flagged. 26 Sep (gateway): 1 held, 32 flagged.
- Numbers the model wrote that the checker flagged (v3 drafts)0 of 533test splitn = 5334 of 299 sentences held, one of them a false hold by the judge.
- Invented facts in randomly sampled checked sentences (manual read, v2)0 of 70; 2 of 70 (3%) minor unsupported details passeddev (tuned on)n = 70Read by the building agent only; v2 is partly development data.
Dataset
Synthetic case files from the vertical's own generator (fictional people, accounts and bank). Seeds 1-20 used for building, 1000+ for the eval: 200 cases for the numeric check, three draft sets of 15 runs (v1, v2, v3; code changed after v1 and v2, v3 run after the last checker change), 20 reference narratives with one invented sentence each.
Caveats
- Everything is synthetic, written by the same agent that wrote the checker and the prompts: evidence the mechanisms work on this generator, not accuracy on real bank data.
- Hallucinations were judged by the building agent reading the sentences; no second reader.
- Code was changed after v1 and v2, so those sets are partly development data; two small fixes after v3 are not re-measured.
- Comparator claims and percentages are not planted in the numeric check, and reference narratives are template prose.
- The screen's thresholds were set on this generator; 5 of 40 clean cases were flagged for structuring.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 26 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 25 s
- Receipts
- 30
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.017
Self-host verification
Verified on 26 Sep 2026: fresh clone, compose up, sample against local model servers
A fresh clone of a decosa-api pre-release build (not yet merged to main), the api image built from it with DECOSA_SAR_SYNTHETIC_ONLY=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The planted-errors check held exactly the three planted sentences (1.7 s); the structuring draft ran in 7.9 s with 51 of 51 numbers traced; the report verified, a changed status failed, and receipts were attested. Model-server startup itself not re-verified.
Rehearsal bundle: sar-narrative-desk.zip (4 KB, 15 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Synthetic only on the hosted demo. Every number here comes from our own synthetic generator; it has not been run on real case files.
- A date is checked for membership in the cited rows, and tied to its amount when the sentence pairs them; a date moved onto another cited day with no amount beside it can pass. Counts can match another subset of the cited rows.
- The screen's thresholds are ours. It flagged structuring on 5 of 40 clean cash-business cases.
- It checks that what is written is traced to the case file, not that nothing is missing, and it sees nothing outside the case file (watch lists, other institutions, prior SARs).
- The ledger must be CSV; PDF statements are not read.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Red-flag screen, citation check, numeric grounding, filing copy, workpaper and signed report (no model; CPU)decosa-api SAR desk (decosa_api/verticals/sar) with the numeric grounding block (decosa_api/verticals/numeric) and the grounding module (decosa_api/verticals/grounding)AGPL-3.0-or-later
- Section drafting and the grounding judgeQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · numbers and citations only, no GPU (3)
- Correct numbers flagged, reference narratives (200 held-out synthetic cases): 0 of 1,920docs/evals/sar-narrative-desk.md, part A, 26 Sep 2026
- Planted wrong numbers caught (amounts, dates, counts), reference narratives: 1,600 of 1,600 (1,598 before a bug fix)docs/evals/sar-narrative-desk.md, part A
- Planted wrong numbers caught in the model's own sentences: 340 of 356 (95.5%): amounts 133/133, dates 165/174, counts 34/40docs/evals/sar-narrative-desk.md, part D (v3)
Standard · one GPU for the model (hosted demo) (6)
- Drafted sentences traced / flagged / held (v3: 15 synthetic cases, 299 sentences): 272 / 23 / 4 (1 of the 4 a false hold)docs/evals/sar-narrative-desk.md, part B
- Numbers the model wrote that failed the check (v3): 0 of 533docs/evals/sar-narrative-desk.md, part B
- Citation validity (v3): 0 uncited sentences; 1 unknown id of 931docs/evals/sar-narrative-desk.md, part B
- Invented facts in investigator drafts held / flagged (20 drafts): 18 / 2 (all 20 caught); 0 of 172 true sentences held, 19 flagged (28 Sep re-run)docs/evals/sar-narrative-desk.md, part C
- Typology coverage: planted typology found by the screen and named in the draft (12 cases, v3): 12/12 and 12/12; none named that the screen did not finddocs/evals/sar-narrative-desk.md, part B
- Hallucinated facts in traced sentences (manual read, 70 sampled): 0 invented facts; 2 small unsupported details (a state, 'international')docs/evals/sar-narrative-desk.md, part E