Skip to content
decosa

53 · Finance and insurance · Compliance and trust · live

SAR narrative desk

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)

  • Planted wrong numbers caught in reference narratives, first run (no model)1,598 of 1,600 (99.9%)test splitn = 1,6001,600 of 1,600 after a fix made on seeing the two misses; 0 of 1,920 correct numbers flagged.
  • Wrong numbers planted in the model's own sentences caught (v3)340 of 356 (95.5%)test splitn = 356Amounts 133/133, dates 165/174, counts 34/40.
  • Invented sentences held in investigator drafts18 of 20test splitn = 20The other two were flagged (partial), so all 20 were caught. Re-run 28 Sep 2026 on the same weights (direct route) after the judge was given the institution-conclusions and place-name context; the same day before that change: 19 held, 1 flagged; 26 Sep (gateway): 19 held, 1 flagged.
  • True reference sentences held / flagged0 of 172 held; 19 of 172 (11%) flaggedtest splitn = 17228 Sep 2026 re-run (direct route, same weights). Before the change, same day: 0 held, 27 flagged. 26 Sep (gateway): 1 held, 32 flagged.
  • Numbers the model wrote that the checker flagged (v3 drafts)0 of 533test splitn = 5334 of 299 sentences held, one of them a false hold by the judge.
  • Invented facts in randomly sampled checked sentences (manual read, v2)0 of 70; 2 of 70 (3%) minor unsupported details passeddev (tuned on)n = 70Read by the building agent only; v2 is partly development data.

Dataset

Synthetic case files from the vertical's own generator (fictional people, accounts and bank). Seeds 1-20 used for building, 1000+ for the eval: 200 cases for the numeric check, three draft sets of 15 runs (v1, v2, v3; code changed after v1 and v2, v3 run after the last checker change), 20 reference narratives with one invented sentence each.

Caveats

  • Everything is synthetic, written by the same agent that wrote the checker and the prompts: evidence the mechanisms work on this generator, not accuracy on real bank data.
  • Hallucinations were judged by the building agent reading the sentences; no second reader.
  • Code was changed after v1 and v2, so those sets are partly development data; two small fixes after v3 are not re-measured.
  • Comparator claims and percentages are not planted in the numeric check, and reference narratives are template prose.
  • The screen's thresholds were set on this generator; 5 of 40 clean cases were flagged for structuring.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
26 Sep 2026
Latency, this run
n/a
p50 over passed runs
25 s
Receipts
30
Model calls
n/a
Tokens
n/a
Cost per run
$0.017

Self-host verification

Verified on 26 Sep 2026: fresh clone, compose up, sample against local model servers

A fresh clone of a decosa-api pre-release build (not yet merged to main), the api image built from it with DECOSA_SAR_SYNTHETIC_ONLY=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The planted-errors check held exactly the three planted sentences (1.7 s); the structuring draft ran in 7.9 s with 51 of 51 numbers traced; the report verified, a changed status failed, and receipts were attested. Model-server startup itself not re-verified.

Rehearsal bundle: sar-narrative-desk.zip (4 KB, 15 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Synthetic only on the hosted demo. Every number here comes from our own synthetic generator; it has not been run on real case files.
  • A date is checked for membership in the cited rows, and tied to its amount when the sentence pairs them; a date moved onto another cited day with no amount beside it can pass. Counts can match another subset of the cited rows.
  • The screen's thresholds are ours. It flagged structuring on 5 of 40 clean cash-business cases.
  • It checks that what is written is traced to the case file, not that nothing is missing, and it sees nothing outside the case file (watch lists, other institutions, prior SARs).
  • The ledger must be CSV; PDF statements are not read.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Red-flag screen, citation check, numeric grounding, filing copy, workpaper and signed report (no model; CPU)decosa-api SAR desk (decosa_api/verticals/sar) with the numeric grounding block (decosa_api/verticals/numeric) and the grounding module (decosa_api/verticals/grounding)AGPL-3.0-or-later
  • Section drafting and the grounding judgeQwen3.8-27B (NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · numbers and citations only, no GPU (3)
  • Correct numbers flagged, reference narratives (200 held-out synthetic cases): 0 of 1,920docs/evals/sar-narrative-desk.md, part A, 26 Sep 2026
  • Planted wrong numbers caught (amounts, dates, counts), reference narratives: 1,600 of 1,600 (1,598 before a bug fix)docs/evals/sar-narrative-desk.md, part A
  • Planted wrong numbers caught in the model's own sentences: 340 of 356 (95.5%): amounts 133/133, dates 165/174, counts 34/40docs/evals/sar-narrative-desk.md, part D (v3)
Standard · one GPU for the model (hosted demo) (6)
  • Drafted sentences traced / flagged / held (v3: 15 synthetic cases, 299 sentences): 272 / 23 / 4 (1 of the 4 a false hold)docs/evals/sar-narrative-desk.md, part B
  • Numbers the model wrote that failed the check (v3): 0 of 533docs/evals/sar-narrative-desk.md, part B
  • Citation validity (v3): 0 uncited sentences; 1 unknown id of 931docs/evals/sar-narrative-desk.md, part B
  • Invented facts in investigator drafts held / flagged (20 drafts): 18 / 2 (all 20 caught); 0 of 172 true sentences held, 19 flagged (28 Sep re-run)docs/evals/sar-narrative-desk.md, part C
  • Typology coverage: planted typology found by the screen and named in the draft (12 cases, v3): 12/12 and 12/12; none named that the screen did not finddocs/evals/sar-narrative-desk.md, part B
  • Hallucinated facts in traced sentences (manual read, 70 sampled): 0 invented facts; 2 small unsupported details (a state, 'international')docs/evals/sar-narrative-desk.md, part E

How we measure · All tools