Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: SAR narrative desk (53) and the numeric grounding block

Run on our server, 26 Sep 2026, against the pre-release server (127.0.0.1:8453) on the hosted route: Qwen3.8-27B NVFP4 through the shared model gateway, temperature 0, thinking off. Script: scripts/eval_sar.py; raw results in docs/evals/sar-narrative-desk/. Everything is synthetic: case files come from decosa_api/verticals/sar/synth.py (fictional people, accounts and bank), written by the same agent that wrote the checker and the prompts. Treat the numbers as evidence that the mechanisms work on this generator, not as accuracy on real bank data.

What was measured, and how honestly

  • Seeds. 1-20 were used while building the checker and the screen (development). 101-106 are the demo samples. The eval uses 1000+ only. Three draft sets were run with the model: v1 (seeds 1000-1002), v2 (2000-2002) and v3 (3000-3002). I changed the code after v1 and after v2 (see "Fixes made after looking at eval runs"), so v1 and v2 are partly development data. v3 was run after the last checker change that affects it and is the set reported in the stack page; the two small fixes made after v3 are listed and not re-measured.
  • Who judged the hallucinations. I (the building agent) read the sentences myself: all 31 flagged and 4 held sentences of v2, plus 70 randomly sampled "checked" sentences of v2, against the generated ledgers. No second reader.

A. Numeric check on reference narratives with planted errors (no model)

200 held-out cases (40 per typology, seeds 1000-1039). Each case has a reference narrative with correct numbers and citations; each exact number in it (amounts, dates, counts) is changed in turn to a plausible wrong value (a total off by 3-20% or by a digit swap or round step, a date moved 1-9 days, a count off by 1-3).

result
Numbers in the reference narratives flagged (false alarms) 0 of 1,920
Planted wrong numbers caught, first run 1,598 of 1,600 (99.9%): amounts 760/760, dates 400/400, counts 438/440
After the fix below 1,600 of 1,600

The two misses were "1 branches" and "1 cities": the checker counted distinct locations inside a subset already fixed on location, which is always 1. Fixed (skip that subset). Limits: comparator claims ("more than", "within", "each under") and percentages are not planted here (a changed comparator is often still true), and the sentences are template prose, easier than model prose. Part D covers model prose.

B. Model drafts (the full pipeline)

15 runs per set: four typologies x 3 seeds in draft mode, plus 3 clean alerts in no-file mode.

v3 (seeds 3000-3002, final code) result
Sentences 299 (272 checked, 23 flagged, 4 held)
Numbers the model wrote 533; the checker flagged 0
Citation validity 0 uncited sentences; 1 unknown id in 931 cited ids (the model wrote K.institution for A.institution)
Planted typology found by the screen / named in the narrative 12/12 / 12/12; no typology named that the screen did not find
Clean alerts with no typology flagged 2 of 3 (one diner case trips the structuring screen; its rationale then addresses it)
Time per draft, gateway shared with other workloads median 24.7 s (11.7-61 s); no-file rationale about 12 s
Model calls per draft median 30 (8 sections + one judge call per sentence)

The 4 held v3 sentences: one bad citation (K.institution), one true contradiction the judge caught (the model said a May 9 debit was a wire to the payee; it was an electric-utility ACH), one unsupported judgement ("in close proximity to ... income deposits"), and one false hold by the judge ("four withdrawals within three days of out-of-state deposits", which the cited screen fact states; the judge saw only the withdrawal rows). The fact now carries the deposit rows too (fix after v3, not re-measured).

Earlier sets, for the record: v1 had 18 held sentences. 13 were number flags and all 13 were checker false alarms (8 on "$10,000" threshold phrasings such as "together exceeded $10,000" and "neither exceeding $10,000", 3 on "SSA TREAS 310", 1 on "within one to two days", 1 on "the remaining two incoming wires"); the first 12 are fixed, and re-scoring v1's saved sentences with the final checker leaves only the last one. The other 5: 2 bad citations (K.institution), 2 true contradictions caught by the judge, 1 judge false hold. v2 (after those fixes): 534 numbers, 0 flagged; 4 held (1 bad cite, 3 judge false holds on spans and "within three days").

C. Invented facts in investigator drafts (check mode)

20 reference narratives (typologies x seeds 1000-1003), each with one invented sentence that cites real ids ("the customer said the cash came from the sale of a car [N3]", "holds an account at another bank in Miami", "sent to a sanctioned country", "works for a money services business" ...).

result
Invented sentences held 19 of 20; the other was flagged (partial)
True reference sentences held (false holds) 1 of 172
True reference sentences flagged 32 of 172 (19%): mostly details the reference cited loosely (an account type or a city without the field id). The generator's citations were tightened afterwards.

D. Wrong numbers planted in the model's own sentences

Every exact number in the checked and flagged sentences of a draft set, changed one at a time as in part A, checked offline against that case.

v2 (before the date-pairing check) v3 (final)
Caught 329 of 346 (95.1%) 340 of 356 (95.5%)
Amounts 132/132 133/133
Dates 167/178 165/174
Counts 27/33 34/40
Durations, percentages 3/3, - 6/7, 2/2

Why dates and counts slip: a date is checked for membership in the cited rows, so a date moved onto another cited day ("same-day splits on May 16 and May 24" becoming "May 17") passes unless an amount is tied to it ("$22,785.00 on February 5" must come from a row dated that day; added after v2). A count can match another subset of the cited rows ("12 credits" when there are 14 credits but 12 ACH credits). Both are listed as known limits.

E. Hallucinated facts (manual read, v2)

  • Of 70 randomly sampled checked sentences: no invented transaction, amount, date, party or statement. Two carried a small unsupported detail the judge passed: a state inferred from a branch name ("Laredo, Texas") and "international" for wires to a named foreign beneficiary. 2 of 70 (3%) minor unsupported details passed; 0 of 70 invented facts.
  • Of 31 flagged sentences: 10 were "the bank is filing this report" (the judge wanted a source; fixed in v3 with a context line, which is why v3 has 23 flagged); the rest were real small inferences (a state from a branch name 6x, "domestic" counterparties 4x, "unusual"/"significantly lower" judgements, "the difference remained in the account").
  • Of 4 held: 1 bad citation, 3 judge false holds (see B).

F. 28 Sep 2026: invented amounts and fewer nitpick flags (before/after)

A blind tester playing a BSA analyst found two problems on the demo: the invented $25,000 wire to Mexico was reported as a number mismatch against a payroll credit ("the cited rows give $1,994.00"), and several flags were noise ("Laredo, TX" flagged because the rows say "Laredo"; "suspected structuring" and "not consistent with income", which are the institution's own conclusions). Two changes, the grounding prompt untouched:

  • Not in the case file. An amount that no row or field of the whole case file holds, that is not stated as a total or other aggregate, and that is not within 50% of the single cited figure it was compared with, is now reason not_in_source ("not in the case file (possibly invented)") with no nearest figure. A wrong total ($112,789 for $111,789) and a shifted date stay number_mismatch; an amount another row holds says which row. Code only: part A's recall is unchanged (a sentence is held either way; only the label differs).
  • Two context spans for the judge (S3.2, S3.3): the institution's characterisations (suspicious, suspected structuring, consistent or not with income or profile) are supported when the facts they rest on are cited, while a new fact beside them still needs a source; and a US state after a city the spans name is part of the place name.

Both runs on 28 Sep 2026, same pre-release server, direct route to the same Qwen3.8-27B weights (127.0.0.1:8114, not the gateway), temperature 0, DECOSA_SAR_WORKERS=2. "Before" is the code without the two changes.

Set Before After
C. Invented sentences (20): held / flagged / missed 19 / 1 / 0 18 / 2 / 0
C. True reference sentences (172): held / flagged 0 / 27 0 / 19
v3 drafts re-checked in check mode (12 drafts, 269 of the model's own sentences): flagged / held / not checked 22 / 2 / 1 14 / 3 / 0
  • Every invented sentence is still caught. One ("works for a money services business", the file says cash-intensive business) moved from held to flagged.
  • On the re-checked v3 sentences, 11 flags went away (states after Laredo and Raleigh, "the customer initiated" transfers the rows show, "generated on" an alert date) and 3 sentences that passed before were flagged ("international" wires to Dubai, "within three days"): run-to-run judge variance is of that size. The third hold is new and false ("four wire transfers", where the rows say 4 debits by wire).
  • The planted-errors sample now holds the invented wire as not in the case file (tests/test_sar.py).
  • Results: invented-28sep-{before,after}*.json, recheck-28sep-{before,after}*.json (scripts/eval_sar.py invented|recheck --tag=...). Not re-measured: parts B and E (drafting runs).

Screen alone (no model), 200 held-out cases

Planted typology indicated in 160 of 160 cases. Clean cases: 35 of 40 with nothing indicated, 5 of 40 (12.5%) flagged for structuring (a cash business whose daily deposits happen to land in the $8,000-$9,999 band). Funnel cases also trip structuring in 13 of 40 (funnel deposits are often under the threshold, as FIN-2014-A005 says). The screen's thresholds are ours and were set on this generator, so these numbers mostly show the rules fire as designed.

Cost

From the recorded demo runs (gateway list price $0.30 / $1.50 per million tokens): a structuring draft 31 calls, about 45,700 tokens, $0.017; an elder draft 29 calls, $0.013; a clean no-file rationale 15 calls, $0.009; the planted-errors check 7 calls, $0.003.

Fixes made after looking at eval runs

  • After v1: threshold comparisons ("over $10,000") are references to the declared threshold; a negation flips a comparator ("neither exceeding" is "at most") and makes it apply to each row; a number after an upper-case token is an identifier ("SSA TREAS 310"); "within one to two days" is a bound; "incoming"/"outgoing" map to the direction; the screen's typology fact no longer says "the screen marks"; distinct counts skip a subset fixed on the same column.
  • After v2: "began on ... ended on" is a range; an amount tied to a date must come from that day's rows; rows on a named day can be summed; the judge gets one context line ("this text is the SAR narrative the institution is filing").
  • After v3 (not re-measured): the funnel fact carries the deposits each debit followed; the prompt names A.institution.

Expected properties of the sample runs (for the rehearsal kit)

  1. planted-errors (check mode): status held; exactly the sentences what.1 (total $112,789.00; the rows give $111,789.00), when.1 (February 28, 2026; the rows give February 26, 2026) and when.2 (the invented $25,000.00 wire) are held, the first two with reason number_mismatch, the wire with not_in_source (since 28 Sep 2026); numbers.mismatch = 2 and numbers.unverified = 1; no drafting call (7 receipts).
  2. structuring (draft): 8 sections; screen.indicated = ["structuring"]; typologies.named_in_text contains structuring; numbers.mismatch = 0 in our runs.
  3. clean (no_file): 4 sections (alert, review, findings, decision); report.document = no_file_rationale; screen.indicated = []; the filing copy contains "No SAR is recommended".
  4. Every run: every receipt id resolves at /receipts/{id}; POST /sar/verify with the report, narrative_md, narrative_plain and the CSV returns valid_signature, both hash matches and ledger_matches true; the filing copy contains no [ and none of the held sentences.
  5. The hosted service refuses a case without "synthetic": true (HTTP 400 naming 31 U.S.C. 5318(g)(2)).

Rehearsal bundle: rehearsal/sar-narrative-desk/ (the planted-errors case in check mode; 13 checks, properties 1, 4 and the filing-copy rules above). python scripts/rehearse.py sar-narrative-desk passed 13/13 against the pre-release server on 26 Sep 2026 in 2.7 s.