57 · Finance and insurance · Compliance and trust · live
Filing tie-out and MD&A grounding
Eval results
Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)
- False flags on untouched held-out 10-K MD&As (frozen rules)11 in 6,039 figures (1.8 per 1,000)test splitn = 6,03913 flags in 6 of 40 filings; 2 were real inconsistencies (stale note numbers in Harmonic). 6 flags, 2 real, after post-test fixes (not a clean test).
- Note table that disagrees with a statement, caught79/79test splitn = 79
- Scale error (million/billion) caught88/142 (62%)test splitn = 142
- Another period's figure caught38/150 (25%)test splitn = 150
- Flipped direction caught in code38/130 (29%)test splitn = 130
- Wrong percentage change caught14/91 (15%)test splitn = 91
- Changed figure caught22/153 (14%)test splitn = 153Figures that tied in the clean pass; over all MD&A amounts 10/156.
- Flipped direction claims shown as a claim mismatch20/40test splitn = 403 of the same 40 sentences unflipped were also shown as mismatches; flips the judge called contradicted but was not sure of go to review. Measured 30 Sep on the direct route (same weights).
- MD&A figures tied or computed34.6%test splitn = 6,039
Dataset
FY2025 10-Ks of US large accelerated filers from the SEC Financial Statement Data Sets 2026q1, picked by hash of the accession number and fetched from EDGAR: 100 for development (two rounds) and 40 fetched after the rules were frozen (test). Errors planted in code in the real MD&A text; the untouched MD&A is the clean set.
Caveats
- One agent wrote the rules, planted the errors and judged which flags on clean filings were real errors.
- The planted errors are ours; real draft errors may differ.
- Recall is low by design; most figures are not in the tagged tables at all and come back untraced.
- The claim-judge eval forced the judge on every sentence; the product's sentence selection changed afterwards.
- MD&A tables and 10-Q quarters are not measured.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 26 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 9.5 s
- Receipts
- 1
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- <$0.001
Self-host verification
Verified on 26 Sep 2026: fresh clone, compose up, sample against local model servers
A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_TIEOUT_PUBLIC_ONLY=0, on the direct route to the running local Qwen3.8-27B. Re-run on commit 2562f1c with the rehearsal bundle (16 of 16 checks). The planted sample caught all seven plants in 0.21 s (0.61 s with one claim read), the clean sample had no flags, the record verified and a changed status failed, and an EDGAR fetch worked from the container.
Rehearsal bundle: filing-tieout.zip (3 KB, 16 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Only MD&A prose is tied; tables inside MD&A are not.
- Recall is low by design: most wrong figures in single-figure sentences come back untraced rather than flagged.
- Figures for segments, non-GAAP measures or narrower scopes that share a line's label can still be flagged against the consolidated line.
- Drafts must be iXBRL or pasted text with CSV tables; Word and PDF are not read.
- 10-Q quarter periods are handled but were not evaluated.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- iXBRL parser, figure tie-out, period and scale rules, direction and cross-reference checks, workpaper and signed record (no model; CPU)decosa-api tie-out (decosa_api/verticals/tieout) with the numeric grounding block's table-cell matcher (decosa_api/verticals/numeric/cells.py)AGPL-3.0-or-later
- Claims in words (optional): the grounding judge reads a sentence against the named lines' figuresQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · figures only, no GPU (4)
- False flags on untouched held-out 10-K MD&As (40 filings, frozen rules): 11 in 6,039 figures (1.8 per 1,000); 2 more flags were real errorsdocs/evals/filing-tieout.md, 26 Sep 2026
- Planted errors caught, held-out: note table vs statement / scale / period / direction / % change / changed figure: 79/79 / 88/142 / 38/150 / 38/130 / 14/91 / 22/153docs/evals/filing-tieout.md, 26 Sep 2026
- MD&A figures tied or computed (held-out): 34.6%docs/evals/filing-tieout.md, 26 Sep 2026
- Parser vs SEC's own extracted statement figures: 10,808 of 10,813docs/evals/filing-tieout.md, 26 Sep 2026
Standard · one GPU for claims in words (hosted demo) (2)
- Flipped direction claims shown as mismatches (40 held-out sentences): 20/40docs/evals/filing-tieout/claims-results-p1-1.json, 30 Sep 2026
- True sentences shown as mismatches (the same 40, unflipped): 3/40docs/evals/filing-tieout/claims-results-p1-1.json, 30 Sep 2026