Skip to content
decosa

57 · Finance and insurance · Compliance and trust · live

Filing tie-out and MD&A grounding

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)

  • False flags on untouched held-out 10-K MD&As (frozen rules)11 in 6,039 figures (1.8 per 1,000)test splitn = 6,03913 flags in 6 of 40 filings; 2 were real inconsistencies (stale note numbers in Harmonic). 6 flags, 2 real, after post-test fixes (not a clean test).
  • Note table that disagrees with a statement, caught79/79test splitn = 79
  • Scale error (million/billion) caught88/142 (62%)test splitn = 142
  • Another period's figure caught38/150 (25%)test splitn = 150
  • Flipped direction caught in code38/130 (29%)test splitn = 130
  • Wrong percentage change caught14/91 (15%)test splitn = 91
  • Changed figure caught22/153 (14%)test splitn = 153Figures that tied in the clean pass; over all MD&A amounts 10/156.
  • Flipped direction claims shown as a claim mismatch20/40test splitn = 403 of the same 40 sentences unflipped were also shown as mismatches; flips the judge called contradicted but was not sure of go to review. Measured 30 Sep on the direct route (same weights).
  • MD&A figures tied or computed34.6%test splitn = 6,039

Dataset

FY2025 10-Ks of US large accelerated filers from the SEC Financial Statement Data Sets 2026q1, picked by hash of the accession number and fetched from EDGAR: 100 for development (two rounds) and 40 fetched after the rules were frozen (test). Errors planted in code in the real MD&A text; the untouched MD&A is the clean set.

Caveats

  • One agent wrote the rules, planted the errors and judged which flags on clean filings were real errors.
  • The planted errors are ours; real draft errors may differ.
  • Recall is low by design; most figures are not in the tagged tables at all and come back untraced.
  • The claim-judge eval forced the judge on every sentence; the product's sentence selection changed afterwards.
  • MD&A tables and 10-Q quarters are not measured.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
26 Sep 2026
Latency, this run
n/a
p50 over passed runs
9.5 s
Receipts
1
Model calls
n/a
Tokens
n/a
Cost per run
<$0.001

Self-host verification

Verified on 26 Sep 2026: fresh clone, compose up, sample against local model servers

A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_TIEOUT_PUBLIC_ONLY=0, on the direct route to the running local Qwen3.8-27B. Re-run on commit 2562f1c with the rehearsal bundle (16 of 16 checks). The planted sample caught all seven plants in 0.21 s (0.61 s with one claim read), the clean sample had no flags, the record verified and a changed status failed, and an EDGAR fetch worked from the container.

Rehearsal bundle: filing-tieout.zip (3 KB, 16 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Only MD&A prose is tied; tables inside MD&A are not.
  • Recall is low by design: most wrong figures in single-figure sentences come back untraced rather than flagged.
  • Figures for segments, non-GAAP measures or narrower scopes that share a line's label can still be flagged against the consolidated line.
  • Drafts must be iXBRL or pasted text with CSV tables; Word and PDF are not read.
  • 10-Q quarter periods are handled but were not evaluated.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • iXBRL parser, figure tie-out, period and scale rules, direction and cross-reference checks, workpaper and signed record (no model; CPU)decosa-api tie-out (decosa_api/verticals/tieout) with the numeric grounding block's table-cell matcher (decosa_api/verticals/numeric/cells.py)AGPL-3.0-or-later
  • Claims in words (optional): the grounding judge reads a sentence against the named lines' figuresQwen3.8-27B (NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · figures only, no GPU (4)
  • False flags on untouched held-out 10-K MD&As (40 filings, frozen rules): 11 in 6,039 figures (1.8 per 1,000); 2 more flags were real errorsdocs/evals/filing-tieout.md, 26 Sep 2026
  • Planted errors caught, held-out: note table vs statement / scale / period / direction / % change / changed figure: 79/79 / 88/142 / 38/150 / 38/130 / 14/91 / 22/153docs/evals/filing-tieout.md, 26 Sep 2026
  • MD&A figures tied or computed (held-out): 34.6%docs/evals/filing-tieout.md, 26 Sep 2026
  • Parser vs SEC's own extracted statement figures: 10,808 of 10,813docs/evals/filing-tieout.md, 26 Sep 2026
Standard · one GPU for claims in words (hosted demo) (2)
  • Flipped direction claims shown as mismatches (40 held-out sentences): 20/40docs/evals/filing-tieout/claims-results-p1-1.json, 30 Sep 2026
  • True sentences shown as mismatches (the same 40, unflipped): 3/40docs/evals/filing-tieout/claims-results-p1-1.json, 30 Sep 2026

How we measure · All tools