Eval: filing tie-out and MD&A grounding (57)
Run on our server, 26 Sep 2026. Scripts: scripts/tieout_build_set.py (data), scripts/eval_tieout.py (code and
claims). Result files: docs/evals/filing-tieout/.
Data
- Filings: FY2025 10-Ks of US large accelerated filers with inline XBRL, listed in the SEC Financial Statement
Data Sets 2026q1 (
sub.txt), ordered by SHA-256 of the accession number (no hand-picking), primary documents fetched from EDGAR with a declared User-Agent at two requests a second at most. SEC filings and data sets are US government works (public domain). Three sets, fetched at different times:- dev (hash positions 0-59, 60 filings): the rules were written and tuned on these.
- dev2 (positions 60-99, 40 filings): fetched as a hold-out, run once, then its false flags were used for a second round of fixes. It is development data now.
- test (positions 100-139, 40 filings): fetched after the rules were frozen at commit
abc8856. The headline numbers are the frozen run. After seeing it we fixed three bug classes (fiscal years named after the calendar year they end in, "in comparison to the year ended ...", run-to-run ordering), so a second run of the current code on the same set is shown too; it is not a clean test.
- Clean text: each filing's own MD&A (Item 7), untouched. Every flag on it is counted as a false flag unless a person (the building agent) checked the filing and found the figure really is inconsistent; those are listed.
- Planted errors, made in code with a seeded RNG per filing on real MD&A sentences, up to four per type per filing,
each checked on its own (one edited sentence at a time):
changed_figure: an amount that tied or computed, moved by 4-25% up or down at its stated rounding;wrong_pct_change: a percentage change that computed, moved by 3-12 points;swapped_period: an amount that tied, replaced by the same line's other-year value in the same style;scale_error: million / billion / thousand swapped on an amount that tied;direction_flip: "increased" / "decreased" (and rose/fell, grew/declined, higher/lower) swapped in a sentence whose figure tied or computed;note_vs_statement: the second showing of a fact that a note table repeats from another table (same concept, period and members) changed by 7% in the HTML, then the filing re-parsed;changed_any: likechanged_figurebut on any amount, including ones the clean pass could not trace (so it measures recall over all of MD&A, where most figures are not in the statements).
- A plant is caught only when the edited figure itself (or the flipped verb, or the changed fact) comes back as a mismatch. "Not tied" (mismatch or untraced) is recorded too.
Parser check
Every face-statement USD fact of the filing in SEC's own num.txt (Financial Statement Data Sets), compared with our
iXBRL parse: dev 16,725 of 16,725; dev2 10,885 of 10,902; test 10,808 of 10,813 agree to the dollar.
Results (code only, no model)
| dev (60) | dev2 (40) | test, frozen (40) | test, current code | |
|---|---|---|---|---|
| MD&A figures | 8,786 | 6,118 | 6,039 | 6,039 |
| tied or computed | 34.7% | 32.4% | 34.6% | 34.9% |
| flags on the clean MD&A | 6 (2 real) | 4 (1 real) | 13 (2 real) | 6 (2 real) |
| false flags per 1,000 figures | 0.46 | 0.49 | 1.8 | 0.66 |
| filings with any flag | 6 of 60 | 3 of 40 | 6 of 40 | 5 of 40 |
| changed_figure caught | 48/230 (21%) | 30/151 (20%) | 22/153 (14%, CI 10-21%) | 25/153 (16%) |
| wrong_pct_change caught | 40/117 (34%) | 19/74 (26%) | 14/91 (15%, 9-24%) | 15/91 (16%) |
| swapped_period caught | 75/221 (34%) | 45/142 (32%) | 38/150 (25%, 19-33%) | 46/150 (31%) |
| scale_error caught | 139/206 (67%) | 72/128 (56%) | 88/142 (62%, 54-70%) | 91/142 (64%) |
| direction_flip caught | 67/167 (40%) | 26/95 (27%) | 38/130 (29%, 22-38%) | 45/130 (35%) |
| note_vs_statement caught | 120/120 | 77/77 | 79/79 (100%, 95-100%) | 79/79 |
| changed_any caught | 19/236 (8%) | 10/152 (7%) | 10/156 (6%) | 11/156 (7%) |
95% Wilson intervals on the test column. The dev columns are the current code (deterministic ordering).
Real inconsistencies found in filed 10-Ks (checked by hand against the EDGAR documents):
- Applied Optoelectronics (0001437749-26-005875): MD&A says the accumulated deficit at 31 Dec 2025 was $491.0 million; the balance sheet shows $490,078 thousand, and the prior year ($451.9M) and the year's net loss ($38.2M) agree with the balance sheet.
- Vaxcyte: MD&A refers to "Note 10, Equity Incentive Plans"; Equity Incentive Plans is Note 9.
- NETGEAR: "Note 3, Balance Sheet Components"; Balance Sheet Components is Note 4 (Note 3 is Business Acquisition).
- Harmonic (two): "Note 3, Leases" and "Note 9, Debt"; Leases is Note 4 and Debt is Note 8 (a Discontinued Operations note was inserted as Note 3).
What the false flags were (test, frozen): G-III Apparel ×5 (it calls the year ended Jan 2026 "fiscal 2026" while the tags call it 2025; fixed after the test), First Commonwealth ×3 ("in comparison to the year ended December 31, 2024" read as a 2024 figure; fixed), MSCI (a 7.0% change in a cost line the MD&A measures differently), Plug Power (a component in an explanation), Rocket (a narrower measure with the same label). Current code on dev, dev2 and test still flags segment or scope figures whose label matches a consolidated line (Southwest Gas, Capital One, Schwab, Boeing, ICE, Levi's quarterly figure, Banc of California).
Throughput (one CPU core, parse plus tie, no model): median 0.44-0.5 s per 10-K; 740-1,020 sentences a second.
Claims in words (the grounding judge, Qwen3.8-27B through the gateway)
40 direction_flip plants from the test set and the same 40 sentences unflipped, each sent to the judge with the
named lines' figures (this forces the judge on every sentence; the product sends only sentences whose subject is a
line, at most 12 per run). 80 receipted calls, 8.5 s mean per call on the shared gateway.
| flipped (should be flagged) | original (should not) | |
|---|---|---|
| contradicted | 31/40 (78%) | 2/40 (5%) |
| partial or unsupported | 8 | 23 |
| supported | 1 | 15 |
| code direction check alone | 18/40 | 0/40 |
| code or judge | 31/40 |
The two false contradictions are period errors: a 2024-vs-2023 sentence and a "fiscal 2026" sentence read against the FY2025 columns. The judge calls explanations it cannot see in the tables partial or unsupported; the product shows those as untraced, not as flags.
Re-measured 2026-09-30 after P1-1 and the batch16 code fixes (filing-tieout/claims-results-p1-1.json; the same 40
pairs, direct route to the same weights, so latency is not comparable). Since P1-1 a contradicted verdict is shown as
a claim mismatch only when the judge is sure (certainty high) and compared a line the sentence names; otherwise the
sentence stays untraced (a person looks). The code's own direction check now also reads a sentence's stated "to X from Y".
| flipped (should be flagged) | original (should not) | |
|---|---|---|
| judge verdict contradicted | 32/40 | 3/40 |
| shown as a claim mismatch | 20/40 | 3/40 |
| code direction check | 23/40 | 0/40 |
| code direction check or claim mismatch | 27/40 |
The rule trades judge recall for fewer false alarms on real filings: on the Microsoft FY2026 10-K the hosted default path went from 6-8 claim mismatches (all false) to none. The judge's own run-to-run variation on these pairs is a sentence or two either way (an earlier re-run the same day: 20/40 and 2/40). The page quotes this file.
Hosted and self-host runs
- Hosted route (pre-release server, gateway): the planted sample with claims 7.2-9.8 s (p50 9.5 s), one receipted call, about 970 tokens ($0.0004 at the gateway list price); without claims 0.1-0.2 s. The AAOI filing (whole MD&A, 316 sentences) 7.9 s with two claim calls. An EDGAR fetch of a filing not in the cache (Eastman Chemical) 2.4-2.8 s.
- Self-host: fresh clone of the branch, api image built from it, direct route to the local Qwen3.8-27B vLLM (following the site's assemble prompt): planted sample 0.21 s without claims and 0.61 s with one, the clean sample no flags, the record verified and a changed status failed, the EDGAR fetch worked from the container.
Cost per filing (30 Sep 2026, filing-tieout/cost-per-filing.json)
The page used to quote the planted sample (one claim read), which understates a real 10-K: each claim in words is one
judge call, up to the claims cap per run. scripts/tieout_cost.py runs whole filed 10-Ks (AMD and AAOI as filed, and
Microsoft FY2026 from EDGAR) with the claims read and keeps each run's own usage at list price; the pricing file reads
the mean from that file, and the page gives the range in words. Direct route to the same weights.
Checkable properties of the sample runs (rehearsal: rehearsal/filing-tieout/)
amd-fy2025-planted: statusmismatches; figure mismatches with reasons value, value, period, scale; one direction flag on the Embedded sentence; one wrong note number (Note 6); onesame_fact_two_valuescross-reference; 44 of 57 figures traced.amd-fy2025(the same text as filed): no mismatch.aaoi-fy2025: exactly one mismatch, the $491.0 million accumulated deficit, naming $490,078,000.synthetic-tables: the 2024 net income written as $64.2 million is a value mismatch naming $62,400,000; the inventories accounts do not roll up to the balance sheet; cash and receivables do.- Every record verifies with its Markdown and CSV workpapers, and a changed status breaks the signature.
Caveats
- Recall is low by design: a near miss is only called when the sentence is surely about that line (another figure in it ties to the line, or the line's label is exactly its subject) and the period is explicit. Most planted changes in single-figure sentences come back untraced, not flagged. About two thirds of MD&A figures are not in the tagged tables at all (non-GAAP measures, segment detail, deal terms, statistics).
- The planted errors are ours; real draft errors may be easier (stale figures from a prior draft) or harder.
- One agent wrote the rules, planted the errors and judged which flags were real.
- MD&A tables are not tied (prose only). 10-Q quarters are handled by the period rules but not evaluated.
- The claim eval forces the judge on every sentence; the product's selection was changed after the eval (it now sends only sentences whose subject is a line), so the product's own judge precision on clean filings is not measured beyond the four recorded sample runs (no contradicted verdict on clean text).