74 · Healthcare · Finance and insurance · live
No Surprises Act IDR packet and eligibility screen
Eval results
Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)
- Eligibility verdict right from documents, held out61 of 61test splitn = 61One model call per case reads every fact; split fixed before any run and run once; templated synthetic documents
- Planted ineligible disputes caught on the planted check32 of 32test splitn = 3217 kinds: late or missing negotiation, early or late initiation, state law, opt-in, consent, public coverage, ground ambulance, in network, coverage denial, cooling-off, four batching faults
- Eligible disputes passed, including traps29 of 29test splitn = 29Traps: ancillary consent, self-insured in state-law states, day-30 and day-34 boundaries, a moved cooling-off window, air ambulance
- Facts read right from the documents496 of 497test splitn = 497The miss: an issue date read as the receipt date; the verdict was still right
- Eligibility, rules only with full facts91 of 91syntheticn = 91No model; gold clocks from a separate numpy implementation
- Planted unsupported brief sentences removed25 of 25syntheticn = 25Written once before the run; 22 of 22 supported sentences kept
- Clock comparisons agreeing with a second implementation160,000 of 160,000syntheticn = 160,00020,000 random dates, 8 clocks, 2022-2027
Dataset
91 synthetic disputes from scripts/idr_eval_cases.py (seed 74): dev 30, test 61; three hand-written briefs with 25 planted sentences; 20,000 random dates for the clocks. All CC0.
Caveats
- Synthetic, templated documents with one fact per line, written by the same author as the rules and prompts; real remittances, scanned EOBs and letters are harder. These numbers do not predict accuracy on real claim files.
- The rules score shows the code matches the author's reading of 45 CFR 149.510 and CMS guidance, not that the reading is right; it was not checked against CMS's recorded eligibility outcomes.
- The planted-sentence set is small (47 sentences) and the planted errors are the obvious kind.
- The extraction prompt was adjusted on the demo samples before the dev run; distractor dates were added after the first dev run; test was run once.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 27 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 52 s
- Receipts
- 23
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.020
Self-host verification
Verified on 27 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after
The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): window 2026-10-06 to 2026-10-09, the late case failed on initiation with no model, er-anesthesia-pa likely eligible with a brief at 198.48% of the QPA in 20 s (23 attested receipts), late-initiation-az and plan-consent-defense likely ineligible with no brief, record verified; the rehearsal bundle passed 15/15. Model-server startup was not re-run.
Rehearsal bundle: nsa-idr-packet.zip (4 KB, 15 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.
- Measured on 91 synthetic, templated disputes written by the building agent; not on real remittances or against CMS's recorded eligibility outcomes.
- State law comes from CMS's chart (state information current as of 11 Jan 2023); a state-regulated plan in one of the 21 listed states is marked needs review unless the file says whether the state law applies.
- The pre-1 Nov 2026 similar-condition batching test is left to the certified IDR entity; the CPT section ranges for the 2026 test are our approximation (the guidance was not found).
- The 2026 final rule's new portal steps are not modelled yet; the clocks will need an update when the Departments announce them.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Reads the claim facts from the EOB, remittance and notices (each with a quote), gathers the evidence for each factor the arbiter must consider, drafts the offer brief, and judges every brief sentence (the grounding judge)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Does it catch the disputes that should not be filed?
- Held-out test, verdict right from documents: 61 of 61 (32 ineligible caught on the planted check, 29 eligible passed)
- Facts read right from documents: 496 of 497 (test; the miss read an issue date as the receipt date)
- Planted unsupported brief sentences removed: 25 of 25 (and 22 of 22 supported sentences kept)
- Cost: about $0.001 per screen, $0.02 per packet (1,662 tokens (1 call) and 52,263 tokens (23 calls) on the samples, gateway list price)
Source: decosa-api docs/evals/nsa-idr-packet.md, 27 Sep 2026
Lite · one 48 GB card (1)
- eligibility and brief grounding: not measured yet
Standard · the hosted demo, one 96 GB card (6)
- eligibility verdict from documents alone, held-out test (61 synthetic disputes): 61/61; 32/32 planted ineligible caught on the planted check; 29/29 eligible, including traps, passeddecosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27, gateway route; split fixed before any run; templated synthetic documents
- eligibility, rules only with full facts (91 planted cases): 91/91; gold clocks from a separate implementationdecosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27
- facts read right from the documents (dev / test): 243/243 / 496/497 (the miss: an issue date read as the receipt date)decosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27
- planted unsupported brief sentences removed / supported sentences kept: 25/25 / 22/22 (changed figures, invented credentials and statistics, contradictions, moved dates, billed charges, Medicare and UCR rates)decosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27; small n
- business-day clocks against a second implementation (numpy over the OPM holiday list): 160,000/160,000 comparisons agreedecosa-api docs/evals/nsa-idr-packet.md, 2026-09-27
- real remittances, scanned EOBs and IDR outcomes: not measured yet
Best · DeepSeek-V4-Flash on two more cards (1)
- eligibility and brief grounding: not measured yet
Wanted · two large judges from different families (1)
- eligibility and brief grounding: not measured yet