Skip to content
decosa

74 · Healthcare · Finance and insurance · live

No Surprises Act IDR packet and eligibility screen

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)

  • Eligibility verdict right from documents, held out61 of 61test splitn = 61One model call per case reads every fact; split fixed before any run and run once; templated synthetic documents
  • Planted ineligible disputes caught on the planted check32 of 32test splitn = 3217 kinds: late or missing negotiation, early or late initiation, state law, opt-in, consent, public coverage, ground ambulance, in network, coverage denial, cooling-off, four batching faults
  • Eligible disputes passed, including traps29 of 29test splitn = 29Traps: ancillary consent, self-insured in state-law states, day-30 and day-34 boundaries, a moved cooling-off window, air ambulance
  • Facts read right from the documents496 of 497test splitn = 497The miss: an issue date read as the receipt date; the verdict was still right
  • Eligibility, rules only with full facts91 of 91syntheticn = 91No model; gold clocks from a separate numpy implementation
  • Planted unsupported brief sentences removed25 of 25syntheticn = 25Written once before the run; 22 of 22 supported sentences kept
  • Clock comparisons agreeing with a second implementation160,000 of 160,000syntheticn = 160,00020,000 random dates, 8 clocks, 2022-2027

Dataset

91 synthetic disputes from scripts/idr_eval_cases.py (seed 74): dev 30, test 61; three hand-written briefs with 25 planted sentences; 20,000 random dates for the clocks. All CC0.

Caveats

  • Synthetic, templated documents with one fact per line, written by the same author as the rules and prompts; real remittances, scanned EOBs and letters are harder. These numbers do not predict accuracy on real claim files.
  • The rules score shows the code matches the author's reading of 45 CFR 149.510 and CMS guidance, not that the reading is right; it was not checked against CMS's recorded eligibility outcomes.
  • The planted-sentence set is small (47 sentences) and the planted errors are the obvious kind.
  • The extraction prompt was adjusted on the demo samples before the dev run; distractor dates were added after the first dev run; test was run once.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
27 Sep 2026
Latency, this run
n/a
p50 over passed runs
52 s
Receipts
23
Model calls
n/a
Tokens
n/a
Cost per run
$0.020

Self-host verification

Verified on 27 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): window 2026-10-06 to 2026-10-09, the late case failed on initiation with no model, er-anesthesia-pa likely eligible with a brief at 198.48% of the QPA in 20 s (23 attested receipts), late-initiation-az and plan-consent-defense likely ineligible with no brief, record verified; the rehearsal bundle passed 15/15. Model-server startup was not re-run.

Rehearsal bundle: nsa-idr-packet.zip (4 KB, 15 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.
  • Measured on 91 synthetic, templated disputes written by the building agent; not on real remittances or against CMS's recorded eligibility outcomes.
  • State law comes from CMS's chart (state information current as of 11 Jan 2023); a state-regulated plan in one of the 21 listed states is marked needs review unless the file says whether the state law applies.
  • The pre-1 Nov 2026 similar-condition batching test is left to the certified IDR entity; the CPT section ranges for the 2026 test are our approximation (the guidance was not found).
  • The 2026 final rule's new portal steps are not modelled yet; the clocks will need an update when the Departments announce them.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Reads the claim facts from the EOB, remittance and notices (each with a quote), gathers the evidence for each factor the arbiter must consider, drafts the offer brief, and judges every brief sentence (the grounding judge)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Does it catch the disputes that should not be filed?

  • Held-out test, verdict right from documents: 61 of 61 (32 ineligible caught on the planted check, 29 eligible passed)
  • Facts read right from documents: 496 of 497 (test; the miss read an issue date as the receipt date)
  • Planted unsupported brief sentences removed: 25 of 25 (and 22 of 22 supported sentences kept)
  • Cost: about $0.001 per screen, $0.02 per packet (1,662 tokens (1 call) and 52,263 tokens (23 calls) on the samples, gateway list price)

Source: decosa-api docs/evals/nsa-idr-packet.md, 27 Sep 2026

Lite · one 48 GB card (1)
  • eligibility and brief grounding: not measured yet
Standard · the hosted demo, one 96 GB card (6)
  • eligibility verdict from documents alone, held-out test (61 synthetic disputes): 61/61; 32/32 planted ineligible caught on the planted check; 29/29 eligible, including traps, passeddecosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27, gateway route; split fixed before any run; templated synthetic documents
  • eligibility, rules only with full facts (91 planted cases): 91/91; gold clocks from a separate implementationdecosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27
  • facts read right from the documents (dev / test): 243/243 / 496/497 (the miss: an issue date read as the receipt date)decosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27
  • planted unsupported brief sentences removed / supported sentences kept: 25/25 / 22/22 (changed figures, invented credentials and statistics, contradictions, moved dates, billed charges, Medicare and UCR rates)decosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27; small n
  • business-day clocks against a second implementation (numpy over the OPM holiday list): 160,000/160,000 comparisons agreedecosa-api docs/evals/nsa-idr-packet.md, 2026-09-27
  • real remittances, scanned EOBs and IDR outcomes: not measured yet
Best · DeepSeek-V4-Flash on two more cards (1)
  • eligibility and brief grounding: not measured yet
Wanted · two large judges from different families (1)
  • eligibility and brief grounding: not measured yet

How we measure · All tools