Eval: No Surprises Act IDR packet and eligibility screen (use case 74)
Run 27 Sep 2026 on the pre-release server (127.0.0.1:8474, on a pre-release build) against the hosted Qwen3.8-27B through the model gateway. Every model call has a signed gateway receipt. All data is synthetic (CC0); the rules were read from primary sources on the same day (see "The law" below).
Four measurements, each on data written before the run it scores:
| What | Set | Result |
|---|---|---|
| Eligibility, rules only (full facts, no model) | 91 planted cases (all) | 91 / 91 verdicts right; 51 / 51 planted ineligible caught with the right check; 40 / 40 eligible (incl. 24 traps) passed |
| Eligibility from documents (the model reads every fact from EOB and notice text) | dev, 30 cases | 30 / 30; 243 / 243 scored facts read right |
| Eligibility from documents | test, 61 cases (held out) | 61 / 61 verdicts right; 32 / 32 ineligible caught with the right check; 29 / 29 eligible passed; 496 / 497 facts read right |
| Brief grounding: planted unsupported sentences | 3 briefs, 47 sentences | 25 / 25 planted removed; 22 / 22 supported kept |
| Deadline math against a second implementation | 20,000 random dates, 160,000 comparisons | 160,000 / 160,000 agree; holiday lists 2021-2028 identical |
Plus end to end on the three eligible samples (real drafting, not planted; the runs recorded for the Watch player at 6e71f71): 36 claim-bearing sentences, 35 kept and 1 flagged partial (a "these facts go to ..." summary line), none removed. I read each kept sentence against the documents: all are supported. (An earlier recording had one sentence citing only one of the two lines it rested on.) In the first runs, before the brief prompt asked for facts without adjectives, 13 of 25 claim sentences were flagged partial for added framing ("demonstrating specialized expertise") and 2 removed: a percentage of the QPA the model made up when no QPA had been read (right to remove), and a sentence whose "March 2025" the number guard misread as a stray year (a bug, fixed: the guard now counts the parts of dates on the cited line). One terminology slip ("Qualified Payment Amount") was fixed in the prompt.
1. Eligibility screen
Cases (scripts/idr_eval_cases.py, seed 74, written to docs/evals/nsa-idr-packet/cases.json): 91 synthetic disputes.
- 16 plain eligible cases (emergency medicine, anesthesiology, radiology, pathology, neonatology; 16 states without a specified state law; fully insured, individual, self-insured and FEHB coverage).
- 24 traps, eligible but easy to call ineligible: a signed consent form for an ancillary service (anesthesia: consent can't waive it), self-insured plans in Texas, California, New York or Florida (state law, but Federal IDR applies), a state law that does not reach the claim, initiation on exactly business day 34, open negotiation sent on exactly business day 30, a negotiation that ended inside a cooling-off period with initiation in the later 30-business-day window, a correct pre-November batch, and air ambulance (no notice and consent).
- 51 planted ineligible cases, 3 of each: open negotiation sent after business day 30; no open negotiation notice; initiated before business day 31 (negotiation not exhausted); initiated after business day 34; a state law that sets the payment (NY, TX, NJ, CA, FL, WA, GA, VA; fully insured or individual); a self-insured plan that opted into NJ, VA, WA, GA, NV or ME's process; a valid notice-and-consent waiver (non-ancillary surgery); Medicaid, Medicare Advantage or TRICARE; ground ambulance; an in-network provider; a coverage (medical necessity) denial; initiation inside a cooling-off period; a batch with two TINs; a batch across two self-insured plans; a batch spread over more than 30 business days; a 51-line batch after 1 Nov 2026; mixed codes after 1 Nov 2026.
The gold clocks come from numpy.busday_offset over an OPM holiday list typed by hand, not from the screen's own clock
module. Documents are rendered from the facts in four date styles with distractor dates on the same pages (the plan's
claim-received date, the remittance issue date, the notice draft date, the plan's reply date, the IDR draft date).
The split was fixed before any run: dev = the first 30 in shuffled order, test = the other 61.
Runs (scripts/idr_eval.py):
--mode rules:POST /idr/screenwith the full facts (no model). Result above; this tests the rules and the clocks, not reading.--mode documents:POST /idr/packet {mode: "screen"}with only the rendered documents (batches and earlier determinations stay structured, as a billing system would hold them). One model call per case. Median latency 6.2 s (max 12.7 s) on test under shared gateway load; 1 signed receipt per case; 61 / 61 signed packets verified.
Tuning. The extraction prompt was changed twice before the dev run (per-claim amounts for batches; a receipt date counts for the open-negotiation date when the send date is not written) after reading the first sample runs, not the eval cases. Distractor dates were added to the renderer after the first dev run (which also scored 30 / 30) to make reading harder; dev was re-run, then test was run once. Nothing changed after the test run.
The one miss on test (c064): the model gave the remittance issue date (5 Jul) instead of the receipt date (7 Jul) for
payment_received. The verdict was still right because the negotiation notice fell inside the window either way. A miss
like this moves the start-window deadline two days earlier, so it errs on the safe side for a provider; it would not for
a plan arguing a late start.
What this does not show. The documents are templated synthetic text with one fact per line, so reading them is much easier than reading real remittances (835s, scanned EOBs, portal screenshots) or free-text correspondence. The cases and the checker were written by the same author from the same reading of the rules, so the rules-only 91 / 91 shows the code does what I think the rules say, not that my reading is right. The CMS public use files record real IDR outcomes, including ineligibility, but they do not include the claim documents, so they could not be used to score this; they are the next step for checking the rule reading at scale (see "Not done").
2. Brief grounding
scripts/idr_grounding_eval.py runs three hand-written briefs (one per eligible sample, written once before the first run)
through the packet's own sentence check: the grounding judge (vertical 17), the numeric block against the lines the
sentence cites, and the prohibited-factor rule (45 CFR 149.510(c)(5)(v)). 22 sentences are supported by the documents; 25
are planted, 2-6 of each kind: a changed figure (6), a credential not in the documents (3), an invented statistic (3), a
contradicted fact (3), a moved date (3), billed charges (2), Medicare rates (3), usual and customary charges (2).
Result: 25 / 25 planted sentences removed (none only flagged), 22 / 22 supported sentences kept. By kind: number 6 / 6,
credential 3 / 3, statistic 3 / 3, contradicted 3 / 3, date 3 / 3, billed 2 / 2, medicare 3 / 3, ucr 2 / 2 removed.
47 model calls, all receipted. n is small and the planted sentences are the obvious kind; a subtle overstatement
("demonstrating superior quality") is what the judge marks partial and the packet shows as [CHECK], which the first
sample runs did before the brief prompt was told to state facts without adjectives.
3. Deadline math
scripts/idr_clock_crosscheck.py: 20,000 random start dates from 2022 to 2027, eight clocks each (open-negotiation last
day, IDR window open and close, business-day index, the single-dispute cooling-off window, the 2026 batched cooling-off
and window, the batch span before and after 1 Nov 2026), compared with numpy's busday functions over the hand-typed OPM
list. 160,000 / 160,000 agree, and the two holiday lists match for 2021-2028. Both implementations share my reading of
the counting conventions (a period "beginning on" a day counts it as day 1, CMS's timeline says so; a weekend start counts
the next business day as day 1, which is my reading), so this checks the arithmetic, not the reading.
4. Sample runs and cost
| Run | Receipts | Tokens | Cost at list price | Time |
|---|---|---|---|---|
| Full packet, emergency anesthesia sample | 23 | 52,263 | $0.0199 | 30-98 s over 7 runs (median 52 s) under shared load |
| Screen only, late-initiation sample | 1 | 1,662 | $0.0012 | 4.4 s |
Expected properties of the sample runs (for the rehearsal kit)
POST /idr/clockwith receipt 13 Aug 2026 and notice 24 Aug 2026 gives the window 6 to 9 Oct 2026 (Labor Day skipped).er-anesthesia-pais screenedlikely_eligiblefrom its documents, the QPA $1,184.00 comes with its EOB quote, and the $2,350.00 offer is 198.48% of it.- The billed-charges line (D1.L11), the Medicare line (D5.L3) and the FAIR Health line (D5.L4) are in
left_out, and the brief does not mention Medicare or billed charges. late-initiation-azislikely_ineligiblewithfailed == ["initiation"]and no brief.plan-consent-defenseislikely_ineligiblewithfailed == ["consent"].- The signed record verifies at
POST /record/verifyand fails once the statement's verdict is changed.
The law (read 27 Sep 2026)
- 45 CFR 149.510 on eCFR (current to 24 Sep 2026, last amended 28 Aug 2026), and its text in force on 1 May 2026.
- Federal IDR Operations final rules, 91 FR 33900 (4 Jun 2026), effective 3 Aug 2026; correction 91 FR 55462 (28 Aug 2026). Most process changes apply to disputes whose negotiation starts 90 days after each IDR Gateway function is announced (149.510(h)); only batching has been announced (CMS notice, 3 Aug 2026: negotiations starting on or after 1 Nov 2026). The $15 fee applies to disputes initiated on or after 11 Jun 2026. CMS's timeline guide (7 Aug 2026) expects the other functions from Spring 2027. The 2023 proposal (88 FR 75744) is superseded.
- CMS guidance for disputing parties (Dec 2023), batching FAQs (Nov 2023), FAQs Part 63, the state applicability chart (state information current as of 11 Jan 2023), notices of 3 Aug, 13 Aug and 15 Sep 2026.
- TMA II, No. 23-40217 (5th Cir. 2 Aug 2024), and TMA III en banc, No. 23-40605 (5th Cir. 11 Aug 2026): read. TMA I
(E.D. Tex. 23 Feb 2022) and TMA IV (E.D. Tex. 3 Aug 2023): not read; described as CMS's guidance describes them and
marked unverified in
/idr/info.
Not done
- Real documents: 835 remittances, scanned EOBs and portal PDFs (the document reader block would feed them in).
- State law beyond CMS's chart: each state's statute and enforcement letter, and the 2023 chart date.
- The CMS IDR public use files: checking the rule reading against real ineligibility determinations by category.
- The 2026 final rule's open negotiation, initiation and registry steps, once the Departments announce them.
- The "similar condition" test before 1 Nov 2026 is left to the certified IDR entity; the typed judgment is advisory and was not evaluated.
- The CPT Category I section ranges the 2026 batching rule refers to "as specified in guidance" were not found; ours are marked unverified.