Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: claim denial appeal packet (60)

Run 26 Sep 2026 on the pre-release server (127.0.0.1:8468) against the hosted Qwen3.8-27B through the model gateway, with every call receipted. The gateway was shared with other workloads' work (and down for about 30 minutes, 20:20 to 20:50 UTC; no run from that window is counted), so latencies were measured under load.

What was measured

The input is a synthetic denial, chart excerpts and a payer policy. Everything comes from scripts/appeal_cases.py:

  • denials in three layouts (a Medicare remittance notice, a payer letter, an "explanation of decision") with CARC/RARC codes and the notice date in five date formats; half of the requests give notice_date, half make the model read it;
  • charts built from a structured truth: which facts the chart documents, states as absent ("denies daytime sleepiness"), or leaves out;
  • policies: excerpts of CMS NCD 240.4 (CPAP), 100.1 (bariatric surgery) and 50.3 (cochlear implants), public domain, and an invented commercial policy for lumbar MRI (Harbor Health Plan MP-117);
  • plans: Original Medicare (initial and redetermination), Medicare Advantage, ERISA and ACA individual (initial and final internal), Medicaid managed care; plus billing denials (CO-16 with MA130, CO-18, CO-4, CO-29).

Each policy is split by hand into requirements ("units", e.g. for CPAP: clinical evaluation, a positive sleep test, ordered by the treating physician, the AHI criterion). The truth gives each unit's gold status (met, not met, undocumented), the gold recommendation, and the gold deadline, computed in the generator with its own arithmetic and hard-coded 2026-2027 federal holidays (not with deadlines.py).

Set Cases Gold: appeal / don't / get docs / fix claim Use
dev (seed 1) 18 4 / 6 / 7 / 1 all prompt and rule work happened on this set
test (seed 2) 40 15 / 11 / 11 / 3 run once, after the prompts were frozen (commit 062e510)
test2 (seed 3) 40 11 / 15 / 11 / 3 a fresh set, generated and run once after the date-window fix (commit 1876e1d)
samples 5 2 / 1 / 1 / 1 the demo cases on the site

Held-out phrasing. Odd-numbered test and test2 cases ("B") use chart and denial sentences that never appeared while the prompts were written ("Falls asleep unintentionally during the day, including in meetings", "Axial LBP x3 months", "EXPLANATION OF DECISION issued ..."). Even-numbered cases ("A") use the dev templates with new values and combinations. The policies are the same four in every set, so criteria mapping is not held out on policy text.

Scoring. A model criterion is matched to a gold unit by word overlap between its policy quote and the unit's anchor text. A unit's status is what the packet's own combination rule gives it (alternatives: any met; several criteria for one unit: the worst). "Planted unsupported" cases are those whose gold is don't appeal or get documentation first; they pass when the packet does not recommend an appeal and drafts no letter.

Changes made after looking at dev results (all before the test run):

  • The criteria prompt returns requirements with explicit options; lines ending in "or" become alternatives in code; a qualifier split off as its own option ("with documented symptoms ...") is merged back into the option before it.
  • The judge prompt: not_met needs a statement that contradicts the criterion (absence is undocumented); only the listed qualifying findings count; a CHECK line lists each part; the policy text is shown for definitions.
  • A quote stitched from the chart and the policy keeps only the chart's pieces as evidence.
  • Model calls retry with backoff when the shared gateway sheds load. Two generator bugs were fixed: an "axial only" MRI case no longer has a neurological deficit on exam (which meets the criterion), and "past medical history: none reported" was dropped from the undocumented co-morbidity case (the model fairly read it as no co-morbidity).

Change made after the test run: the test run had one wrong appeal: the model read a 53-day gap between an exam and a request as 23 days. Date windows ("within 30 days before the request") are now re-checked in code against the two dates the model compares; a mismatch makes the criterion undocumented. That is a fix motivated by a test failure, so test2 was generated fresh (new seed) and run once to measure it.

Changes after test2, not re-measured on a held-out set (checked on the five samples and the unit tests only): a template opening line in the letter, stray reference markers like [P2.L2] stripped, the letter prompt wording options as "the policy accepts X, among other tests" instead of "requires X", source titles counted by the number guard, and the /appeal/deadlines error messages naming its own fields.

Results

Recommendation and "don't appeal" correctness

dev (final) test (run once) test2 (fresh, run once)
Recommendation right 18/18 36/40 36/40
Planted unsupported: no appeal and no letter 13/13 21/22 26/26
Supported: appeal with a letter 4/4 14/15 11/11
Billing denials routed to "fix the claim" 1/1 3/3 3/3

Confusion (gold -> packet):

  • test: appeal->appeal 14, appeal->gather_first 1, do_not_appeal->do_not_appeal 10, do_not_appeal->appeal 1, gather_first->gather_first 9, gather_first->do_not_appeal 2, fix_claim 3/3.
  • test2: appeal->appeal 11, do_not_appeal->do_not_appeal 15, gather_first->gather_first 7, gather_first->do_not_appeal 4, fix_claim 3/3.

By phrasing: test A 18/20, B 18/20; test2 A 20/20, B 16/20. The held-out phrasing costs accuracy on test2: all four misses are B cases.

The errors.

  • Undocumented read as not met (6 of the 8 misses): the chart leaves a criterion out, and the model reads the silence as a contradiction. Examples: a past medical history that lists only allergic rhinitis, read as "no hypertension"; an exam "notable for paraspinal tenderness only", read as a neurological exam without the required findings. The advice becomes "don't appeal" where the gold says "get the documentation first". It never produces a letter, but a denials team would miss a chance to fix the chart. In several of these the gold is itself a judgement call.
  • A wrong appeal (test-035, before the fix): the exam was 53 days before the request (the window is 30); the model wrote "23 days". Fixed by the code-side window check; test2 had no wrong appeal.
  • A missed appeal (test-037): "furnished under appropriate physician supervision" was split into its own criterion and judged undocumented for an in-lab PSG interpreted by a board-certified physician.

Criteria mapping

dev test test2
Gold units found in the model's criteria 61/61 133/133 133/133
Unit status right 60/61 125/133 126/133
Model criteria returned / matching no unit or only the coverage statement 116 / 8 259 / 16 257 / 16

The "spurious" criteria are the NCDs' opening coverage statements ("CPAP is covered for adults with OSA"), which the model lists as a criterion and which the chart always meets. A recurring unit error that does not change the recommendation: for an AHI under 5, "a positive sleep test" is marked met because the test type matches (3 cases in test, 2 in test2).

Deadlines

test test2
Next level and date right 40/40 40/40
Notice date read from the denial (when not given) 20/20 20/20

The deadline rules are code, so this mostly measures the notice-date read and the level routing. The arithmetic itself is also unit-tested against the rules' own examples (30 Oct + 4 months = 1 Mar; weekend and Presidents' Day roll-forward for external review; Medicare's 5-day presumption).

Letters: unsupported assertions

Draft letters were written for 15 test and 11 test2 cases. Every sentence went through the grounding judge (vertical 17) and the number guard:

test test2
Sentences with a claim (judge verdict other than no_claim) 157 103
Removed (unsupported, contradicted or a number not in the sources) 9 6
Kept but marked [CHECK] (partly supported) 8 9

What was removed: opening lines that assert "the patient meets all coverage criteria" (the judge treats it as a claim), an AHI attributed to the consultation note instead of the sleep study, "conducted under the supervision of Dr. X", restatements that said the policy requires an attended PSG, and "NCD 50.3" (a number only in the policy's title; titles now count).

The kept sentences were read by the building agent against the chart, the policy and the denial: 0 of 245 kept sentences (148 test, 97 test2) state a clinical fact the chart does not support. Two kept, unflagged sentences overstate the policy ("The policy requires an attended PSG performed in a sleep laboratory", where the NCD lists four accepted tests); the letter prompt now words options as accepted alternatives (not re-measured). This rating is the author's own, not an independent rater's.

Cost and latency

  • The CPAP demo case with a letter: 21 model calls, 34,908 tokens, about $0.014 at the gateway list price. Without a letter (don't appeal): 12 calls, 18,489 tokens, $0.008. A billing denial: one call.
  • Per packet, under a shared gateway: median 12.9 s, max 32.7 s (test2); median 41.3 s, max 96.2 s (test, a busier hour).
  • Self-hosted (fresh clone, direct route to the local Qwen3.8-27B, attested receipts): the CPAP appeal in 11.9 s, the no-appeal case in 7.6 s.

Expected properties of the sample run (the rehearsal bundle checks these)

  1. POST /appeal/deadlines {"plan_type":"medicare_ab","level":"initial","notice_date":"2026-09-08"} gives 2027-01-11 (5 days' presumed receipt, then 120 days, 42 CFR 405.942), with no model call.
  2. cpap-appeal (AHI 11.2 with daytime sleepiness and hypertension) recommends appeal, reads the notice date from the denial, and the AHI 5-14 criterion is met with a chart quote containing "11.2"; the letter cites the AHI.
  3. cpap-no-appeal (AHI 10.4; the chart denies sleepiness, insomnia, hypertension, heart disease and stroke; a physician letter asserts the criteria are met) recommends do_not_appeal and drafts no letter.
  4. admin-missing-npi (CO-16, MA130) recommends fix_claim with no criteria and no letter, after one model call.
  5. The signed packet verifies at POST /record/verify, and fails once its recommendation is changed.
  6. Every model call has a signed receipt (gateway) or an attested one (self-host).

Limits

  • Synthetic, templated cases written by the same author as the prompts and the generator. Real charts are long, scanned and inconsistent, commercial policies are denser than these excerpts, and real denials often give vaguer reasons.
  • Four policies. Nested alternatives (an option that is itself a list of options, like NCD 240.2's Group II) are read as one option; non-covered indications are not mapped.
  • The gold for "undocumented" versus "not met" is a judgement call in a few chart phrasings; the misses lean towards "don't appeal", never towards an unsupported letter after the date fix.
  • Deadlines are federal minimums; state programs, plan documents and provider contracts are not modelled beyond a contract window the caller gives.
  • No real denials, and no denials specialist has rated the letters.

Verdict

Would a buyer pay? A denials team would pay for the two things that are hard to get today: a recommendation that says "don't appeal" when the chart doesn't support it (48 planted unsupported cases, 47 correct, and the one miss is now caught in code), and a letter where each clinical sentence is traceable to the chart. The deadline table with the CFR cite is a real time-saver for Medicare and ERISA work. At about $0.014 per packet it is far below the ~$20 per appeal that AI appeal generators charge and the ~$118 staff cost of a formal appeal (MGMA figure cited second-hand, wiki page 40).

What's missing before a customer relies on it:

  1. Measurement on real, de-identified denials with a denials specialist's labels, including the "undocumented versus not met" calls, where the model leans to "don't appeal".
  2. Commercial payer policies as customers use them (long InterQual-style or payer CPB criteria with nested options), and an OCR step for scanned charts and EOBs.
  3. Payer-specific filing details (addresses, portals, forms) and state external-review rules; the packet deliberately leaves these to a person.