Eval: prior-auth pre-check (102)
Run 28 Sep 2026 on the pre-release server (127.0.0.1:8477) against Qwen3.8-27B through the model gateway (every call receipted), with other workloads sharing the gateway, so latencies are under load. Dev rounds used our server's direct route to the same weights. Numbers are as run; nothing below was re-run to look better.
What was measured
The input is a real, published payer medical policy, a synthetic chart and the requested service. The output is a decision (ready to send / get records first / not supported on this chart / can't check licensed criteria), each criterion's status with quotes, form answers, a peer-to-peer brief, and a letter of medical necessity only when ready.
Policies (all real, published, quoted in excerpt with their source; code tables removed; no InterQual or MCG text).
- Dev (all prompt and rule work): Aetna CPB 0004 (CPAP), Cigna 0106 (CGM), Cigna 0051 (bariatric surgery), and the UnitedHealthcare panniculectomy policy (it points to InterQual). 13 cases.
- Held-out test (run once, prompts frozen at commit bb4cf18): Aetna CPB 0660 (knee arthroplasty), 0050 (varicose veins), 0194 (spinal cord stimulation), 0236 (lumbar MRI), 0017 (breast reduction); Cigna 0045 (blepharoplasty), 0119 (septoplasty), 0027 (panniculectomy); UnitedHealthcare epidural steroid injections and vagus nerve stimulation; CMS LCD L33797 (home oxygen) and L38803 (facet joint interventions). 12 policies x 5 cases = 60.
- Fresh held-out test2 (run once, at commit 9069217, after the guards below): Aetna CPB 0016 (sacroiliac joint injection), 0084 (upper-lid blepharoplasty), 0157 (bariatric surgery), 0325 (physical therapy); Cigna 0053 (hyperbaric oxygen), 0152 (reduction mammoplasty); UnitedHealthcare OSA treatment (hypoglossal nerve stimulation). 7 policies x 5 cases = 35. (Cigna 0063 was dropped: its current text states no coverage criteria.)
- Fresh test3 (run once, at commit 1da15a0, after the fixes from two blind cold-user tests): UnitedHealthcare bariatric surgery, Aetna CPB 0660 section I.B (knee revision), Cigna 0119 rhinoplasty section, Cigna 0045 brow ptosis section. 4 policies x 4 cases = 16. (Cigna 0158 and the Aetna gynecomastia section were dropped: no criteria in their text.)
- Guard set: 3 cases on UnitedHealthcare policies that send reviewers to InterQual (breast reduction, CGM).
Charts. Written by blind sub-agents (Claude Code, Opus) that never saw the tool or its prompts, from a brief: per policy, 5 cases: all met; all met with a trap (a less common option, a value only in a test report, an irrelevant negative); one criterion undocumented by silence; one undocumented by vagueness (mentioned without the value, duration or date; or only claimed in a physician's letter); one not met by a stated contradiction. Each author also wrote the gold status of every requirement with chart evidence. A second blind sub-agent labelled every case from the chart and the policy alone: it agreed with the author on 400 of 400 requirement labels (test), 240 of 240 (test2) and 104 of 104 (test3).
Scoring. Each model criterion is matched to a gold requirement by word overlap between its policy quote and the requirement's anchor quotes; a gold requirement's predicted status is the status of the model requirement holding its matched criteria (the worst, if several). The gold decision is computed from the gold statuses with the product's rule (any not met: not supported; any requirement undocumented: get records first; an exclusion the chart doesn't address alone: ready, listed to confirm). Letter sentences were rated by a separate blind sub-agent against the chart and the policy excerpt.
Results
| Held-out test (60, before guards) | Fresh test2 (35, after guards) | Fresh test3 (16, after cold-user fixes) | |
|---|---|---|---|
| Decision right | 45/60 | 24/35 | 14/16 |
| Unsupported requests (gold: get records / not supported) that came back ready | 10/31 | 4/21 | 1/11 |
| Supported requests that came back ready (with a letter) | 26/29 | 8/14 | 4/5 |
| Requirement status right (of gold requirements the model found) | 316/347 (91.1%) | 179/195 (91.8%) | 88/94 (93.6%) |
| Gold requirements found in the model's criteria | 347/400 (86.8%) | 195/240 (81.2%) | 94/104 (90.4%) |
| Undocumented read as not met (units / cases) | 0 / 0 | 0 / 0 | 0 / 0 |
| Not met read as undocumented (cases) | 2 | 0 | 0 |
| Letters: kept claim sentences rated unsupported (blind) | 6/375 | 2/91 | 0/46 |
| Letters: sentences that overstate the policy (blind) | 14/375 | 1/91 | 2/46 |
| Form questions answered with a verified chart quote | 151/180 | 93/105 | 39/48 |
| Time per check (wall, 4 at a time, shared gateway): p50 / p95 | 73.8 s / 148.9 s | 57.5 s / 109.7 s | 41.3 s / 105.5 s |
| Cost per check at list price: p50 / p95 | $0.0168 / $0.0456 | $0.0121 / $0.0259 | $0.0136 / $0.0270 |
| Model calls per check (median) | 20 | 17 | about 22 |
Decision confusion (gold -> tool):
- test: ready->ready 26, ready->gather 3; gather->gather 12, gather->ready 7; not supported->not supported 7, not supported->gather 2, not supported->ready 3.
- test2: ready->ready 8, ready->gather 4, ready->not supported 2; gather->gather 9, gather->ready 4, gather->not supported 1; not supported->not supported 7.
- test3: ready->ready 4, ready->gather 1; gather->gather 7; not supported->not supported 3, not supported->ready 1 (a knee revision whose vascular study shows an ankle-brachial index of 0.46: read as met, and the letter left it out).
Guard set (InterQual policies): before the rule below, 1 of 3 right (one came back ready from sub-headings of the InterQual reference, one not supported on a branch the InterQual criteria govern); after it, 3 of 3. The rule was written on these 3 cases, so this is not a held-out number.
Policy reader on the full source text (the whole CPB page or PDF text, 31k to 390k characters; one model call when the text is long): the chosen section held every gold requirement for 17 of 19 policies (123 of 128 requirements). The misses: Aetna 0084's photo and visual-field footnotes, and Aetna 0325's appendix. The accuracy numbers above use the criteria section as a coordinator would paste it.
Against the targets on wiki page 78
- Criteria status at least 95%: not met (91.1%, 91.8%, then 93.6% on the small test3), and 10% to 19% of requirements aren't found at all.
- "Undocumented vs not met" errors halved against the appeal eval's 6 of 80: met (0 of 95 cases in either direction of undocumented read as not met; 2 cases the other way on the first set).
- 0 unsupported letter sentences: not met overall (6 of 375, then 2 of 91 after the guards; 0 of 46 on test3). Of the 8, 4 carried a fact that only a physician's letter in the chart stated. Policy overstatements remain (2 of 46 on test3: "the policy requires X" where X is one accepted option).
Where it fails
- Saying ready too often (the error that matters). Before the guards, 10 of 31 unsupported requests came back ready, with a letter: the model accepted a vague mention ("tried different creams") or a letter's claim as meeting a criterion, missed exclusions the chart was silent on, and once read a steroid injection 8 weeks before surgery as outside a 12-week window. After the guards, 4 of 21: all from requirements the model never extracted (footnotes on photograph quality, "wound care continues during treatment") or one vague mention read as met.
- Too cautious after the guards. 6 of 14 supported test2 requests were held back, 2 as not supported: Cigna 0152's resection-weight table (a Schnur-scale lookup by body surface area) was misread in 3 cases.
- Coverage. 13% to 19% of requirements aren't extracted, mostly footnotes, appendix tables and qualifiers.
- Letters. A few sentences restate the policy too narrowly ("the policy requires NSAIDs" where it accepts any 3 of 6 drug classes).
Changes made after looking at held-out results (then measured on the fresh test2)
From the test errors (all general rules, none specific to a test policy): a met answer that rests only on a letter of medical necessity counts as not documented; a criterion with a measured threshold or duration needs a value in its quote (a date window may use the document's date); an exclusion answered met must quote a line that says the condition is absent (a second check); an exclusion with a date window is not met when the compared dates fall inside it. Plus, on the dev set only: a lead-in condition that applies to a list goes into each item; age limits are requirements. test2 was written, labelled and run once after these were frozen.
After test2, from two blind cold-user tests (a prior-auth coordinator persona, each run on the local site), then measured on the fresh test3: a policy that points to licensed criteria never gets ready or not supported (guard set above); a final check before any "ready" re-reads the policy for requirements the checklist missed and checks them; form answers must be stated by their chart words (one answer-check call; "No" needs words that say no; document headings count); the brief and findings work per requirement; alternatives are grouped by their list's lead-in line wherever the model put them; one re-quote before a not-met answer is called undocumented; the letter is told which options the policy accepts.
After test3 (checked on the samples, not re-measured): an item the final check adds from an "any of" list joins that list as an option. Before this, the lumbar MRI sample came back "not supported" on the gateway when the model listed only 10 of the policy's 19 accepted reasons and the final check added the rest as requirements; after it, 3 of 3 runs came back ready.
Cost and time
Median $0.012 to $0.017 per check at the gateway list price; p95 $0.026 to $0.046 (long policies with a letter). Median 41 to 74 s per check with 4 running at once on a shared gateway.
Against doing it by hand (an estimate, not a measurement). Two blind cold-user sub-agents playing a prior-auth coordinator at a 6-physician orthopedic practice estimated 25 to 40 minutes of hands-on time for a clean MRI or injection request today (40 to 60 for a joint replacement), and put the saving at about 0 minutes for an MRI or injection and 10 to 15 minutes for a joint replacement as the tool stood, rising to 8 to 12 (MRI, injection) and 15 to 20 (joint) minutes if the chart could be pasted from an EHR export or PDF and the payer's policies came from a library. Both said they would not pay yet; both named about $150 to $300 a month for a 6-physician practice once the fixes land and a BAA-covered hosted option exists.
Expected properties of the sample run (the rehearsal bundle checks these)
cpap-ready(Aetna CPB 0004; in-lab study, AHI 26.5 over 6 h of sleep) isready, with a letter of medical necessity that cites 26.5.cpap-not-supported(home sleep test index 3.6; a physician letter asserts the criteria) isnot_supported, no letter, and a not-met criterion quotes 3.6.proprietary-criteria(UnitedHealthcare panniculectomy, points to InterQual) isproprietary_criteria, no letter.bariatric-missing-eval(Cigna 0051; no mental health clearance) isgather_first.- The signed record verifies at
POST /record/verify, and fails once its decision is changed. - Every model call has a signed receipt (gateway) or an attested one (self-host).
Limits
- Synthetic charts written against real policy excerpts, one author per policy; real charts are longer, scanned and inconsistent. No prior-auth coordinator has labelled these cases; the gold is two blind model labellers in agreement.
- Accuracy is below the page-78 target: treat "ready" as "nothing obviously missing", not as a guarantee. A person checks every criterion before sending.
- Licensed criteria (InterQual, MCG) are out of scope by design; many medical-service policies depend on them.
- Decision clocks are federal maximums; state law, plan documents and contracts can be shorter.
Cold-user tests (blind, before merge)
- Test 1: "would not use it yet": a false ready on a made-up lumbar MRI case (an "ALL of" list read as "one of"), form answers the chart didn't state ("completed PT" from a referral; "No" from silence), a verdict that changed during the run, noisy missing items (unused alternatives, red flags), jargon citations. All fixed as listed above.
- Test 2 (after the fixes): still "not for real requests yet": the lumbar MRI sample flipped from ready (recorded) to not supported (live) twice; form answers withdrawn although the line held the value; "No neoplasm: Yes" reads backwards when pasted; long option lists cut off. It never called a request ready that should have waited, got 5 of its own 7 made-up cases right, and valued the letter-overclaim warning. Fixed after: the any-of join above, form answers checked with the document heading and date, exclusions worded "ruled out / present", full option lists, a "check each line" caveat under every ready.
Verdict
Would a practice pay? The "get this record first" and "not supported" calls with the exact policy line are the useful part, and the cost is cents. But the ready call is wrong often enough (4 of 21 unsupported requests after the guards) that it can only be sold as a checklist that finds what's missing, not as a gate. What's missing before a customer relies on it: requirement extraction from footnotes and tables (the M12 criterion-evidence judge on page 78, plus a table reader), labels from real coordinators on real, de-identified requests, and a second-pass check on every "ready".