Skip to content
decosa

102 · Healthcare · preview

Prior-auth pre-check and packet

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 28 Sep 2026Eval write-up (decosa-api, access required)

  • Decision right, latest fresh set14 of 16test splitn = 164 real policies never used while building; run once after the cold-user fixes
  • Requests that should have waited, called ready1 of 11test splitn = 11latest set; earlier sets 4 of 21 and 10 of 31
  • Decision right, earlier fresh set24 of 35test splitn = 357 real policies; run once after the guards
  • Decision right, first held-out set45 of 60test splitn = 6012 real policies; run once before the guards
  • Undocumented read as not met0 of 111test splitn = 111all three sets; the appeal engine's main error was 6 of 80
  • Letter sentences rated unsupported8 of 512test splitn = 512blind reviewer, all three sets; 0 of 46 on the latest

Dataset

Synthetic charts written blind against 23 real published payer policies and sections (Aetna, Cigna, UnitedHealthcare, CMS LCDs); charts CC0, policies quoted with their sources

Caveats

  • Synthetic charts; real charts are longer and messier.
  • Each set was run once; fixes were made after each and measured on the next, fresh set. The latest set is small (16).
  • Requirement status right 91% to 94%, below the 95% target; 10% to 19% of requirements not found.

Nightly smoke check

Loading the nightly status…

Result
partial
Run
28 Sep 2026
Latency, this run
n/a
p50 over passed runs
41 s
Receipts
n/a
Model calls
n/a
Tokens
n/a
Cost per run
$0.014

Self-host verification

Verified on 28 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, run with a named data volume, direct route to the local Qwen3.8-27B, local signing; torn down after

The rehearsal bundle passed 8/8 (CPAP ready with a letter citing the AHI 26.5, CPAP not supported with no letter, the not-met criterion quoting 3.6, the record verifies and fails once changed, every receipt attested) in 28.2 s; the lumbar MRI sample came back ready in 17.5 s and the InterQual sample can't-check in 2.9 s, all receipts attested. Model-server startup was not re-run.

Rehearsal bundle: prior-auth-check.zip (6 KB, 8 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Time and cost from the latest held-out run (4 checks at a time on the shared gateway); replaced by production measurements after launch.
  • Accuracy is below the target set for this tool (95% of requirements right): a person checks every criterion before sending.
  • Synthetic charts written against real policy excerpts; not measured on real charts or with coordinators' labels.
  • Licensed criteria (InterQual, MCG) are out of scope; a policy that points to them gets can't check.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Picks the criteria section of a long policy, splits the policy into requirements, alternatives and exclusions, checks each against the chart, re-checks every not-met answer (and every exclusion answered met), answers the payer's form questions, drafts the letter of medical necessity when the chart supports every criterion, and judges every letter sentence (the grounding judge)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Does it say "don't send" when it should?

  • Decision right, latest set: 14 of 16 (earlier sets: 24 of 35, 45 of 60)
  • Unsupported requests it called ready: 1 of 11 (earlier sets: 4 of 21, 10 of 31)
  • Supported requests it called ready: 4 of 5 (earlier sets: 8 of 14, 26 of 29)
  • Undocumented read as not met: 0 of 111 cases (the appeal engine's main error was 6 of 80)
  • Cost per check: about $0.014 (median on the latest set, gateway list price; p95 $0.027)

Source: decosa-api docs/evals/prior-auth-check.md, 28 Sep 2026

Lite · one 48 GB card (1)
  • decision and criteria accuracy: not measured yet
Standard · the hosted demo, one 96 GB card (5)
  • latest fresh held-out set (16 synthetic charts, 4 real policies, run once after the cold-user fixes): decision right: 14/16; unsupported requests called ready 1/11; supported called ready 4/5decosa-api docs/evals/prior-auth-check.md, measured on our server 2026-09-28, gateway route
  • earlier fresh set (35 charts, 7 policies, after the guards) / first set (60 charts, 12 policies, before them): 24/35 (unsupported called ready 4/21) / 45/60 (10/31)decosa-api docs/evals/prior-auth-check.md
  • requirement status right (of requirements the model found): 91.1%, 91.8%, 93.6%; 81% to 90% of gold requirements founddecosa-api docs/evals/prior-auth-check.md
  • undocumented read as not met (the appeal engine's main error, 6 of 80 cases): 0 of 111 casesdecosa-api docs/evals/prior-auth-check.md
  • letter sentences rated unsupported by a blind reviewer: 6/375, 2/91, 0/46decosa-api docs/evals/prior-auth-check.md
Best · DeepSeek-V4-Flash on two more cards (1)
  • decision and criteria accuracy: not measured yet
Wanted · two large judges from different families (1)
  • decision and criteria accuracy: not measured yet

How we measure · All tools