Skip to content
decosa

60 · Healthcare · Finance and insurance · live

Claim denial appeal packet

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)

  • Recommendation right, fresh held-out set36 of 40test splitn = 40test2, run once after the date check moved into code; 4 misses said don't appeal where gold says get documentation first
  • Unsupported cases with no appeal and no letter26 of 26test splitn = 26test2; the earlier test set: 21 of 22
  • Supported cases with an appeal and a letter11 of 11test splitn = 11test2; the earlier test set: 14 of 15
  • Recommendation right, first test set36 of 40test splitn = 40Run once before the date check; one wrong appeal from a date misread
  • Criteria status right per policy requirement126 of 133test splitn = 133test2; test 125 of 133; every requirement was found in the policy
  • Deadline right (next level and date)80 of 80test splitn = 80test and test2; gold computed by separate code
  • Notice date read from the denial40 of 40test splitn = 40cases where the request did not give it
  • Kept letter sentences with an unsupported clinical fact0 of 245test splitn = 245Read by the building agent; 2 kept sentences overstated the policy
  • Held-out phrasing cases, recommendation right34 of 40held outn = 40Sentences never seen while the prompts were written: test 18 of 20, test2 16 of 20 (template phrasing: 18 of 20 and 20 of 20)

Dataset

Synthetic denials, chart excerpts and policies from scripts/appeal_cases.py: dev 18 cases, test 40 and test2 40 (half with held-out phrasing), plus 5 demo samples; CMS NCD excerpts (public domain) and an invented commercial policy.

Caveats

  • Synthetic, templated cases written by the same author as the prompts; real charts and commercial policies are longer and messier. These numbers do not predict accuracy on a provider's denials.
  • The unsupported-sentence count is the building agent's own reading of the letters, not an independent rater's.
  • After the first test run, date windows were moved into code; the fresh test2 set measures that change. Later small changes were checked on the demo samples only.
  • Four policies and five denial types; the gold for undocumented versus not met is a judgement call in a few chart phrasings.
  • Federal deadline rules only; state programs, plan documents and provider contracts are not modelled.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
26 Sep 2026
Latency, this run
n/a
p50 over passed runs
13 s
Receipts
21
Model calls
n/a
Tokens
n/a
Cost per run
$0.014

Self-host verification

Verified on 26 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): deadline 2027-01-11, cpap-appeal appeal with a letter in 11.9 s (21 attested receipts, AHI criterion quoting 11.2), cpap-no-appeal don't appeal with no letter, admin-missing-npi fix the claim, record verified; the rehearsal bundle passed 10/10. Model-server startup was not re-run.

Rehearsal bundle: denial-appeal-packet.zip (6 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.
  • Measured on 98 synthetic, templated cases written by the building agent; not on real denials or with a denials specialist's labels.
  • Federal deadline rules only; state external review, plan documents and provider contracts can differ. Deadlines other than external review are not moved off weekends.
  • Criteria come only from the policy text pasted in; non-covered indications are not mapped, and nested alternatives (an option with its own list of options) are read as one option.
  • Changes after the test2 run (a template opening line, stray reference markers stripped, alternatives worded as accepted options, source titles in the number guard) were checked on the demo samples only.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Reads the denial (notice date, reasons, service), splits the payer policy into criteria, checks each criterion against the chart, drafts the letter when the chart supports it, and judges every letter sentence (the grounding judge)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Does it say no when it should?

  • Unsupported cases with no appeal and no letter: 47 of 48 (test2 26/26, test 21/22)
  • Supported cases with an appeal and a letter: 25 of 26 (test2 11/11, test 14/15)
  • Deadlines right: 80 of 80 (next level and date, test and test2)
  • Cost per packet with a letter: about $0.014 (21 model calls, 34,908 tokens on the CPAP demo case, gateway list price)

Source: decosa-api docs/evals/denial-appeal-packet.md, 26 Sep 2026

Lite · one 48 GB card (1)
  • recommendation and criteria accuracy: not measured yet
Standard · the hosted demo, one 96 GB card (7)
  • fresh held-out set (test2, 40 synthetic cases, run once): recommendation right: 36/40; all 4 misses said don't appeal where the gold says get documentation firstdecosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26, gateway route; prompts frozen on an 18-case dev set; half the cases use phrasing never seen while writing the prompts
  • planted unsupported cases (test2 / test): no appeal and no letter: 26/26 / 21/22 (the one miss: the model read a 53-day gap as 23; date windows are now checked in code, which test2 measures)decosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26
  • supported cases (test2 / test): appeal with a letter: 11/11 / 14/15decosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26
  • criteria status right, per policy requirement (test2 / test): 126/133 / 125/133; every requirement was found in the policy (133/133)decosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26
  • deadline (next level and date) right / notice date read from the denial: 80/80 / 40/40decosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26; gold deadlines computed by separate code with hard-coded holidays
  • letter sentences kept after the check that state a clinical fact the chart doesn't support (read by the building agent): 0/245; 2 kept sentences overstated the policy (said PSG is required where the NCD lists alternatives); 32 of 260 draft sentences were cut or marked [CHECK] by the checkdecosa-api docs/evals/denial-appeal-packet.md, test and test2 letters, 2026-09-26
  • real, de-identified denials rated by a denials specialist: not measured yet
Best · DeepSeek-V4-Flash on two more cards (1)
  • recommendation and criteria accuracy: not measured yet
Wanted · two large judges from different families (1)
  • recommendation and criteria accuracy: not measured yet

How we measure · All tools