Skip to content
decosa
LiveHostedSelf-hostSelf-host first for real data

Appeal a denial

A recommendation (appeal, don't, get documents first or fix the claim), the next filing deadline, a criteria checklist with quotes, and a letter when it's supported.

Held-out test36 of 40Right recommendation on whether and how to appeal (held-out test)
On production13 smedian on production (2026-09-26); slower when the service is busy
List price~$0.014 per denialmeasured, at list price

Built on: Grounding, Typed judgment, Signed record, Form filling

1. Pick a sample

Sample

Real patient data: request confidential access, or run the tool on your own hardware. The demo takes samples or made-up data only.

2. Run it

On production the sample took 13 s (median, 2026-09-26). Slower when the service is busy.

Result

The answer appears here first, then what it found, the draft, and how long it took. Sample: CPAP denied, AHI 11 with sleepiness and hypertension (synthetic).

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the deadline rules and the signed packet need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the claim denial appeal packet API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
denial-appeal-packet

Use the hosted API

# Decosa claim denial appeal packet: use the hosted API

You are wiring Decosa's denial appeal packet into this project. It takes a claim denial (EOB, remittance or letter with
its CARC/RARC codes), chart excerpts and the payer's coverage criteria, and returns an appeal packet with a signed record.
Every model call has a signed receipt. Use only what is listed below; if you need something else, stop and ask me.

The packet has:
- the next appeal level's filing deadline, with the federal rule and the arithmetic (ERISA and ACA plans, Medicare Parts A
  and B, Medicare Advantage, Medicaid managed care);
- each coverage criterion marked `met`, `not_met` or `undocumented`, with the policy quote and a chart quote;
- a recommendation: `appeal`, `do_not_appeal`, `gather_first` or `fix_claim` (or `review` when no criterion could be read);
- a draft letter, only when the recommendation is `appeal`, with every sentence checked against the chart and the policy.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **Hosted use is for synthetic cases only.** Charts are protected health information: real patients belong on a
  self-hosted box (see the self-host prompt). Say so wherever this is wired in.
- This is a drafting aid. Never auto-file, never present the letter as final, and never tell a user an appeal will win.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "denial-appeal-packet"}` returns `{"token", "expires_at", "budget"}`.
   - The demo allows a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz) and a token allowance per session (the `budget` in the session response).
   - Over a limit you get HTTP 429 with `Retry-After`.
   - A demo token runs one packet at a time (409 otherwise).
3. A packet needs about 10,400 generated tokens left in the budget before it starts (402 otherwise); it usually uses
   fewer, and only what it uses is charged.

## Endpoints
- `POST /appeal/packet` (token). Body: `{"denial": {"text", "plan_type": "erisa_group"|"aca_individual"|"medicare_ab"|"medicare_advantage"|"medicaid_mc"|"other", "level"?, "notice_date"?: "YYYY-MM-DD", "received_date"?, "claim_type"?: "post_service"|"pre_service", "urgent"?, "grandfathered"?, "two_levels"?, "contract_days"?, "amount_in_dispute"?, "payer"?}, "records": [{"kind"?: "note"|"study"|"lab"|"imaging"|"order"|"letter"|"other", "title"?, "date"?, "text"}], "policy": {"title"?, "text", "source"?}, "service"?, "patient_ref"?, "title"?, "today"?}` or `{"sample_id": "..."}`.
  - Levels by plan: ERISA `initial`, `internal_appeal`, `final_internal`; ACA individual `initial`, `final_internal`;
    Medicare A/B `initial`, `redetermination`, `reconsideration`, `alj`, `council`; Medicare Advantage `initial`, `ire`;
    Medicaid managed care `initial`, `plan_appeal`.
  - Limits: a denial of 8,000 characters; 8 records of 12,000 characters each, 40,000 in total; a policy of 16,000; 256 KB
    of JSON. Paste the policy's criteria text: URLs are not fetched.
  - The JSON response has:
    - `recommendation`: `{decision, headline, why, fixes?, missing?, not_met?, deadline_passed?}`;
    - `deadlines`: `{next: {step, deadline, arithmetic, start_is, status, days_left, rules: [{cite, url, text}]}, rows, notes, conventions}`;
    - `criteria`: `[{id, text, quote, ref, group, status, reason, evidence: [{ref, quote}], note?}]` (same `group` = alternatives);
    - `codes` (CARC/RARC with groups), `denial` (notice date, reasons, route), `letter` (text or null),
      `letter_sentences` (each with its grounding verdict and whether it was kept, flagged or removed), `packet_md`,
      `record`, `record_check`, `receipts`, `steps` and `note`.
  - With `Accept: text/event-stream` (or `"stream": true`), the events are `ready`, `codes`, a `receipt` per model call,
    `denial`, `deadlines`, `criteria`, a `criterion` per criterion, `recommendation`, a `sentence` per letter sentence,
    `letter`, then `result`, `budget` and `done`.
- `POST /appeal/deadlines` (no token, no model) `{"plan_type", "level"?, "notice_date", "received_date"?, "urgent"?, "claim_type"?, "contract_days"?, "amount_in_dispute"?}`
  returns the deadline part alone.
- `POST /record/verify` (no token) `{"record": {...}}` returns `{ok, summary, checks, first_bad}`.
- `GET /appeal/info`, `GET /appeal/samples` and `GET /attest/signing-key` need no token.

## Example: build a packet and act on the recommendation (Python, `pip install httpx`)
```python
import httpx, json, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {"denial": {"text": open("denial.txt").read(), "plan_type": "medicare_ab", "level": "initial"},
        "records": json.load(open("chart-excerpts.json")), "policy": {"title": "NCD 240.4", "text": open("criteria.txt").read()}}
r = httpx.post(f"{API}/appeal/packet", json=body, headers=H, timeout=600)
r.raise_for_status()
js = r.json()
print(js["recommendation"]["headline"], "| file by", (js["deadlines"]["next"] or {}).get("deadline"))
for c in js["criteria"]:
    print(f"{c['status']:>12}  {c['text']}")
if js["letter"]:
    open("appeal-letter.txt", "w").write(js["letter"])   # a draft: a person edits and files it
open("appeal-packet.md", "w").write(js["packet_md"])
json.dump(js["record"], open("appeal-packet.json", "w"))   # keep with the claim
```

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa claim denial appeal packet: run it yourself (containers)

You are setting up the Decosa denial appeal packet on this machine, so charts never leave it. It reads a claim denial,
chart excerpts and the payer's coverage criteria, and returns:
- the next appeal level's filing deadline with the federal rule and the arithmetic;
- each criterion marked met, not met or undocumented, with quotes;
- a recommendation (appeal, don't appeal, get documentation first, fix the claim);
- a draft letter only when the chart supports every criterion, each sentence checked.

It seals a signed packet. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/denial-appeal-packet.zip (6 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py denial-appeal-packet` (the api image carries the same bundle under /app/rehearsal/denial-appeal-packet/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py denial-appeal-packet --bundle denial-appeal-packet.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the redetermination deadline is the notice date + 5 days presumed receipt + 120 days (no model)", "the supported case is recommended for appeal", "the notice date is read from the denial and gives the same deadline"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - keep its data on a named volume;
   - bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check. The first start
   downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/appeal/info` lists every deadline rule with its citation, link and the date
   it was read. `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify
   my packets.
5. Smoke test:
   - `POST /appeal/deadlines {"plan_type":"medicare_ab","level":"initial","notice_date":"2026-09-08"}` needs no model and
     should give `2027-01-11`.
   - Get a token with `POST /demo/session {"vertical":"denial-appeal-packet"}`, then send
     `POST /appeal/packet {"sample_id": "cpap-appeal"}`. Expect `recommendation.decision` `appeal`, a letter, the deadline
     `2027-01-11`, and every receipt `attested`.
   - `{"sample_id": "cpap-no-appeal"}` should come back `do_not_appeal` with `letter: null`.
   - `POST /record/verify {"record": <record>}` should give `ok: true`.
6. Report back: the public key and key id, the smoke-test results, and how long each packet took.

Charts are protected health information. This is a drafting aid: it is not legal or medical advice, it never predicts an
outcome, and a person checks the criteria, the letter and the deadline against the plan's documents before filing.

Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds
patient charts. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVIDIA NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns

    The standard tier fits (57.6 of 192 GB). The best tier fits too.

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"denial-appeal-packet"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py denial-appeal-packet

Download the mock-data bundle (6 KB, 10 checks)expected.json

Two synthetic Medicare CPAP denials (CO-50 N386, AHI under 15) against CMS NCD 240.4. In the first, the home sleep test shows an AHI of 11.2 and the chart documents daytime sleepiness and hypertension, which meets criterion 5b: the packet must recommend an appeal, draft a letter, and compute the redetermination deadline (notice date + 5 days presumed receipt + 120 days). In the second, the AHI is 10.4 and the chart says the patient has no daytime sleepiness, insomnia, hypertension, heart disease or stroke, while a physician letter asserts the criteria are met: the packet must recommend not appealing and draft no letter. The signed packet must verify, and fail once changed.

What the rehearsal checks
  • the redetermination deadline is the notice date + 5 days presumed receipt + 120 days (no model)
  • the supported case is recommended for appeal
  • the notice date is read from the denial and gives the same deadline
  • the AHI 5 to 14 criterion is met with the study's AHI quoted
  • a letter is drafted and cites the AHI
  • the unsupported case is not recommended for appeal
  • and no letter is drafted for it
  • the signed packet verifies
  • a packet whose recommendation was changed no longer verifies
  • every model call has a signed receipt

Licence: Synthetic cases (CC0): the patients, notes, providers and payer are invented (scripts/appeal_cases.py). Policy: an excerpt of CMS NCD 240.4 (US government work, public domain). Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa claim denial appeal packet: run it yourself (containers)

You are setting up the Decosa denial appeal packet on this machine, so charts never leave it. It reads a claim denial,
chart excerpts and the payer's coverage criteria, and returns:
- the next appeal level's filing deadline with the federal rule and the arithmetic;
- each criterion marked met, not met or undocumented, with quotes;
- a recommendation (appeal, don't appeal, get documentation first, fix the claim);
- a draft letter only when the chart supports every criterion, each sentence checked.

It seals a signed packet. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/denial-appeal-packet.zip (6 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py denial-appeal-packet` (the api image carries the same bundle under /app/rehearsal/denial-appeal-packet/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py denial-appeal-packet --bundle denial-appeal-packet.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the redetermination deadline is the notice date + 5 days presumed receipt + 120 days (no model)", "the supported case is recommended for appeal", "the notice date is read from the denial and gives the same deadline"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - keep its data on a named volume;
   - bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check. The first start
   downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/appeal/info` lists every deadline rule with its citation, link and the date
   it was read. `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify
   my packets.
5. Smoke test:
   - `POST /appeal/deadlines {"plan_type":"medicare_ab","level":"initial","notice_date":"2026-09-08"}` needs no model and
     should give `2027-01-11`.
   - Get a token with `POST /demo/session {"vertical":"denial-appeal-packet"}`, then send
     `POST /appeal/packet {"sample_id": "cpap-appeal"}`. Expect `recommendation.decision` `appeal`, a letter, the deadline
     `2027-01-11`, and every receipt `attested`.
   - `{"sample_id": "cpap-no-appeal"}` should come back `do_not_appeal` with `letter: null`.
   - `POST /record/verify {"record": <record>}` should give `ok: true`.
6. Report back: the public key and key id, the smoke-test results, and how long each packet took.

Charts are protected health information. This is a drafting aid: it is not legal or medical advice, it never predicts an
outcome, and a person checks the criteria, the letter and the deadline against the plan's documents before filing.

Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds
patient charts. If I ask for it later, follow the Provide page instead of improvising.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsClaim denial appeal packet on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

Standard · the hosted demo, one 96 GB card: what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Reads the denial: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Claim denial appeal packet, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Claim denial appeal packet on my hardware

Fetch https://decosa.ai/prompts/denial-appeal-packet-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=denial-appeal-packet)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Reads the denial: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/denial-appeal-packet-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 13 s · ~$0.014 per run · 21 receipts

Loading the nightly status…

Self-host: verified 26 Sep 2026 · fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

Measured cost to run: about $0.014 per denial (hosted, 26 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): deadline 2027-01-11, cpap-appeal appeal with a letter in 11.9 s (21 attested receipts, AHI criterion quoting 11.2), cpap-no-appeal don't appeal with no letter, admin-missing-npi fix the claim, record verified; the rehearsal bundle passed 10/10. Model-server startup was not re-run.

Known limits (5)
  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.
  • Measured on 98 synthetic, templated cases written by the building agent; not on real denials or with a denials specialist's labels.
  • Federal deadline rules only; state external review, plan documents and provider contracts can differ. Deadlines other than external review are not moved off weekends.
  • Criteria come only from the policy text pasted in; non-covered indications are not mapped, and nested alternatives (an option with its own list of options) are read as one option.
  • Changes after the test2 run (a template opening line, stray reference markers stripped, alternatives worded as accepted options, source titles in the number guard) were checked on the demo samples only.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the deadline rules and the signed packet need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the claim denial appeal packet API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.
The open stack

A denial, the chart and the payer's criteria in; the appeal deadline, a criteria checklist with quotes, and a letter only when the chart supports it.

For billing, revenue-cycle and denials staff, and patient advocates. Give it the denial (EOB, remittance or letter with its CARC/RARC codes), the chart excerpts and the payer's coverage criteria. It reads the codes, computes the next appeal level's filing deadline from the notice date with the federal rule and the arithmetic (ERISA and ACA plans, Medicare Parts A and B, Medicare Advantage, Medicaid managed care), splits the policy into criteria and marks each one met, not met or undocumented, with the policy quote and a chart quote the code finds word for word. The recommendation is a rule: every criterion met means appeal; a criterion the chart contradicts means don't appeal on this record; a criterion the chart doesn't show means get the documentation first; a billing or eligibility code means fix the claim. Only an appeal gets a letter, and every sentence in it is checked against the chart, the policy and the denial by the grounding judge and a number guard; unsupported sentences are cut and listed. Everything is sealed in a signed, hash-chained packet. A drafting aid: a person checks it and files.

Deployment
Self-host first
Regulatory
Charts are protected health information under HIPAA: run it on the provider's own hardware; the hosted demo takes synthetic cases only. Deadline rules read on 26 Sep 2026 from eCFR (titles current to 24 Sep 2026): ERISA group health plans, at least 180 days after receipt of the denial to appeal (29 CFR 2560.503-1(h)(3)(i)), with plan decisions due in 72 hours (urgent), 30 days (pre-service) or 60 days (post-service) (2560.503-1(i)(2)); ACA individual coverage, the same rules with one internal level (45 CFR 147.136(b)(3)); federal external review, four months after receipt of the final internal denial, the first day of the next month when that day does not exist, moved off weekends and federal holidays (29 CFR 2590.715-2719(d)(2)(i), 45 CFR 147.136(d)(2)(i)); state processes must allow at least four months. Medicare Parts A and B: redetermination 120 days, reconsideration 180 days, ALJ hearing, Council review and court 60 days each, all from receipt presumed 5 days after the notice (42 CFR 405.942, 405.962, 405.1002, 405.1102, 405.1130/405.1136); amount in controversy $200 for an ALJ hearing and $1,960 for court in 2026, $2,000 for court from 1 Jan 2027 (Federal Register, 4 Dec 2025 and 16 Sep 2026). Medicare Advantage: reconsideration 60 days after presumed receipt (42 CFR 422.582(b)); ALJ 60 days (422.602(b)). Medicaid managed care: plan appeal 60 days from the date on the notice (42 CFR 438.402(c)(2)(ii)); state fair hearing 90 to 120 days as the state sets (438.408(f)(2)). These are federal minimums: plan documents, state programs and provider contracts can differ. Criteria come only from the policy text pasted in; the demo uses CMS National Coverage Determinations (public domain) and an invented payer policy. CARC/RARC labels are our own summaries; the official descriptions belong to X12. Not legal or medical advice, and it never predicts an outcome.
Architecture
Text description

A denial (EOB, remittance or letter with CARC/RARC codes), chart excerpts, the payer's coverage criteria and the plan type go into decosa-api. Code reads the codes and computes the next appeal level's deadline under the cited federal rule. Qwen3.8-27B reads the denial, splits the policy into criteria and checks each against the chart; code keeps a met or not-met answer only when its chart quote is found word for word. A rule turns the checklist into a recommendation: appeal, don't appeal, get documentation first, or fix the claim. Only when every criterion is met, the model drafts a letter, and the grounding judge plus a number guard check each sentence; unsupported sentences are cut. Outputs: the recommendation, the deadline, the checklist, the letter and a signed hash-chained packet. Hosted calls get gateway-signed receipts; a self-hosted box signs with its own key.

Architecture

At a glance

What it gives you
A recommendation (appeal, don't appeal, get documentation first, fix the claim), the next level's filing deadline with its CFR rule and arithmetic, each criterion met, not met or undocumented with quotes, a draft letter when the chart supports it, a Markdown packet and a signed record.
What it does not do
It does not file, fax or submit to a portal, predict outcomes, read plan documents or provider contracts, know state rules, or fetch payer policies (paste the criteria). Non-covered indications are not mapped. Scanned charts need OCR first.
Data retention
Nothing kept on the server. Denials, charts and packets live in memory for the request; logs carry counts only. You keep the signed packet with the claim.
What leaves the box (hosted demo)
The denial, chart excerpts and policy go to Qwen3.8-27B through the Decosa API, whose receipts hold hashes, not text. Self-hosted, nothing leaves. The hosted demo is for synthetic cases only.
Model calls per packet
One denial read, one criteria map, one check per criterion; with a letter, one draft and one grounding call per letter sentence: about 1 to 21 calls on the demo cases.
Typical run cost
A fraction of a cent at the gateway list price for the CPAP case with a letter (a couple of dozen calls); less when it says don't appeal. Each run shows its own measured cost.
Deadlines
Federal minimums for ERISA and ACA plans, Medicare A/B, Medicare Advantage and Medicaid managed care, from the notice date (Medicare presumes receipt 5 days later). Give the receipt date and any shorter provider-contract window; the earlier date wins.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 48 GB card

    The same pipeline on a smaller mixture-of-experts model. Faster and cheaper; accuracy on this task unknown.

    Models
    • Gemma 4 26B A4B (instruction-tuned)
    Hardware
    1x L40S or RTX 6000 Ada 48 GB (not measured)
    Quality evidence
    • recommendation and criteria accuracynot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
  • In the hosted demo

    Standard

    the hosted demo, one 96 GB card

    Qwen3.8-27B reads the denial, maps and checks the criteria and drafts the letter; the deadline, the quote checks and the recommendation are plain code. Every model call receipted.

    Models
    • Qwen3.8-27B (NVIDIA NVFP4)
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • fresh held-out set (test2, 40 synthetic cases, run once): recommendation right36/40; all 4 misses said don't appeal where the gold says get documentation firstdecosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26, gateway route; prompts frozen on an 18-case dev set; half the cases use phrasing never seen while writing the prompts
    • planted unsupported cases (test2 / test): no appeal and no letter26/26 / 21/22 (the one miss: the model read a 53-day gap as 23; date windows are now checked in code, which test2 measures)decosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26
    • supported cases (test2 / test): appeal with a letter11/11 / 14/15decosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26
    • criteria status right, per policy requirement (test2 / test)126/133 / 125/133; every requirement was found in the policy (133/133)decosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26
    • deadline (next level and date) right / notice date read from the denial80/80 / 40/40decosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26; gold deadlines computed by separate code with hard-coded holidays
    • letter sentences kept after the check that state a clinical fact the chart doesn't support (read by the building agent)0/245; 2 kept sentences overstated the policy (said PSG is required where the NCD lists alternatives); 32 of 260 draft sentences were cut or marked [CHECK] by the checkdecosa-api docs/evals/denial-appeal-packet.md, test and test2 letters, 2026-09-26
    • real, de-identified denials rated by a denials specialistnot measured yet
    Latency
    measured on our server under a shared gateway: seconds per packet, under a minute at the slow end; a packet with a letter makes a couple of dozen model calls
    Verification
    Proof: strongHosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.
  • Best

    DeepSeek-V4-Flash on two more cards

    A larger model for long charts and dense payer policies.

    Models
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    Hardware
    2x RTX PRO 6000 96 GB
    Quality evidence
    • recommendation and criteria accuracynot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyNot a hosted model: calls are attested by the box's key only.
  • Needs more compute

    Wanted: the best setup

    two large judges from different families

    DeepSeek-V4-Flash and GLM-5.3-Flash each read the file, and a finding stands when they agree; disagreements go to the reviewer. Patient records stay on your own hardware, never on community providers. Not served yet.

    Models
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    • GLM-5.3-Flash
    Hardware
    Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash plus 2x 96 GB for GLM-5.3-Flash, or one Mac Studio with 512 GB holding both 4-bit builds (156 + 165 GB, sizes from our Mac; not run together yet). Estimate.
    Quality evidence
    • recommendation and criteria accuracynot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
After the session
Reads the denial (notice date, reasons, service), splits the payer policy into criteria, checks each criterion against the chart, drafts the letter when the chart supports it, and judges every letter sentence (the grounding judge)Qwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
Lite tier: the same pipeline on a 48 GB cardGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
25.2B (3.8B active)No proof yetSelf-host only
Best tier: a larger model for long charts and dense payer policiesDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active) · 192 GBNo proof yetSelf-host only
Other
Second judge, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only

Measured

Does it say no when it should?

98 synthetic cases from a generator with a structured truth: Medicare CPAP, cochlear-implant and bariatric denials against CMS coverage determinations, lumbar MRI denials against an invented payer policy, and billing denials. Some charts support every criterion; others contradict one or leave one out. Prompts were written on an 18-case dev set; a 40-case test set was run once, a date check was added in code, and a fresh 40-case set was run once.

Unsupported cases with no appeal and no letter
47 of 48test2 26/26, test 21/22
Supported cases with an appeal and a letter
25 of 26test2 11/11, test 14/15
Deadlines right
80 of 80next level and date, test and test2
Cost per packet with a letter
about $0.01421 model calls, 34,908 tokens on the CPAP demo case, gateway list price

Where it fails

It reads a criterion the chart leaves out as not met rather than undocumented (6 of 8 wrong recommendations across both sets): the advice becomes don't appeal where it should be get the documentation first. Once, before the date check moved into code, it read a 53-day gap as 23 and recommended an appeal the chart did not support.

What it does not show

The cases are templated and synthetic, written by the same author as the prompts. Real charts are longer, scanned and messier, and commercial policies are denser than these excerpts. Measure it on your own closed denials before relying on it.

Source: decosa-api docs/evals/denial-appeal-packet.md, 26 Sep 2026

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    Intake, the codes, the deadline rules, the quote checks, the recommendation rule, grounding, signing and the HTTP API (/appeal/*). No GPU. Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server.

  • 1x L40S / RTX 6000 Ada 48 GB

    Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite).

  • CPU only Fits

    The deadline rules (POST /appeal/deadlines), the codes, the signed packet and verification need no GPU; reading the chart needs the model.

Latency per lane

  • one packet, busy shared gateway12.9 s

    Measureddecosa-api docs/evals/denial-appeal-packet.md, test2 median on our server 2026-09-26 (max 32.7 s; 41.3 s median on the earlier, busier test run), gateway route

  • one packet with a letter, self-hosted direct route11.9 s

    Measuredmeasured on our server 2026-09-26, self-host sandbox, cpap-appeal (21 model calls)

  • deadline rules only50 ms

    Estimateestimate: no model call

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

denial-appeal-packet/assemble-prompt.md192 lines
# Assemble the Decosa claim denial appeal packet on this machine

You are setting up a self-hosted denial appeal packet on this Linux machine for a provider's billing, revenue-cycle or
denials team. It reads a claim denial (EOB, remittance or letter with CARC/RARC codes), the chart excerpts and the payer's
coverage criteria, and returns:
- the next appeal level's filing deadline, with the federal rule and the arithmetic (ERISA and ACA plans, Medicare Parts A
  and B, Medicare Advantage, Medicaid managed care);
- each coverage criterion marked met, not met or undocumented, with the policy quote and a chart quote found word for word;
- a recommendation: appeal, don't appeal on this record, get documentation first, or fix the claim;
- a draft letter, only when the chart supports every criterion, with every sentence checked against the chart and policy;
- a signed, hash-chained packet.

Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- Charts are protected health information. Keep everything on this machine: the model route stays local (`direct`), and
  nothing goes to a hosted service.
- This is a drafting aid. It is not legal or medical advice, it never predicts an outcome, and a person checks the
  criteria, the letter and the deadline against the plan's documents, then files.

Repeat both points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/denial-appeal-packet.zip (6 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py denial-appeal-packet` (the api image carries the same bundle under /app/rehearsal/denial-appeal-packet/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py denial-appeal-packet --bundle denial-appeal-packet.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the redetermination deadline is the notice date + 5 days presumed receipt + 120 days (no model)", "the supported case is recommended for appeal", "the notice date is read from the denial and gives the same deadline"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
   - Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main`. On a 48 GB card, also
     set `LLM_MAX_LEN=32768`. Not measured.
   - Under 48 GB: stop and tell me it will not fit.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the
   NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories. Then run
   `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 60 GB of free disk.

## 2. Get the images

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`.

If a pull fails, build from source once the `decosa-api` source is published:
- clone it;
- in the clone, run `docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`;
- run `docker compose build llm` from its compose file.

If neither works, stop and tell me.

## 3. Write the compose file

Create `~/decosa-appeal/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these packets, e.g. Example Clinic revenue cycle>"
```

Create `~/decosa-appeal/docker-compose.yml` with exactly these services:

```yaml
name: decosa-appeal
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"                # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "200000"        # per session; one packet needs up to about 10,000 generated tokens
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_APPEAL_MAX_CONCURRENT: "3"
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

Use the named volume `decosa-data` exactly as written. A host bind mount owned by root makes the API fail on
`/data/keys.sqlite`.

Run `docker compose up -d`, then poll `docker compose ps` until both services are healthy. The LLM takes 5-10 minutes
the first time. `curl -s localhost:8445/healthz` should show `"llm": true` (`asr` is false here, which is fine).

## 4. Smoke test

The deadline rules need no model:

```bash
API=localhost:8445
curl -s $API/appeal/deadlines -H 'content-type: application/json' \
  -d '{"plan_type":"medicare_ab","level":"initial","notice_date":"2026-09-08"}' | jq '.next | {step, deadline, arithmetic}'
```

Pass if the deadline is `2027-01-11` (5 days' presumed receipt plus 120 days, 42 CFR 405.942).

Then the packets (synthetic cases; the policies are CMS coverage determinations or an invented payer policy):

```bash
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"denial-appeal-packet"}' | jq -r .token)
for s in cpap-appeal cpap-no-appeal admin-missing-npi; do
  curl -s $API/appeal/packet -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
    -d "{\"sample_id\":\"$s\"}" > /tmp/appeal-$s.json
  jq -c '{decision: .recommendation.decision, deadline: .deadlines.next.deadline, letter: (.letter != null), receipts: (.receipts|length), statuses: ([.receipts[].status]|unique)}' /tmp/appeal-$s.json
done
jq -r '.criteria[] | "\(.status)\t\(.text)"' /tmp/appeal-cpap-appeal.json
jq '{record}' /tmp/appeal-cpap-appeal.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```

Pass if:
- `cpap-appeal` is `appeal` with a letter, deadline `2027-01-11`, and the AHI criterion is met with a quote containing
  "11.2";
- `cpap-no-appeal` is `do_not_appeal` with no letter (the chart says there are no qualifying symptoms);
- `admin-missing-npi` is `fix_claim` with no criteria and no letter;
- every receipt has `"status": "attested"`, and the record verifies (`ok: true`).

## 5. Point the app at the local API

- Base URL: `http://localhost:8445`. For a web app, set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`. Add other origins to
  `DECOSA_CORS_ORIGINS`.
- `POST /appeal/packet` takes `{denial: {text, plan_type, level?, notice_date?, received_date?, claim_type?, urgent?,
  contract_days?}, records: [{kind?, title?, date?, text}], policy: {title, text}, service?, patient_ref?}`. It returns JSON,
  or streams Server-Sent Events when asked with `Accept: text/event-stream`.
- `GET /appeal/info` lists each deadline rule with its citation, link and the date it was read, the CARC/RARC groups and
  the limits.
- Paste the policy's criteria section, not a URL: nothing is fetched. Give `received_date` when you know it (ERISA and ACA
  count from receipt), and `contract_days` when a provider contract sets a shorter dispute window.
- The server stores nothing. Keep each signed packet (JSON) with the claim. Anyone can re-check it with
  `POST /record/verify` against the key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1. For other users on the LAN, put a TLS reverse proxy with authentication in front and set
  `DECOSA_TRUSTED_PROXIES`.

## 6. Keep the direct route

`DECOSA_LLM_ROUTE=gateway` would send prompts to the hosted Decosa
API, and those prompts contain the chart and the denial.
Never use it for real patients. At most, use it for synthetic training material.

Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

18 laws, rules and guidance pages cited; 18 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A denial, the chart and the payer's criteria in; the appeal deadline, a criteria checklist with quotes, and a letter only when the chart supports it.
Who it's for
Billing, revenue-cycle and denials staff at providers and DME suppliers, and patient advocates.
Where it runs
Self-host for real patients (hosted demo: synthetic cases only)
Key numbers
  • 34 of 40 Held-out phrasing cases, recommendation right (held out, n = 40)
  • 36 of 40 Recommendation right, fresh held-out set (test split, n = 40)
  • 26 of 26 Unsupported cases with no appeal and no letter (test split, n = 26)
  • 12.9 s Median end-to-end run, hosted (QA sweep 2026-09-26)
All results, datasets and caveats
Models
Qwen3.8-27B
Where
Self-host for real patients (hosted demo: synthetic cases only)
Checks
Receipt per model call; every criterion quote found word for word in the chart; every letter sentence grounded; signed hash-chained packet
Output
Notes, reports and drafts · Signed record or verdict
Data
Patient data (PHI)
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

Does it always write an appeal letter?

No. It writes one only when the chart supports every criterion in the policy. If the chart contradicts a criterion it recommends not appealing on this record; if the chart leaves one out it lists what to get first. On a fresh set of 26 unsupported synthetic cases it drafted no letter for any of them.

How is the appeal deadline worked out?

From the notice date under the federal rule for the plan type and level, with the arithmetic shown: for example Medicare redetermination is 120 days after receipt, presumed 5 days after the notice (42 CFR 405.942). Plan documents, state programs and provider contracts can set other windows; give a contract window and the earlier date wins.

Can the letter invent clinical facts?

Each criterion answer needs a chart quote found word for word, and every letter sentence is checked against the chart, the policy and the denial by the grounding judge and a number guard. Unsupported sentences are cut and listed. In 245 kept sentences from the test letters, none stated a clinical fact the chart did not support.

Which payer policies does it use?

The criteria text you paste; it fetches nothing. The demo uses CMS National Coverage Determinations, which are public domain, and an invented commercial policy.

Where does patient data go?

Self-hosted, nothing leaves your hardware. The hosted demo takes synthetic cases only; there, the text goes to Qwen3.8-27B through our gateway, whose receipts hold hashes, not text. Nothing is kept on the server.

Will it tell me an appeal will win?

No. It never predicts an outcome. It is a drafting aid: a person checks the criteria, the letter and the deadline, then files.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Claim denial appeal packet

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.