Skip to content
decosa
LiveHostedSelf-hostSelf-host first for real data

Screen a claim for Federal IDR

An eligibility screen for Federal IDR with the dated rule and arithmetic for each check, and a draft offer brief checked against the documents.

Held-out test61 of 61Eligibility verdicts right from the documents (held-out, templated)
On production52 smedian on production (2026-09-27); slower when the service is busy
List price~$0.020 per claimmeasured, at list price

Built on: Numeric grounding, Grounding, Typed judgment, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the clocks, the screen from structured facts and the signed packet need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the no surprises act idr packet and eligibility screen API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
nsa-idr-packet

Use the hosted API

# Decosa No Surprises Act IDR packet: use the hosted API

You are wiring Decosa's No Surprises Act IDR packet into this project. It takes an out-of-network dispute's documents (the
EOB or remittance, the open negotiation notice, any notice of IDR initiation, and supporting records) and returns an
eligibility screen and an offer packet with a signed record. Every model call has a signed receipt. Use only what is listed
below; if you need something else, stop and ask me.

The packet has:
- the claim facts read from the documents, each with a quote and the line it came from;
- an eligibility screen for the Federal IDR process: `likely_eligible`, `likely_ineligible` or `needs_review`, with each
  check (`pass`, `fail`, `unknown`, `review`, `na`), its rule, the date the rule was read and the business-day arithmetic;
- the offer as a dollar amount and a percentage of the QPA, and the evidence for each factor the arbiter must consider;
- a draft brief, only when no check fails, with every sentence checked against the documents.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **Hosted use is for synthetic disputes only.** EOBs and clinical records are protected health information: real claims
  belong on a self-hosted box (see the self-host prompt). Say so wherever this is wired in.
- This is a preparation aid. Never auto-file, never present the brief as final, and never tell a user they will win: a
  person files, and the IDR entity decides eligibility and payment.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "nsa-idr-packet"}` returns `{"token", "expires_at", "budget"}`.
   - The demo allows a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz) and a token allowance per session (the `budget` in the session response).
   - Over a limit you get HTTP 429 with `Retry-After`.
   - A demo token runs one packet at a time (409 otherwise).
3. A full packet needs about 9,500 generated tokens left in the budget before it starts, screen mode about 1,220 (402
   otherwise); only what it uses is charged.

## Endpoints
- `POST /idr/packet` (token). Body: `{"documents": [{"kind"?: "eob"|"remittance"|"on_notice"|"idr_notice"|"correspondence"|"contract_history"|"provider_profile"|"market_data"|"clinical_note"|"qpa_disclosure"|"other", "title"?, "date"?, "text"}], "claim"?: {...facts you already hold...}, "batch"?: [{"claim_id"?, "service_code", "date_of_service", "npi"?, "tin"?, "plan_id"?, "payer"?, "patient_ref"?, "claim_form"?}], "prior_determinations"?: [{"determination_date", "same_other_party"?, "same_or_similar"?, "batched"?}], "offer"?: {"amount", "qpa"?}, "side"?: "provider"|"plan", "facility_claim"?, "mode"?: "packet"|"screen", "title"?, "today"?}` or `{"sample_id": "..."}`.
  - `mode: "screen"` makes one model call and stops after the screen: use it on every remittance before opening a dispute.
  - Facts in `claim` win over what the model reads; disagreements come back in `conflicts`.
  - Limits: 8 documents of 12,000 characters each, 48,000 in total; 60 batch lines; 256 KB of JSON. Paste text: URLs are
    not fetched.
  - The JSON response has `screen` (`verdict`, `checks`, `failed`, `open`, `clocks`, `fee`, `notes`), `claim` and
    `provenance`, `figures` (`offer`, `qpa`, `pct`), `factors`, `left_out`, `decision`
    (`draft`|`draft_with_open_items`|`no_brief`|`screen_only`), `brief` (text or null), `brief_sentences`, `packet_md`,
    `record`, `record_check`, `receipts`, `steps` and `note`.
  - With `Accept: text/event-stream` (or `"stream": true`), the events are `ready`, a `receipt` per model call, `facts`,
    `screen`, `batch_judgment` (sometimes), `factors`, a `sentence` per brief sentence, `brief`, then `result`, `budget`
    and `done`.
- `POST /idr/screen` (no token, no model) takes the facts (`claim`, `batch`, `prior_determinations`, `today`) and returns
  the screen alone. `POST /idr/clock` (no token, no model) returns the business-day windows for `payment_received`,
  `on_sent` and `idr_initiated`.
- `POST /record/verify` (no token) `{"record": {...}}` returns `{ok, summary, checks, first_bad}`.
- `GET /idr/info` (every rule with its citation, link, date read and status), `GET /idr/samples` and
  `GET /attest/signing-key` need no token.

## Example: screen every remittance, build a packet only for the eligible ones (Python, `pip install httpx`)
```python
import httpx, json, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
docs = [{"kind": "remittance", "text": open("remittance.txt").read()}, {"kind": "on_notice", "text": open("open-negotiation.txt").read()}]
s = httpx.post(f"{API}/idr/packet", json={"documents": docs, "mode": "screen"}, headers=H, timeout=300).json()
print(s["screen"]["headline"])
for c in s["screen"]["checks"]:
    print(f"{c['status']:>8}  {c['label']}: {c['why']}")
if s["screen"]["verdict"] != "likely_ineligible":
    p = httpx.post(f"{API}/idr/packet", json={"documents": docs, "offer": {"amount": 2350.00}, "side": "provider"}, headers=H, timeout=900).json()
    print(p["decision_is"], p["figures"]["pct"], "% of the QPA")
    if p["brief"]:
        open("idr-brief.txt", "w").write(p["brief"])   # a draft: a person edits and files it
    open("idr-packet.md", "w").write(p["packet_md"])
    json.dump(p["record"], open("idr-packet.json", "w"))   # keep with the claim
```

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa No Surprises Act IDR packet: run it yourself (containers)

You are setting up the Decosa No Surprises Act IDR packet on this machine, so claim files never leave it. It reads a
dispute's documents (the EOB or remittance, the open negotiation notice, any notice of IDR initiation, supporting records)
and returns:
- the claim facts with quotes;
- a Federal IDR eligibility screen, each check with its rule, the date read and the business-day arithmetic;
- the offer as a percentage of the QPA and the evidence for each factor the arbiter must consider;
- a draft brief only when no check fails, each sentence checked.

It seals a signed packet. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/nsa-idr-packet.zip (4 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py nsa-idr-packet` (the api image carries the same bundle under /app/rehearsal/nsa-idr-packet/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py nsa-idr-packet --bundle nsa-idr-packet.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the IDR window opens on business day 31 of open negotiation, 6 Oct 2026 (no model)", "and closes on business day 34, 9 Oct 2026", "a self-insured plan that opted into New Jersey's law fails the state-law check (no model)"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - keep its data on a named volume;
   - bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check. The first start
   downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/idr/info` lists every rule with its citation, link, the date it was read and
   whether it applies yet. `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to
   verify my packets.
5. Smoke test:
   - `POST /idr/clock {"claim":{"payment_received":"2026-08-13","on_sent":"2026-08-24"}}` needs no model and should give
     the window `2026-10-06` to `2026-10-09`.
   - Get a token with `POST /demo/session {"vertical":"nsa-idr-packet"}`, then send
     `POST /idr/packet {"sample_id": "er-anesthesia-pa"}`. Expect `screen.verdict` `likely_eligible`, `figures.pct` 198.48,
     a brief, the billed-charges and Medicare lines in `left_out`, and every receipt `attested`.
   - `{"sample_id": "late-initiation-az"}` should come back `likely_ineligible` with `failed: ["initiation"]` and
     `brief: null`.
   - `POST /record/verify {"record": <record>}` should give `ok: true`.
6. Report back: the public key and key id, the smoke-test results, and how long each packet took.

EOBs are protected health information. This is a preparation aid: it is not legal advice, it never predicts an outcome, a
person checks the screen, the clocks and the brief before filing, and the IDR entity decides.

Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds claim
files. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVIDIA NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns

    The standard tier fits (57.6 of 192 GB). The best tier fits too.

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"nsa-idr-packet"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py nsa-idr-packet

Download the mock-data bundle (4 KB, 15 checks)expected.json

Two synthetic out-of-network disputes. The first (Pennsylvania, fully insured plan, emergency anesthesia) had its open negotiation notice sent on 24 Aug 2026, so the 4-business-day IDR window is 6 to 9 Oct 2026 (Labor Day skipped). The packet must read the facts from the EOB with quotes, screen it likely eligible, express the $2,350.00 offer as 198.48% of the $1,184.00 QPA, leave out the billed-charges, Medicare and FAIR Health lines (factors the arbiter may not consider), draft a brief, and seal a signed record that fails once changed. The second (Arizona) was initiated on 20 Aug 2026 after its window closed on 18 Aug: the screen must find it likely ineligible on the initiation check, from the documents alone.

What the rehearsal checks
  • the IDR window opens on business day 31 of open negotiation, 6 Oct 2026 (no model)
  • and closes on business day 34, 9 Oct 2026
  • a self-insured plan that opted into New Jersey's law fails the state-law check (no model)
  • the eligible dispute is screened likely eligible from its documents
  • the QPA is read from the EOB with its quote
  • the offer is expressed as a percentage of the QPA
  • the billed-charges line is left out as a prohibited factor
  • a brief is drafted
  • and it does not mention Medicare rates or billed charges
  • the signed packet verifies
  • a packet whose verdict was changed no longer verifies
  • the late dispute is screened likely ineligible
  • on the initiation check
  • and no brief is drafted
  • every model call has a signed receipt

Licence: Synthetic (CC0): invented providers, plans, patients and figures (scripts/idr_samples.py). Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa No Surprises Act IDR packet: run it yourself (containers)

You are setting up the Decosa No Surprises Act IDR packet on this machine, so claim files never leave it. It reads a
dispute's documents (the EOB or remittance, the open negotiation notice, any notice of IDR initiation, supporting records)
and returns:
- the claim facts with quotes;
- a Federal IDR eligibility screen, each check with its rule, the date read and the business-day arithmetic;
- the offer as a percentage of the QPA and the evidence for each factor the arbiter must consider;
- a draft brief only when no check fails, each sentence checked.

It seals a signed packet. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/nsa-idr-packet.zip (4 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py nsa-idr-packet` (the api image carries the same bundle under /app/rehearsal/nsa-idr-packet/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py nsa-idr-packet --bundle nsa-idr-packet.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the IDR window opens on business day 31 of open negotiation, 6 Oct 2026 (no model)", "and closes on business day 34, 9 Oct 2026", "a self-insured plan that opted into New Jersey's law fails the state-law check (no model)"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - keep its data on a named volume;
   - bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check. The first start
   downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/idr/info` lists every rule with its citation, link, the date it was read and
   whether it applies yet. `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to
   verify my packets.
5. Smoke test:
   - `POST /idr/clock {"claim":{"payment_received":"2026-08-13","on_sent":"2026-08-24"}}` needs no model and should give
     the window `2026-10-06` to `2026-10-09`.
   - Get a token with `POST /demo/session {"vertical":"nsa-idr-packet"}`, then send
     `POST /idr/packet {"sample_id": "er-anesthesia-pa"}`. Expect `screen.verdict` `likely_eligible`, `figures.pct` 198.48,
     a brief, the billed-charges and Medicare lines in `left_out`, and every receipt `attested`.
   - `{"sample_id": "late-initiation-az"}` should come back `likely_ineligible` with `failed: ["initiation"]` and
     `brief: null`.
   - `POST /record/verify {"record": <record>}` should give `ok: true`.
6. Report back: the public key and key id, the smoke-test results, and how long each packet took.

EOBs are protected health information. This is a preparation aid: it is not legal advice, it never predicts an outcome, a
person checks the screen, the clocks and the brief before filing, and the IDR entity decides.

Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds claim
files. If I ask for it later, follow the Provide page instead of improvising.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsNo Surprises Act IDR packet and eligibility screen on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

Standard · the hosted demo, one 96 GB card: what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Reads the claim facts from the EOB, remittanc...: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for No Surprises Act IDR packet and eligibility screen, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up No Surprises Act IDR packet and eligibility screen on my hardware

Fetch https://decosa.ai/prompts/nsa-idr-packet-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=nsa-idr-packet)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Reads the claim facts from the EOB, remittanc...: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/nsa-idr-packet-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 27 Sep 2026 · measured 27 Sep 2026: · p50 52 s · ~$0.020 per run · 23 receipts

Loading the nightly status…

Self-host: verified 27 Sep 2026 · fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

Measured cost to run: about $0.020 per claim (hosted, 27 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): window 2026-10-06 to 2026-10-09, the late case failed on initiation with no model, er-anesthesia-pa likely eligible with a brief at 198.48% of the QPA in 20 s (23 attested receipts), late-initiation-az and plan-consent-defense likely ineligible with no brief, record verified; the rehearsal bundle passed 15/15. Model-server startup was not re-run.

Known limits (5)
  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.
  • Measured on 91 synthetic, templated disputes written by the building agent; not on real remittances or against CMS's recorded eligibility outcomes.
  • State law comes from CMS's chart (state information current as of 11 Jan 2023); a state-regulated plan in one of the 21 listed states is marked needs review unless the file says whether the state law applies.
  • The pre-1 Nov 2026 similar-condition batching test is left to the certified IDR entity; the CPT section ranges for the 2026 test are our approximation (the guidance was not found).
  • The 2026 final rule's new portal steps are not modelled yet; the clocks will need an update when the Departments announce them.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the clocks, the screen from structured facts and the signed packet need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the no surprises act idr packet and eligibility screen API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.
The open stack

The EOB and notices in; a Federal IDR eligibility screen with the business-day arithmetic, and an offer brief built only from checked facts.

For out-of-network emergency, anesthesia, radiology and pathology groups, their billing companies, and the health plans on the other side. Give it the remittance or EOB, the open negotiation notice and the supporting documents. It reads the claim facts with quotes found word for word, then screens the claim for the Federal IDR process: coverage type, qualified service, state law, notice and consent, the 30-business-day open negotiation start, the 4-business-day initiation window, cooling-off periods and batching, each check with its dated rule and the arithmetic. When nothing fails, it gathers the evidence for each factor the arbiter must consider, states the offer as a dollar amount and a percentage of the QPA, and drafts a brief whose sentences are each checked against the documents; lines about billed charges, usual and customary charges or Medicare rates are left out by rule. Sealed in a signed, hash-chained packet. A preparation aid: a person files, and a certified IDR entity decides.

Deployment
Self-host first
Regulatory
EOBs and clinical records are protected health information under HIPAA: run it on your own hardware; the hosted demo takes synthetic disputes only. Rules read on 27 Sep 2026 from primary sources. 45 CFR 149.510 (eCFR, current to 24 Sep 2026): open negotiation within 30 business days of receiving the initial payment or denial ((b)(1)); initiation in the 4 business days starting on business day 31 of negotiation ((b)(2)(i)); no IDR after a valid notice-and-consent waiver ((b)(2)(ii)), which never covers ancillary services (149.420(b)); a 90-calendar-day cooling-off period after a determination ((c)(4)(vi), text in force before 3 Aug 2026). The Federal IDR Operations final rules (91 FR 33900, 4 Jun 2026, effective 3 Aug 2026; correction 91 FR 55462, 28 Aug 2026) apply in stages: the $15 administrative fee to disputes initiated on or after 11 Jun 2026; the new batching rules (at most 50 line items, three similar-condition tests, a 30-business-day batched cooling-off) to negotiations starting on or after 1 Nov 2026 (CMS notice, 3 Aug 2026); the new open negotiation, initiation and registry steps only 90 days after the Departments announce each IDR Gateway function (expected from Spring 2027), so until then the earlier text applies (149.510(h)). State law: CMS's applicability chart (state information current as of 11 Jan 2023) and each state's enforcement letter. Courts: TMA II (5th Cir., 2 Aug 2024) affirmed striking the rule on how arbiters weigh the QPA; TMA III (5th Cir. en banc, 11 Aug 2026) held that ghost rates may not be counted in the QPA and incentive payments may not be left out; CMS said guidance would follow (13 Aug 2026). TMA I and TMA IV are described from CMS guidance (opinions not read). It does not recompute QPAs, read state statutes or file anything. Not legal advice, and it never predicts an outcome.
Architecture
Text description

The remittance or EOB, the open negotiation notice, any notice of IDR initiation and the supporting documents go into decosa-api. Qwen3.8-27B reads the claim facts; code keeps a fact only when its quote is on a document line and its date or amount is written there. A rule engine screens eligibility (coverage, qualified service, state law, notice and consent, the 30- and 4-business-day windows, cooling-off, batching) with the arithmetic and the dated rule for each check. If nothing fails, the model gathers the evidence for each factor the arbiter must consider, lines about prohibited factors are left out by rule, and the model drafts the brief; the grounding judge, the numeric block and the prohibited-factor rule check each sentence. Outputs: the screen, the clocks, the offer as a dollar amount and a percentage of the QPA, the brief and a signed hash-chained packet. Hosted calls get gateway-signed receipts; a self-hosted box signs with its own key.

Architecture

At a glance

What it gives you
An eligibility screen (likely eligible, likely ineligible or needs review) with each check, its dated rule and the business-day arithmetic; the claim facts with quotes; the factor evidence; the offer as a dollar amount and a percentage of the QPA; a checked draft brief when nothing fails; a Markdown packet and a signed record.
What it does not do
It does not file in the IDR portal, predict or promise an outcome, recompute or audit a QPA, read state statutes (it uses CMS's chart and says when to check the state's letter), or decide the similar-condition batching test, which is the certified IDR entity's call. Scanned EOBs need OCR first.
Data retention
Nothing kept on the server. Documents and packets live in memory for the request; logs carry counts only. You keep the signed packet with the claim.
What leaves the box (hosted demo)
The documents go to Qwen3.8-27B through the Decosa API, whose receipts hold hashes, not text. Self-hosted, nothing leaves. The hosted demo is for synthetic disputes only.
Model calls per run
Screen mode: one call. Full packet: one fact read, one factor-evidence call, one brief and one grounding call per brief sentence, about 15 to 23 calls on the samples. The clocks and a screen from structured facts need no model.
Typical run cost
A fraction of a cent at the gateway list price for a screen from documents; a few cents for the emergency anesthesia packet with its brief. Each run shows its own measured cost.
Rules and dates
45 CFR 149.510 read on 27 Sep 2026 with the staged applicability of the 2026 final rule: $15 fee from 11 Jun 2026, new batching rules for negotiations from 1 Nov 2026, new portal steps only after each is announced. Plans and providers use the same screen.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 48 GB card

    The same pipeline on a smaller mixture-of-experts model. Faster and cheaper; accuracy on this task unknown.

    Models
    • Gemma 4 26B A4B (instruction-tuned)
    Hardware
    1x L40S or RTX 6000 Ada 48 GB (not measured)
    Quality evidence
    • eligibility and brief groundingnot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
  • In the hosted demo

    Standard

    the hosted demo, one 96 GB card

    Qwen3.8-27B reads the facts, gathers the factor evidence and drafts the brief; the screen, the clocks, the offer arithmetic and the prohibited-factor rule are plain code. Every model call receipted.

    Models
    • Qwen3.8-27B (NVIDIA NVFP4)
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • eligibility verdict from documents alone, held-out test (61 synthetic disputes)61/61; 32/32 planted ineligible caught on the planted check; 29/29 eligible, including traps, passeddecosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27, gateway route; split fixed before any run; templated synthetic documents
    • eligibility, rules only with full facts (91 planted cases)91/91; gold clocks from a separate implementationdecosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27
    • facts read right from the documents (dev / test)243/243 / 496/497 (the miss: an issue date read as the receipt date)decosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27
    • planted unsupported brief sentences removed / supported sentences kept25/25 / 22/22 (changed figures, invented credentials and statistics, contradictions, moved dates, billed charges, Medicare and UCR rates)decosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27; small n
    • business-day clocks against a second implementation (numpy over the OPM holiday list)160,000/160,000 comparisons agreedecosa-api docs/evals/nsa-idr-packet.md, 2026-09-27
    • real remittances, scanned EOBs and IDR outcomesnot measured yet
    Latency
    measured on our server under a shared gateway: a screen from documents in seconds; a full packet in under a minute to a couple of minutes (a dozen or two model calls)
    Verification
    Proof: strongHosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.
  • Best

    DeepSeek-V4-Flash on two more cards

    A larger model for long files and dense correspondence.

    Models
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    Hardware
    2x RTX PRO 6000 96 GB
    Quality evidence
    • eligibility and brief groundingnot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyNot a hosted model: calls are attested by the box's key only.
  • Needs more compute

    Wanted: the best setup

    two large judges from different families

    DeepSeek-V4-Flash and GLM-5.3-Flash each read the file, and a fact or brief sentence stands when they agree; disagreements go to the reviewer. Claim files stay on your own hardware, never on community providers. Not served yet.

    Models
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    • GLM-5.3-Flash
    Hardware
    Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash plus 2x 96 GB for GLM-5.3-Flash, or one Mac Studio with 512 GB holding both 4-bit builds (156 + 165 GB, sizes from our Mac; not run together yet). Estimate.
    Quality evidence
    • eligibility and brief groundingnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
After the session
Reads the claim facts from the EOB, remittance and notices (each with a quote), gathers the evidence for each factor the arbiter must consider, drafts the offer brief, and judges every brief sentence (the grounding judge)Qwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
Lite tier: the same pipeline on a 48 GB cardGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
25.2B (3.8B active)No proof yetSelf-host only
Best tier: a larger model for long files and dense correspondenceDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active) · 192 GBNo proof yetSelf-host only
Other
Second judge, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only

Measured

Does it catch the disputes that should not be filed?

91 synthetic disputes from a generator with a structured truth: 40 eligible (24 of them traps that look ineligible) and 51 planted ineligible, 3 of each of 17 kinds. The gold clocks come from a separate implementation. In the documents run the model reads every fact from EOB and notice text with distractor dates; a 61-case test split was fixed before any run and run once.

Held-out test, verdict right from documents
61 of 6132 ineligible caught on the planted check, 29 eligible passed
Facts read right from documents
496 of 497test; the miss read an issue date as the receipt date
Planted unsupported brief sentences removed
25 of 25and 22 of 22 supported sentences kept
Cost
about $0.001 per screen, $0.02 per packet1,662 tokens (1 call) and 52,263 tokens (23 calls) on the samples, gateway list price

What it does not show

The documents are templated synthetic text with one fact per line; real 835 remittances, scanned EOBs and letters are much harder to read. The same author wrote the cases and the rules, so a perfect rules score shows the code matches that reading, not that the reading is right. Check it on your own closed disputes before relying on it.

Where the rules are still moving

The 2026 final rule's new open negotiation and initiation steps apply only after the Departments announce each IDR Gateway function; the batching changes apply to negotiations starting on or after 1 Nov 2026. Every check shows which text it applies and when it was read.

Source: decosa-api docs/evals/nsa-idr-packet.md, 27 Sep 2026

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    Intake, the codes, the deadline rules, the quote checks, the recommendation rule, grounding, signing and the HTTP API (/appeal/*). No GPU. Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server.

  • 1x L40S / RTX 6000 Ada 48 GB

    Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite).

  • CPU only Fits

    The clocks (POST /idr/clock), the eligibility screen from structured facts (POST /idr/screen), the signed packet and verification need no GPU; reading documents and drafting need the model.

Latency per lane

  • screen from documents (one model call), busy shared gateway6.2 s

    Measureddecosa-api docs/evals/nsa-idr-packet.md, test median on our server 2026-09-27 (max 12.7 s), gateway route

  • full packet with a brief, busy shared gateway52.1 s

    Measuredmeasured on our server 2026-09-27: median of 7 runs of the three eligible samples (30 to 98 s; about 15 to 23 model calls), gateway route

  • full packet with a brief, self-hosted direct route20.0 s

    Measuredmeasured on our server 2026-09-27, self-host sandbox, er-anesthesia-pa (23 model calls)

  • clocks and screen from structured facts50 ms

    Estimateestimate: no model call

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

nsa-idr-packet/assemble-prompt.md194 lines
# Assemble the Decosa No Surprises Act IDR packet on this machine

You are setting up a self-hosted No Surprises Act IDR eligibility screen and offer packet on this Linux machine, for an
out-of-network provider group, a billing company or a health plan. It reads a dispute's documents (the remittance or EOB,
the open negotiation notice, any notice of IDR initiation, and supporting records) and returns:
- the claim facts, each with a quote found word for word in the documents;
- an eligibility screen for the Federal IDR process (coverage, qualified service, state law, notice and consent, the
  30-business-day open negotiation start, the 4-business-day initiation window, cooling-off, batching), each check with
  its dated rule and the business-day arithmetic;
- the evidence for each factor the arbiter must consider, and the offer as a dollar amount and a percentage of the QPA;
- a draft brief, only when no check fails, with every sentence checked against the documents;
- a signed, hash-chained packet.

Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- EOBs and clinical records are protected health information. Keep everything on this machine: the model route stays
  local (`direct`), and nothing goes to a hosted service.
- This is a preparation aid. It is not legal advice, it never predicts an outcome, a person checks the screen, the clocks
  and the brief and files, and a certified IDR entity decides eligibility and payment. The rules were read on 27 Sep 2026;
  check `GET /idr/info` for what applies and re-check after the Departments announce new IDR Gateway steps.

Repeat both points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/nsa-idr-packet.zip (4 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py nsa-idr-packet` (the api image carries the same bundle under /app/rehearsal/nsa-idr-packet/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py nsa-idr-packet --bundle nsa-idr-packet.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the IDR window opens on business day 31 of open negotiation, 6 Oct 2026 (no model)", "and closes on business day 34, 9 Oct 2026", "a self-insured plan that opted into New Jersey's law fails the state-law check (no model)"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
   - Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main`. On a 48 GB card, also
     set `LLM_MAX_LEN=32768`. Not measured.
   - Under 48 GB: stop and tell me it will not fit. (The clocks and the screen from structured facts need no GPU; see step 4.)
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the
   NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories, run
   `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 60 GB of free disk.

## 2. Get the images

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`.

If a pull fails, build from source once the `decosa-api` source is published: clone it, run
`docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .` in the clone, and
`docker compose build llm` from its compose file. If neither works, stop and tell me.

## 3. Write the compose file

Create `~/decosa-idr/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these packets, e.g. Example Emergency Physicians revenue cycle>"
```

Create `~/decosa-idr/docker-compose.yml` with exactly these services:

```yaml
name: decosa-idr
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"                # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "200000"        # per session; one packet needs up to about 9,500 generated tokens
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_IDR_MAX_CONCURRENT: "3"
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

Use the named volume `decosa-data` exactly as written. A host bind mount owned by root makes the API fail on
`/data/keys.sqlite`.

Run `docker compose up -d`, then poll `docker compose ps` until both services are healthy. The LLM takes 5-10 minutes
the first time. `curl -s localhost:8445/healthz` should show `"llm": true` (`asr` is false here, which is fine).

## 4. Smoke test

The clocks and the screen from structured facts need no model:

```bash
API=localhost:8445
curl -s $API/idr/clock -H 'content-type: application/json' \
  -d '{"claim":{"payment_received":"2026-08-13","on_sent":"2026-08-24"},"today":"2026-09-27"}' | jq '.initiation | {opens, closes, arithmetic}'
curl -s $API/idr/screen -H 'content-type: application/json' \
  -d '{"claim":{"payment_received":"2026-06-22","on_sent":"2026-07-01","idr_initiated":"2026-08-20"}}' | jq '{verdict, failed, open}'
```

Pass if the window is `2026-10-06` to `2026-10-09` (Labor Day skipped), and the screen says `likely_ineligible` with
`failed: ["initiation"]` (the window closed on 2026-08-18); the facts it was not given are listed in `open`.

Then the packets (synthetic disputes):

```bash
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"nsa-idr-packet"}' | jq -r .token)
for s in er-anesthesia-pa late-initiation-az plan-consent-defense; do
  curl -s $API/idr/packet -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
    -d "{\"sample_id\":\"$s\"}" > /tmp/idr-$s.json
  jq -c '{verdict: .screen.verdict, failed: .screen.failed, decision, pct: .figures.pct, brief: (.brief != null), receipts: (.receipts|length), statuses: ([.receipts[].status]|unique)}' /tmp/idr-$s.json
done
jq -r '.left_out[] | "\(.ref)\t\(.why)"' /tmp/idr-er-anesthesia-pa.json
jq '{record}' /tmp/idr-er-anesthesia-pa.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```

Pass if:
- `er-anesthesia-pa` is `likely_eligible` with `decision: "draft"`, `pct` 198.48, a brief, and the billed-charges (D1.L11),
  Medicare (D5.L3) and FAIR Health (D5.L4) lines in `left_out`;
- `late-initiation-az` is `likely_ineligible` with `failed: ["initiation"]` and no brief;
- `plan-consent-defense` is `likely_ineligible` with `failed: ["consent"]` and no brief;
- every receipt has `"status": "attested"`, and the record verifies (`ok: true`).

## 5. Point the app at the local API

- Base URL: `http://localhost:8445`. For a web app, set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`. Add other origins to
  `DECOSA_CORS_ORIGINS`.
- `POST /idr/packet` takes `{documents: [{kind, title?, date?, text}], claim?, batch?, prior_determinations?, offer: {amount,
  qpa?}, side?: provider|plan, facility_claim?, mode?: packet|screen}`. `mode: "screen"` makes one model call and stops
  after the screen: use it on every remittance before anyone opens a dispute. JSON, or Server-Sent Events with
  `Accept: text/event-stream`.
- Facts you already hold (from your claim system) go in `claim` and win over what the model reads; disagreements are listed.
  `POST /idr/screen` and `POST /idr/clock` take structured facts with no model at all.
- `GET /idr/info` lists every rule with its citation, link, the date read and whether it applies yet.
- The server stores nothing. Keep each signed packet (JSON) with the claim. Anyone can re-check it with
  `POST /record/verify` against the key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1. For other users on the LAN, put a TLS reverse proxy with authentication in front and set
  `DECOSA_TRUSTED_PROXIES`.

## 6. Keep the direct route

`DECOSA_LLM_ROUTE=gateway` would send prompts to the hosted Decosa
API, and those prompts contain the EOB and the claim file.
Never use it for real claims. At most, use it for synthetic training material.

Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

20 laws, rules and guidance pages cited; 16 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A No Surprises Act IDR eligibility screen that reads the EOB and notices, shows each check with its dated rule and the business-day arithmetic, and drafts the offer brief only from facts it can quote.
Who it's for
Out-of-network emergency, anesthesia, radiology and pathology groups, their billing companies and IDR vendors, and health plans answering disputes.
Where it runs
Self-host for real claims (hosted demo: synthetic disputes only)
Key numbers

On a held-out set of 61 synthetic disputes, read from the documents alone, the screen got 61 of 61 verdicts right; the documents are templated, so real files will be harder.

  • 61 of 61 Eligibility verdict right from documents, held out (test split, n = 61)
  • 32 of 32 Planted ineligible disputes caught on the planted check (test split, n = 32)
  • 29 of 29 Eligible disputes passed, including traps (test split, n = 29)
  • 52.1 s Median end-to-end run, hosted (QA sweep 2026-09-27)
All results, datasets and caveats
Models
Qwen3.8-27B
Where
Self-host for real claims (hosted demo: synthetic disputes only)
Checks
Receipt per model call; every fact's quote found word for word with its date or amount on the line; every brief sentence grounded; signed hash-chained packet
Output
Structured data · Signed record or verdict
Data
Patient data (PHI)
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

How does the No Surprises Act IDR eligibility screen decide?

It runs fixed checks on the claim facts: coverage type, qualified service, state law, notice and consent, the 30-business-day open negotiation start, the 4-business-day initiation window, cooling-off periods and batching. Each check shows its rule, the date the rule was read and the arithmetic. The IDR entity the parties select makes the actual eligibility decision.

Which IDR rules apply after the 2026 final rule?

The final rule took effect on 3 Aug 2026 but applies in stages: the $15 administrative fee from 11 Jun 2026, the new batching rules for negotiations starting on or after 1 Nov 2026, and the new open negotiation and initiation steps only 90 days after the Departments announce each IDR Gateway function. Until then the earlier text applies, and the screen uses it.

Does it count business days and federal holidays?

Yes. Every clock skips weekends and US federal holidays, and shows the arithmetic, such as business day 31 to 34 of open negotiation for initiation. Checked against a second implementation on 160,000 comparisons, all of which agreed.

Can the brief use billed charges or Medicare rates?

No. The IDR entity may not consider usual and customary charges, billed charges or public-payer rates such as Medicare's, so lines about them are left out of the evidence and brief sentences that rely on them are removed.

Can health plans use it?

Yes. The same screen and checks work for a plan answering a dispute, for example to find a signed notice-and-consent waiver or a missed window, and the brief can argue the plan's offer.

Will it tell me I will win the IDR dispute?

No. It never predicts an outcome. It prepares the file; a person checks it and files, and the IDR entity decides.

Where does claim data go?

Self-hosted, nothing leaves your hardware. The hosted demo takes synthetic disputes only; there, the text goes to Qwen3.8-27B through our gateway, whose receipts hold hashes, not text. Nothing is kept on the server.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about No Surprises Act IDR packet and eligibility screen

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.