Skip to content
decosa
LiveHostedSelf-hostSelf-host first for real data

Answer a payer audit

The respond-by date, each claim checked against every documentation requirement of the payer's policy with the chart, page and line, the weak claims first, a draft cover letter and an indexed response packet.

Held-out test9 / 9Weak claims flagged on test audits it had never seen (held-out test)
On production60 smedian on production (2026-09-28); slower when the service is busy
List price~$0.13 per 100 claimsmeasured, at list price

Built on: Grounding, Signed record

1. Pick a sample

Sample

Real patient data: request confidential access, or run the tool on your own hardware. The demo takes samples only.

2. Run it

On production the sample took 60 s (median of 5 runs, 2026-09-28; slowest 81 s). Slower when the service is busy.

Result

The answer appears here first, then what it found, the draft, and how long it took. Sample: Solo therapist: 140 charts requested, due in 10 days (synthetic).

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the deadline, the matching, the packet and the signed record need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the payer audit response API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
payer-audit

Use the hosted API

# Decosa payer audit response: use the hosted API

You are wiring Decosa's "Answer a payer audit" tool into this project. It takes a payer's (or its review contractor's)
medical-records request or post-payment audit letter, the charts, and the payer's documentation policy, and returns:
- the respond-by date, worked out in code from the letter's own words, with the arithmetic and the letter line;
- for every claim in the letter's list: its note (matched by member and date of service), and each requirement of the
  payer's policy marked `found` (with the chart, page and line: `C3 p2:L14`, and the words) or `missing`;
- the claims ordered weak first (`weak`: an element missing, no note, or an entry after the request; `check`: something
  unclear; `supported`: every element found);
- a draft cover letter whose sentences were each checked against the letter and the index, a bookmarked PDF packet
  (letter, index, records), a worksheet, a CSV, the values an upload form asks for with their sources, and a signed record.
Every model call has a signed receipt. Use only what is listed below; if you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`. Health check: `GET https://api.decosa.ai/healthz`.
- **Hosted use is for synthetic cases only.** Charts are protected health information: real records belong on a
  self-hosted box (see the self-host prompt) or confidential access. Say so wherever this is wired in.
- It is a review and drafting aid, not legal, coding or compliance advice. Never auto-send the packet, never present the
  letter as final, and never add to, change or back-date a record because of what it finds.
- Never automate a payer or auditor portal (their terms forbid it). Use the `portal_fields` values to fill an upload form
  by hand.

## Auth
1. Preferred: an API key (`dk_…`) in `DECOSA_API_KEY`, sent as `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "payer-audit"}` gives a short token. A demo session covers
   the 12-claim sample; the 140-claim sample needs a key (402 otherwise). One run per token at a time (409).

## Endpoints
- `POST /payer-audit/run` (token). Body: `{"mode"?: "respond"|"self_audit", "letter": {"text", "date"?, "received_date"?},
  "claims"?: "<table as text>" | [{"claim_id"?, "patient", "dos", "code", "units"?, "provider"?}], "charts": [{"id"?, "title"?, "text"}],
  "pack": "<built-in pack id>" | {<your own pack: GET /payer-audit/packs shows the shape>} | "policy": {"title", "text", "source"?},
  "practice"?: {"name", "npi", "contact", "address"}, "sample_size"?, "seed"?, "today"?}` or `{"sample_id": "therapist-12"}`.
  - Charts are text; split pages with form feeds or lines like `--- Page 2 ---`. When `claims` is left out, the claim list
    is read from the letter (a table with a date of service and a billed code on each row).
  - Prefer a pack (the policy's requirements, quoted, reviewed once by a person). Pasting the policy text makes the model
    read the requirements itself; that is measured as less reliable (see the eval), so review what it read.
  - Response (JSON; SSE with `Accept: text/event-stream` or `"stream": true`): `summary {verdict, counts, weak_first}`,
    `deadline {respond_by, days_left, rows[{kind, date, ref, quote, arithmetic}], notes}`, `claims[{n, claim_id, patient, dos,
    code, status, reasons, elements[{id, label, kind, status, cite, quote, note}], note{chart, pages}}]`, `policy`, `index`,
    `cover_letter`, `portal_fields[{label, value, source}]`, `record`, `record_check`, `usage`, `run_id`, `export`.
- `GET /payer-audit/runs/{run_id}/export?format=packet|worksheet|csv|letter|record` (same token): the PDF packet, the
  Markdown worksheet, the CSV table, the letter, the signed record. Kept in memory for one hour.
- `POST /payer-audit/deadline` (no token, no model) `{"letter": {"text", "date"?, "received_date"?}, "today"?}`.
- `GET /payer-audit/packs`, `/payer-audit/samples`, `/payer-audit/info`; `POST /record/verify {"record"}`.

## Example (Python, `pip install httpx`)
```python
import httpx, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
r = httpx.post(f"{API}/payer-audit/run", json={"sample_id": "therapist-12"}, headers=H, timeout=900)
r.raise_for_status()
out = r.json()
print(out["summary"]["verdict"], "respond by", out["deadline"]["respond_by"])
for c in out["claims"]:
    if c["status"] != "supported":
        print(c["n"], c["dos"], c["status"], "; ".join(c["reasons"]))
pdf = httpx.get(API + out["export"]["packet"], headers=H, timeout=120).content
open("response-packet.pdf", "wb").write(pdf)
```
Show the weak claims first and as plainly as the supported ones. A person, with their compliance lead or counsel, decides
what to send and signs the letter.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa payer audit response: run it yourself (containers)

You are setting up Decosa's "Answer a payer audit" tool on this machine, so charts never leave it. It reads an auditor's
records request, the charts and the payer's documentation policy, and returns the respond-by date (in code), each claim's
policy requirements found with chart, page and line or missing (weak claims first), a draft cover letter checked sentence
by sentence, a bookmarked response packet and a signed record. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/payer-audit.zip (7 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py payer-audit` (the api image carries the same bundle under /app/rehearsal/payer-audit/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py payer-audit --bundle payer-audit.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the respond-by date is the letter date + 10 calendar days (no model)", "the run gives the same date", "12 claims are read from the letter"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin (docs.docker.com/engine/install),
   and the NVIDIA container toolkit; check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the `llm`
   service (Qwen3.8-27B on vLLM) and the `api` service. On `api`: `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`,
   `DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_BUDGET_LLM_TOKENS=400000` (a 140-claim audit needs about 70,000 generated tokens),
   data on a named volume, every port bound to 127.0.0.1.
3. `docker compose pull && docker compose up -d`; wait for the `llm` health check (the first start downloads about 20 GB).
4. Check `curl -fsS http://127.0.0.1:<PORT>/payer-audit/info` (lists the built-in policy packs with their sources) and
   `GET /attest/signing-key` (this box's public key; show it to me).
5. Smoke test:
   - `POST /payer-audit/deadline` with the sample letter (`GET /payer-audit/samples?id=therapist-12`, field `letter`) and
     `"today": "2026-09-24"` needs no model and should give `respond_by` `2026-10-01`.
   - With a token from `POST /demo/session {"vertical":"payer-audit"}`: `POST /payer-audit/run {"sample_id": "therapist-12"}`.
     Expect 4 weak claims (numbers 4, 5, 6 and 10), listed first; every `found` element with a cite; every receipt `attested`.
   - `GET <export.packet>` with the same token returns a PDF; `POST /record/verify {"record": <record>}` gives `ok: true`.
6. Report back: the public key, the smoke results and how long the run took.

Charts are protected health information. This is a review and drafting aid, not legal, coding or compliance advice: a
person reads the weak claims with their compliance lead or counsel, decides what to send and signs. It never suggests
changing a record, and it never touches a payer or auditor portal.

Off by default. Never join as a provider on a box that holds patient charts.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVIDIA NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"payer-audit"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py payer-audit

Download the mock-data bundle (7 KB, 11 checks)expected.json

A synthetic records request from a made-up review contractor, dated 21 Sep 2026, with ten calendar days to respond, for 12 psychotherapy claims of a made-up solo practice, and the practice's progress notes. Checked against BCBSM's published individual-therapy documentation requirements (the built-in pack). Four claims are planted weak: two notes with no start and stop times, one with an empty interventions field, and one whose times appear only in an addendum dated after the request. Two sessions on the same day overlap, which an auditor would ask about. The run must mark exactly the four weak, mark the overlapping pair to check by hand and list them first, cite every found requirement by chart, page and line, give the respond-by date 2026-10-01, and seal a record that verifies.

What the rehearsal checks
  • the respond-by date is the letter date + 10 calendar days (no model)
  • the run gives the same date
  • 12 claims are read from the letter
  • exactly the four planted claims are weak
  • four weak claims
  • and they are listed first
  • the two sessions that overlap (same clinician, same day) are marked check, not supported
  • claim 6's times are only in an addendum dated after the request
  • a draft cover letter names the request reference
  • the signed record verifies
  • every model call has a signed receipt

Licence: Synthetic (CC0): practice, clinician, clients, IDs, plan and auditor are invented (scripts/payer_audit_cases.py). Policy: BCBSM's published documentation requirements, quoted short with the source. Part of decosa-api.

Prompt for your coding agent

# Decosa payer audit response: run it yourself (containers)

You are setting up Decosa's "Answer a payer audit" tool on this machine, so charts never leave it. It reads an auditor's
records request, the charts and the payer's documentation policy, and returns the respond-by date (in code), each claim's
policy requirements found with chart, page and line or missing (weak claims first), a draft cover letter checked sentence
by sentence, a bookmarked response packet and a signed record. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/payer-audit.zip (7 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py payer-audit` (the api image carries the same bundle under /app/rehearsal/payer-audit/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py payer-audit --bundle payer-audit.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the respond-by date is the letter date + 10 calendar days (no model)", "the run gives the same date", "12 claims are read from the letter"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin (docs.docker.com/engine/install),
   and the NVIDIA container toolkit; check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the `llm`
   service (Qwen3.8-27B on vLLM) and the `api` service. On `api`: `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`,
   `DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_BUDGET_LLM_TOKENS=400000` (a 140-claim audit needs about 70,000 generated tokens),
   data on a named volume, every port bound to 127.0.0.1.
3. `docker compose pull && docker compose up -d`; wait for the `llm` health check (the first start downloads about 20 GB).
4. Check `curl -fsS http://127.0.0.1:<PORT>/payer-audit/info` (lists the built-in policy packs with their sources) and
   `GET /attest/signing-key` (this box's public key; show it to me).
5. Smoke test:
   - `POST /payer-audit/deadline` with the sample letter (`GET /payer-audit/samples?id=therapist-12`, field `letter`) and
     `"today": "2026-09-24"` needs no model and should give `respond_by` `2026-10-01`.
   - With a token from `POST /demo/session {"vertical":"payer-audit"}`: `POST /payer-audit/run {"sample_id": "therapist-12"}`.
     Expect 4 weak claims (numbers 4, 5, 6 and 10), listed first; every `found` element with a cite; every receipt `attested`.
   - `GET <export.packet>` with the same token returns a PDF; `POST /record/verify {"record": <record>}` gives `ok: true`.
6. Report back: the public key, the smoke results and how long the run took.

Charts are protected health information. This is a review and drafting aid, not legal, coding or compliance advice: a
person reads the weak claims with their compliance lead or counsel, decides what to send and signs. It never suggests
changing a record, and it never touches a payer or auditor portal.

Off by default. Never join as a provider on a box that holds patient charts.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsPayer audit response on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

Standard · the hosted demo, one 96 GB card: what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Reads the auditor's letter: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Payer audit response, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Payer audit response on my hardware

Fetch https://decosa.ai/prompts/payer-audit-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=payer-audit)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Reads the auditor's letter: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/payer-audit-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 28 Sep 2026 · measured 28 Sep 2026: · p50 60 s · p95 81 s (5 runs) · ~$0.018 per run · 18 receipts

Loading the nightly status…

Self-host: verified 28 Sep 2026 · fresh clone of the branch into a clean directory on our server, api image built from docker/api/Dockerfile, run with a named data volume, direct route to the already-running local Qwen3.8-27B, local signing; torn down after

Measured cost to run: about $0.13 per 100 claims (hosted, 28 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Rehearsal bundle 10/10 in 25.3 s (4 weak claims listed first, respond by 2026-10-01, record verifies, every receipt attested); smoke ok in 21.5 s with 18/18 attested receipts and a PDF packet. Model-server startup was not re-run.

Known limits (6)
  • Hosted timing: 5 runs of the 12-claim sample on the pre-release server over the shared gateway (47.8-81.2 s; p95 is the slowest of 5). Production is re-measured after the merge.
  • Synthetic cases only, written by other workloads from real published policies; no practice manager or auditor has rated the output.
  • One signature per note: documents signed by two people (a therapist and a certifying physician) are read as one.
  • Text charts only: scanned PDFs are not read yet.
  • A pasted policy is read by the model and is not reliable (see the eval); a reviewed pack is.
  • Checks across claims (overlapping sessions by the same clinician, near-identical notes) were added after the blind cold-user test; they mark claims to check by hand and are not validated on real charts (the 92% wording threshold was set on the demo data).

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the deadline, the matching, the packet and the signed record need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the payer audit response API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.
The open stack

An auditor's records request and the charts in; the respond-by date, each claim checked against the payer's policy with page:line cites, weak claims first, and a packet.

For owners and managers of small practices, and the billing services that work for them, when a payer or its review contractor asks for records or audits paid claims. Give it the letter, the charts and the payer's published documentation policy. It works out the respond-by date from the letter in code, reads the claim list, matches each claim to its note, and marks each of the policy's requirements found (with chart, page and line) or missing. Weak claims come first, as plainly as the supported ones. It drafts a cover letter whose sentences are each checked against the sources, builds a bookmarked response packet with an index, lists the values an upload form asks for, and signs a record of what was checked. It never suggests changing a record and never touches a payer portal. The same check runs as a self-audit.

Deployment
Self-host first
Regulatory
Charts are protected health information under HIPAA: run it on the practice's own hardware (confidential access can't take patient data until a business associate agreement is in place); the hosted demo takes synthetic cases only. Not legal, coding or compliance advice: a person, with their compliance lead or counsel, decides what to send and signs. It never suggests adding to, changing or back-dating documentation (a code guard enforces this) and it flags entries dated after the request instead of counting them. It holds no CPT descriptors or AMA guideline text: the practice supplies its codes, and the check uses the payer's own published policy text, quoted with the source. Payer and auditor portals are not automated (their terms forbid it). In a self-audit, rules on identified overpayments may apply (for Medicare, 42 CFR 401.305, text read on eCFR 28 Sep 2026); ask counsel.
Architecture
Text description

The auditor's letter, the charts and the payer's policy go into decosa-api. Code works out the respond-by date, reads the claim list and matches each claim to its note. Qwen3.8-27B points at the times, the signature and any addenda and judges each content requirement; code checks every quote and does every count and date sum. The grounding judge checks each cover-letter sentence. Out come the claim table with page:line cites, weak claims first, the draft cover letter, a bookmarked PDF packet, the upload-form values and a signed record.

Architecture

At a glance

What it gives you
The respond-by date with its arithmetic; each claim's note and each policy requirement found (chart, page, line and the words) or missing; weak claims first; a draft cover letter; a bookmarked PDF packet with an index; the values an upload form asks for with their sources; a worksheet, a CSV and a signed record.
Objective measures
BCBS Michigan asks for objective tools to monitor progress: a scored instrument's result (PHQ-9, GAD-7, PCL-5) counts, the client's own rating (SUDS, a 0-10 score) does not (our reading; BCBSM does not define the term). Optum and Evernorth ask for no instrument in a progress note, and the packs say so with their sources.
Checks across claims
Two sessions by the same clinician that overlap on the same day, and a note whose wording is nearly identical to another claim's note, are flagged to check by hand, with the other claim named.
What it does not do
It does not send, fax or upload anything, or log in to a portal. It does not judge medical necessity or whether a code was right, know your contract or the auditor's sampling, or read scanned charts (text and pages only). It never suggests changing a record.
Data retention
Nothing written to disk. The letter, charts and result live in memory; a finished run is kept one hour so its downloads work, then dropped. Logs carry counts only. The signed record holds hashes, not chart text.
What leaves the box (hosted demo)
The letter and notes go to Qwen3.8-27B through Decosa's gateway, whose receipts hold hashes, not text. Self-hosted, nothing leaves. The hosted demo takes synthetic cases only; real records go through self-host (confidential access can't take patient data until a business associate agreement is in place).
Model calls per audit
One letter read, one call per claim, one cover-letter draft and one grounding call per letter sentence: 18 calls for 12 claims, about 145 for 140.
Typical run cost
A few cents for the small sample and more for a full audit at the gateway list price; a fraction of a cent per claim on the held-out sets. Each run shows its own measured cost.
Codes
You supply the billed codes. The tool holds no code descriptors or code-set guidelines; it checks the chart against the payer's own published policy text, quoted with its source.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 48 GB card

    The same pipeline on a smaller mixture-of-experts model. Accuracy on this task not measured.

    Models
    • Gemma 4 26B A4B (instruction-tuned)
    Hardware
    1x L40S or RTX 6000 Ada 48 GB (not measured)
    Quality evidence
    • recommendation and criteria accuracynot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
  • In the hosted demo

    Standard

    the hosted demo, one 96 GB card

    Qwen3.8-27B reads the letter and each note; the respond-by date, the matching, times, signatures, dates, the claim status and the packet are code. Every model call receipted.

    Models
    • Qwen3.8-27B (NVIDIA NVFP4)
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • weak claims flagged, held-out set B (pack mode)24/26; 0/35 clean claims flagged weak; both misses were plans of care with two signers (the engine reads one signature per note)decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route
    • requirements marked as labelled, held-out set B423/433 (97.7%); the found words were on the labelled line 355/377decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route
    • respond-by date right3/3 held-out letters (and 5/5 dev letters)decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route
    • kept cover-letter sentences a blind judge found unsupported0/26 (6 letters); 0 argued, advised or promiseddecosa-api docs/evals/payer-audit.md; blind judge: Claude Code (Opus 5.5) sub-agent that saw only the sources and the sentences
    • pasted policy text instead of a pack (the model reads the requirements)weak 21/26 but 22/35 clean claims flagged weak: not reliable; review the requirements it read, or use a packdecosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route
    Latency
    measured over the shared gateway: about a minute for the small sample; a few minutes for a full audit depending on load; seconds per claim
    Verification
    Proof: strongGateway route: every model call has a gateway-signed receipt.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
After the session
Reads the auditor's letter (who, dates, reference, policy named), reads a pasted policy's requirements when no pack is given, and for each claim points at the note's times, signature and addenda and judges each content requirement found or missing with the exact words; drafts the cover letter body and judges each of its sentences (the grounding judge)Qwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
Lite tier: the same pipeline on a 48 GB cardGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
25.2B (3.8B active)No proof yetSelf-host only

Measured

Does it flag the weak claims?

Blind-written synthetic audits against real published payer policies: another agent wrote each policy's requirement list, the letters, the charts with planted defects and the labels, without seeing the tool. The engine was frozen before held-out set B was opened, then run once.

Weak claims flagged (held-out B)
24 of 263 policies, 61 claims; both misses were plans of care with two signers
Clean claims flagged weak (held-out B)
0 of 35no false alarms in pack mode
Requirements found/missing as labelled
423 of 433held-out B, pack mode
Cost per audit of about 20 claims
about $0.03held-out B median $0.028, 80 s, gateway list price

Where it fails

A document with two signers (a therapist's plan of care certified by a physician) is read as one signature, so a late or uncredentialed certifying signature was missed twice. Pasting the policy text instead of using a reviewed pack makes the model read the requirements, and that flagged 22 of 35 clean claims: use a pack, or review what it read.

What it will not do

Suggest adding to, changing or back-dating a record; argue the case; judge medical necessity or coding; send anything or touch a portal.

What a practice manager said

A blind test user playing a practice manager put a 140-chart request at about 25 hours by hand (her estimate); the tool took about 4 minutes and $0.15. She would pay $150-300 per audit once she can put her own charts in, and would still call counsel when a lot of money is at stake or the letter mentions fraud.

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    Letter and claim-list reading, the respond-by date, note matching, the checks in code, the cover letter check, the PDF packet, signing and the HTTP API (/payer-audit/*). No GPU. Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server.

  • 1x L40S / RTX 6000 Ada 48 GB

    Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite).

  • CPU only Fits

    The respond-by date (POST /payer-audit/deadline), the claim list, the packet PDF and record verification need no GPU; reading the notes needs the model.

Latency per lane

  • 12-claim sample, busy shared gateway59.9 s

    Measured5 runs on our server 2026-09-28, pre-release branch, gateway route: 58.6, 81.2, 64.3, 59.9, 47.8 s

  • 12-claim sample, self-hosted direct route21.5 s

    Measuredself-host sandbox on our server 2026-09-28 (smoke run)

  • 140-claim sample, shared gateway157.9 s

    Measuredrecorded run on our server 2026-09-28 (final engine), 145 model calls, $0.1513; an earlier run under heavier load took 273.5 s

  • respond-by date only50 ms

    Estimateestimate: no model call

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

payer-audit/assemble-prompt.md188 lines
# Assemble the Decosa payer audit response on this machine

You are setting up a self-hosted "Answer a payer audit" tool on this Linux machine for a practice (or the billing service
working for it) that has received a payer's medical-records request or post-payment audit. It reads the auditor's letter,
the charts and the payer's documentation policy, and returns:
- the respond-by date, worked out in code from the letter's own words, with the arithmetic and the letter line;
- each claim matched to its note, and each requirement of the payer's policy marked found (with chart, page and line and
  the words) or missing, the weak claims listed first;
- a draft cover letter whose sentences are each checked against the letter and the index;
- a bookmarked PDF response packet (letter, index, the records as kept), a worksheet, a CSV and a signed record.

Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- Charts are protected health information. Keep everything on this machine: the model route stays local (`direct`), and
  nothing goes to a hosted service.
- This is a review and drafting aid, not legal, coding or compliance advice. It never suggests adding to, changing or
  back-dating a record, and it never touches a payer or auditor portal. A person reads the weak claims with their
  compliance lead or counsel, decides what to send and signs.

Repeat both points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/payer-audit.zip (7 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py payer-audit` (the api image carries the same bundle under /app/rehearsal/payer-audit/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py payer-audit --bundle payer-audit.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the respond-by date is the letter date + 10 calendar days (no model)", "the run gives the same date", "12 claims are read from the letter"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
   - Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main`. On a 48 GB card, also
     set `LLM_MAX_LEN=32768`. Not measured.
   - Under 48 GB: stop and tell me it will not fit.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the
   NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories. Then run
   `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 60 GB of free disk.

## 2. Get the images

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`.

If a pull fails, build from source once the `decosa-api` source is published:
- clone it;
- in the clone, run `docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`;
- run `docker compose build llm` from its compose file.

If neither works, stop and tell me.

## 3. Write the compose file

Create `~/decosa-payer-audit/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Counseling compliance>"
```

Create `~/decosa-payer-audit/docker-compose.yml` with exactly these services:

```yaml
name: decosa-payer-audit
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"                # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "400000"        # per session; a 140-claim audit used about 71,000 generated tokens
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_PAYERAUDIT_MAX_CONCURRENT: "2"
      DECOSA_PAYERAUDIT_WORKERS: "4"            # claims checked in parallel
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

Use the named volume `decosa-data` exactly as written. A host bind mount owned by root makes the API fail on
`/data/keys.sqlite`.

Run `docker compose up -d`, then poll `docker compose ps` until both services are healthy. The LLM takes 5-10 minutes
the first time. `curl -s localhost:8445/healthz` should show `"llm": true` (`asr` is false here, which is fine).

## 4. Smoke test

The deadline needs no model (the sample letter is synthetic):

```bash
API=localhost:8445
curl -s "$API/payer-audit/samples?id=therapist-12" | jq '{letter, today: "2026-09-24"}' \
  | curl -s $API/payer-audit/deadline -H 'content-type: application/json' -d @- | jq '{respond_by, days_left, rows: [.rows[].arithmetic]}'
```

Pass if `respond_by` is `2026-10-01` (the letter is dated 21 Sep 2026 and gives ten calendar days).

Then the 12-claim sample (a made-up practice; the policy is BCBSM's published individual-therapy documentation list):

```bash
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"payer-audit"}' | jq -r .token)
curl -s $API/payer-audit/run -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
  -d '{"sample_id":"therapist-12"}' > /tmp/pa.json
jq -c '{verdict: .summary.verdict, weak: [.claims[] | select(.status=="weak") | .n], first: .summary.weak_first[0:4], respond_by: .deadline.respond_by, receipts}' /tmp/pa.json
jq -r '.claims[] | select(.status=="weak") | "\(.n)\t\(.reasons | join("; "))"' /tmp/pa.json
jq '{record}' /tmp/pa.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
curl -s "$API$(jq -r .export.packet /tmp/pa.json)" -H "authorization: Bearer $TOKEN" -o /tmp/pa-packet.pdf && head -c 8 /tmp/pa-packet.pdf
```

Pass if:
- claims 4, 5, 6 and 10 are weak and listed first (no start/stop times; an empty interventions field; times only in an
  addendum dated after the request; no start/stop times), and the other 8 are supported;
- every found element has a cite like `C3 p2:L14`; the record verifies (`ok: true`); the packet starts with `%PDF`;
- the receipts in `/tmp/pa.json` are signed by this box (`GET /receipts/<id>` shows `"status": "attested"`).

## 5. Point the app at the local API

- Base URL: `http://localhost:8445`. For a web app, set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`.
- `POST /payer-audit/run` takes `{letter: {text, date?, received_date?}, charts: [{id?, title?, text}], pack | policy,
  claims?, practice?, mode?: respond|self_audit, sample_size?, seed?}`. Charts are text; split pages with form feeds or
  `--- Page 2 ---` lines. Leave `claims` out to read the claim list from the letter.
- Use a pack for your payer's policy (`GET /payer-audit/packs` shows the shape: each requirement quoted from the payer's
  published text, with the source). Pasting the policy text instead makes the model read the requirements; review them.
- Exports: `GET /payer-audit/runs/<run_id>/export?format=packet|worksheet|csv|letter|record`, kept in memory one hour.
- The server stores nothing else. Keep the signed record with the audit file; anyone can re-check it with
  `POST /record/verify` against the key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1. For other users on the LAN, put a TLS reverse proxy with authentication in front.

## 6. Keep the direct route

`DECOSA_LLM_ROUTE=gateway` would send prompts to the hosted Decosa
API, and those prompts contain the charts. Never use it for
real patients.

Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
An auditor's records request and the charts in; the respond-by date, each claim checked against the payer's policy with page:line cites, weak claims first, and a packet.
Who it's for
Owners and managers of small practices that take insurance, and the billing services that work for them.
Where it runs
Self-host for real records (hosted demo: synthetic cases only; confidential access can't take patient data yet)
Key numbers
  • 9 / 9 Weak claims flagged, new blind BCBSM letter (29 Sep) (held out, n = 9)
  • 7 / 9 Weak claims flagged, new blind Optum letter (29 Sep) (held out, n = 9)
  • 23 / 26 Weak claims flagged, held-out B rerun (29 Sep) (held out, n = 26)
  • 59.9 s Median end-to-end run, hosted (QA sweep 2026-09-28)
All results, datasets and caveats
Models
Qwen3.8-27B
Where
Self-host for real records (hosted demo: synthetic cases only; confidential access can't take patient data yet)
Checks
Receipt per model call; every found element's words located in the note by code; times, dates and signatures read in code; every cover-letter sentence grounded; signed hash-chained record
Output
Notes, reports and drafts · Signed record or verdict
Data
Patient data (PHI)
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

What does a payer audit response need?

The records for every claim the letter lists, sent by the date the letter sets. This tool works out that date from the letter's own words, matches each claim to its note, and marks each requirement of the payer's policy found (with chart, page and line) or missing, then builds an indexed packet and a draft cover letter for you to review and sign.

Will it hide the weak claims?

No. Weak claims are listed first, as plainly as the supported ones. On the held-out synthetic audits it flagged 24 of 26 weak claims and flagged none of 35 clean ones as weak. Talk to your compliance lead or a healthcare attorney about the weak ones before you respond.

Does it fix or complete my notes?

No. It never suggests adding to, changing or back-dating a record; a code guard enforces this. Entries dated after the request are flagged and not counted.

Does it upload to the payer portal?

No. Payer and auditor portal terms forbid automated use. It gives you files, and the values an upload form asks for with where each came from, to copy by hand.

Can I use it with real patient records?

Not on the hosted demo, which takes synthetic cases only. Run it on your own hardware; confidential access can't take patient data until a business associate agreement is in place.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Payer audit response

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.