Check a prior auth before you send it
Each criterion of the payer's own policy met, missing or not met with the chart quote, a plain send-or-not call, form values to copy and a peer-to-peer brief.
Built on: Grounding, Typed judgment, Signed record, Form filling
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the policy reader, the decision clock and the signed record need no GPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the prior-auth pre-check and packet API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real patient data belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- prior-auth-check
Use the hosted API
# Decosa prior-auth pre-check: use the hosted API
You are wiring Decosa's prior-auth pre-check into this project. Before a prior-authorization request goes to the payer,
it takes the payer's own published medical policy (text or the PDF), the chart and the requested service, and returns a
check with a signed record. Every model call has a signed receipt. Use only what is listed below; if you need something
else, stop and ask me.
The check has:
- the criteria section it used (code-table rows removed; licensed criteria such as InterQual or MCG are never included);
- each criterion `met`, `not_met` or `undocumented`, with the policy quote and a chart quote found word for word, and the
record to get when it's undocumented;
- a decision: `ready`, `gather_first`, `not_supported`, `proprietary_criteria` (the policy points to licensed criteria it
can't check) or `review`;
- form answers with chart quotes, a peer-to-peer brief, and a letter of medical necessity only when the decision is `ready`.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **Hosted use is for synthetic charts only.** Charts are protected health information: real patients belong on a
self-hosted box (see the self-host prompt) or confidential access. Say so wherever this is wired in.
- This is a drafting aid. Never auto-submit, never type the values into a payer portal for the user (payer portals'
terms generally forbid automation), never present the letter as final, and never tell a user the payer will approve.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in code.
Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "prior-auth-check"}` returns `{"token", "expires_at", "budget"}`.
A demo token runs one check at a time (409 otherwise); over a limit you get HTTP 429 with `Retry-After`.
3. A check needs about 12,000 generated tokens left in the budget before it starts (402 otherwise).
## Endpoints
- `POST /priorauth/check` (token). Body: `{"policy": {"title"?, "text" | "pdf_base64", "source"?, "section"?}, "records": [{"kind"?: "note"|"study"|"lab"|"imaging"|"order"|"letter"|"therapy"|"other", "title"?, "date"?, "text"}], "service", "codes"?: {"procedure": [...], "diagnosis": [...]}, "payer"?, "plan_type"?: "erisa_group"|"aca_individual"|"medicare_advantage"|"medicaid_mc"|"other", "urgent"?, "request_date"?: "YYYY-MM-DD", "form_questions"?: [string], "patient_ref"?}` or `{"sample_id": "..."}`.
- Limits: 8 records of 12,000 characters each, 40,000 in total; a policy of 400,000 characters or a 12 MB PDF with a text
layer; 20 form questions; codes only (a description is refused). URLs are not fetched.
- The JSON response has `recommendation` `{decision, headline, why, requirements, missing: [{text, get, any_of}], not_met, confirm, guidance}`,
`criteria` `[{id, text, quote, ref, group, kind?, status, evidence: [{ref, quote}], get?, note?}]`, `tree`, `policy`,
`clock`, `form` `{standard, criteria, questions}`, `brief`, `letter` (or null), `letter_sentences`, `packet_md`, `record`,
`record_check`, `usage` (calls, tokens, cost at list price), `receipts` and `note`.
- With `Accept: text/event-stream` (or `"stream": true`), the events are `ready`, `policy`, `clock`, `criteria`, a
`receipt` per model call, a `criterion` per criterion, `recommendation`, `form`, `brief`, a `sentence` per letter
sentence, `letter`, then `result`, `budget` and `done`.
- `POST /priorauth/section` (token): the policy reader alone (which section of a long policy or PDF it would use).
- `POST /record/verify` (no token) `{"record": {...}}` returns `{ok, summary, checks, first_bad}`.
- `GET /priorauth/info`, `GET /priorauth/samples` and `GET /attest/signing-key` need no token.
## Example: check a request and act on the decision (Python, `pip install httpx`)
```python
import httpx, json, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {"policy": {"title": "Payer policy", "text": open("policy-criteria.txt").read()},
"records": json.load(open("chart.json")), "service": "MRI of the lumbar spine",
"codes": {"procedure": ["72148"], "diagnosis": ["M54.16"]}, "plan_type": "erisa_group"}
r = httpx.post(f"{API}/priorauth/check", json=body, headers=H, timeout=900)
r.raise_for_status()
js = r.json()
print(js["recommendation"]["headline"])
for m in js["recommendation"]["missing"]:
print("get first:", m["text"], "->", m["get"])
if js["letter"]:
open("letter-of-medical-necessity.txt", "w").write(js["letter"]) # a draft: the clinician edits and signs it
open("prior-auth-check.md", "w").write(js["packet_md"])
json.dump(js["record"], open("prior-auth-check.json", "w")) # keep with the request
```
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa prior-auth pre-check: run it yourself (containers)
You are setting up the Decosa prior-auth pre-check on this machine, so charts never leave it. Before a request goes to the
payer, it reads the payer's published medical policy (text or PDF) and the chart, and returns:
- each criterion met, not met or not documented, with the policy and chart quotes, and the record to get;
- a decision (ready to send, get records first, not supported on this chart, or can't check licensed criteria);
- form answers with chart quotes, a peer-to-peer brief, and a letter of medical necessity only when every criterion is met.
It seals a signed record. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/prior-auth-check.zip (6 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py prior-auth-check` (the api image carries the same bundle under /app/rehearsal/prior-auth-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py prior-auth-check --bundle prior-auth-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the supported request is ready to send", "a letter of medical necessity is drafted and cites the AHI", "the unsupported request is not supported on this chart"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
- set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
- keep its data on a named volume;
- bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check. The first start
downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/priorauth/info` lists the decision-clock rules with their citations, links
and the date they were read. `GET /attest/signing-key` shows this box's public key.
5. Smoke test: get a token with `POST /demo/session {"vertical":"prior-auth-check"}`, then:
- `POST /priorauth/check {"sample_id": "cpap-ready"}`: expect `recommendation.decision` `ready` and a letter;
- `{"sample_id": "cpap-not-supported"}`: expect `not_supported` and `letter: null`;
- `{"sample_id": "proprietary-criteria"}`: expect `proprietary_criteria`;
- `POST /record/verify {"record": <record>}` should give `ok: true`, and every receipt should be `attested`.
6. Report back: the public key and key id, the smoke-test results, and how long each check took.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVIDIA NVFP4) needs a GPU.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns
The standard tier fits (57.6 of 192 GB). The best tier fits too.
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"prior-auth-check"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py prior-auth-check
Download the mock-data bundle (6 KB, 8 checks)expected.json
Two synthetic CPAP requests checked against the CPAP section of Aetna Clinical Policy Bulletin 0004 (Obstructive Sleep Apnea in Adults), before sending. In the first, an attended in-lab sleep study shows an AHI of 26.5 over more than 6 hours of sleep: the check must say ready to send, quote the study, and draft a letter of medical necessity. In the second, the home sleep test index is 3.6 events per hour, below the policy minimum, while a physician letter asserts the criteria are met: the check must say not supported on this chart and draft no letter. The signed record must verify, and fail once changed.
What the rehearsal checks
- the supported request is ready to send
- a letter of medical necessity is drafted and cites the AHI
- the unsupported request is not supported on this chart
- and no letter is drafted for it
- the not-met criterion quotes the 3.6 index
- the signed record verifies
- a record whose decision was changed no longer verifies
- every model call has a signed receipt
Licence: Synthetic charts (CC0): patients, notes and clinicians are invented. Policy: an excerpt of Aetna CPB 0004 (the payer's published medical policy, https://www.aetna.com/cpb/medical/data/1_99/0004.html), quoted for demonstration with its source. Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa prior-auth pre-check: run it yourself (containers)
You are setting up the Decosa prior-auth pre-check on this machine, so charts never leave it. Before a request goes to the
payer, it reads the payer's published medical policy (text or PDF) and the chart, and returns:
- each criterion met, not met or not documented, with the policy and chart quotes, and the record to get;
- a decision (ready to send, get records first, not supported on this chart, or can't check licensed criteria);
- form answers with chart quotes, a peer-to-peer brief, and a letter of medical necessity only when every criterion is met.
It seals a signed record. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/prior-auth-check.zip (6 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py prior-auth-check` (the api image carries the same bundle under /app/rehearsal/prior-auth-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py prior-auth-check --bundle prior-auth-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the supported request is ready to send", "a letter of medical necessity is drafted and cites the AHI", "the unsupported request is not supported on this chart"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
- set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
- keep its data on a named volume;
- bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check. The first start
downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/priorauth/info` lists the decision-clock rules with their citations, links
and the date they were read. `GET /attest/signing-key` shows this box's public key.
5. Smoke test: get a token with `POST /demo/session {"vertical":"prior-auth-check"}`, then:
- `POST /priorauth/check {"sample_id": "cpap-ready"}`: expect `recommendation.decision` `ready` and a letter;
- `{"sample_id": "cpap-not-supported"}`: expect `not_supported` and `letter: null`;
- `{"sample_id": "proprietary-criteria"}`: expect `proprietary_criteria`;
- `POST /record/verify {"record": <record>}` should give `ok: true`, and every receipt should be `attested`.
6. Report back: the public key and key id, the smoke-test results, and how long each check took.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsPrior-auth pre-check and packet on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Standard · the hosted demo, one 96 GB card: what changesuses estimates
- Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Picks the criteria section of a long policy,...: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Prior-auth pre-check and packet, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Prior-auth pre-check and packet on my hardware Fetch https://decosa.ai/prompts/prior-auth-check-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=prior-auth-check) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Picks the criteria section of a long policy,...: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/prior-auth-check-assemble.md
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: partial on 28 Sep 2026 · QA sweep · p50 41 s · p95 106 s (16 runs) · ~$0.014 per run
Loading the nightly status…
Self-host: verified 28 Sep 2026 · fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, run with a named data volume, direct route to the local Qwen3.8-27B, local signing; torn down after
Measured cost to run: about $0.014 per prior-auth request (hosted, 28 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The rehearsal bundle passed 8/8 (CPAP ready with a letter citing the AHI 26.5, CPAP not supported with no letter, the not-met criterion quoting 3.6, the record verifies and fails once changed, every receipt attested) in 28.2 s; the lumbar MRI sample came back ready in 17.5 s and the InterQual sample can't-check in 2.9 s, all receipts attested. Model-server startup was not re-run.
Known limits (4)
- Time and cost from the latest held-out run (4 checks at a time on the shared gateway); replaced by production measurements after launch.
- Accuracy is below the target set for this tool (95% of requirements right): a person checks every criterion before sending.
- Synthetic charts written against real policy excerpts; not measured on real charts or with coordinators' labels.
- Licensed criteria (InterQual, MCG) are out of scope; a policy that points to them gets can't check.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the policy reader, the decision clock and the signed record need no GPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the prior-auth pre-check and packet API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real patient data belongs on your own hardware.
The payer's own policy and the chart in, before you send: each criterion met, missing or not met with quotes, and a plain call on whether to send.
For prior-authorization coordinators, medical assistants and the physicians who take the peer-to-peer call. Give it the payer's published medical policy (pasted, or the PDF) and the chart. It finds the criteria section for the requested service, removes code tables, and splits the criteria into requirements, alternatives and exclusions. Each is checked against the chart: met with a chart quote the code finds word for word, missing with the record to get, or not met with the line that contradicts it. Every not-met answer is re-checked, a met answer resting only on a letter of medical necessity doesn't count, and a criterion with a threshold needs the value in its quote. The call is a rule: every requirement met means ready to send; one the chart contradicts means not supported on this chart; one it doesn't show means get that record first. It answers the payer's form questions with chart quotes, to copy into the payer's form or portal yourself, writes a one-page peer-to-peer brief from the verified quotes, and drafts a letter of medical necessity only when every criterion is met, each sentence checked against the chart and the policy. Licensed criteria (InterQual, MCG) are never reproduced: a policy that points to them gets a plain "can't check". It never suggests codes or changes to a note and never touches a payer portal. A person checks it, the clinician signs, the practice sends it.
- Deployment
- Self-host first
- Regulatory
- Charts are protected health information under HIPAA: run it on the practice's own hardware (confidential access can't take patient data until a business associate agreement is in place); the hosted demo takes synthetic charts only. Decision clocks read on 28 Sep 2026 (Cornell LII, eCFR text): employer group plans must decide a pre-service claim within 15 days of receiving it, extendable once by 15 days (29 CFR 2560.503-1(f)(2)(iii)(A)), and an urgent one within 72 hours ((f)(2)(i)); ACA individual coverage follows the same rules (45 CFR 147.136(b)(3)(i)); Medicare Advantage: 7 calendar days for a standard request for an item or service subject to its prior-authorization rules from 1 Jan 2026 (42 CFR 422.568(b)(1)(ii)) and 72 hours expedited (422.572(a)); Medicaid managed care: the state's time frame, no more than 7 calendar days for rating periods from 1 Jan 2026 (42 CFR 438.210(d)(1)(i)(B)) and 72 hours expedited ((d)(2)(i)). These are federal maximums; state law, plan documents and provider contracts can be shorter. Payer medical policies are quoted from the payer's published text with its source; licensed criteria sets (InterQual, MCG) are not included and never reproduced. CPT is an AMA code set: the tool takes codes as you enter them, removes descriptor rows from policy text and never suggests codes. Payer portals' terms of use generally forbid automation, so it never logs in to, fills or submits a portal. Not medical, coding or legal advice, and it never predicts the payer's decision.
Text description
The payer's medical policy (text or PDF), the chart, the requested service and codes, and the plan type go into decosa-api. Code reads the PDF, removes code tables and, for a long policy, Qwen3.8-27B picks the criteria section. The model splits it into requirements, alternatives and exclusions and checks each against the chart; code keeps a met or not-met answer only when its chart quote is found word for word, re-checks every not-met answer, and never judges licensed criteria. A rule turns the checklist into the call: ready to send, get records first, not supported, or can't check. Then the form answers, a peer-to-peer brief built in code, and, only when ready, a letter of medical necessity whose sentences the grounding judge checks. Outputs are sealed in a signed hash-chained record. Hosted calls get gateway-signed receipts; a self-hosted box signs with its own key.
At a glance
- What it gives you
- A call (ready to send, get records first, not supported on this chart, can't check), each criterion met, missing or not met with the policy and chart quotes, what record to get, form values with their chart lines to copy, a peer-to-peer brief, a letter of medical necessity when the chart supports it, the federal decision clock, a Markdown packet and a signed record.
- What it does not do
- It doesn't submit, fax or touch a payer portal, predict the payer's decision, know which services need authorization, suggest codes, or fetch policies (paste them or upload the PDF). It never includes InterQual or MCG criteria. It never suggests adding to or changing a note. Scanned charts and scanned policy PDFs need OCR first.
- Data retention
- Nothing kept on the server. Policies, charts and results live in memory for the request; logs carry counts only. You keep the signed record with the request.
- What leaves the box (hosted demo)
- The policy section and the chart go to Qwen3.8-27B through Decosa's gateway, whose receipts hold hashes, not text. Self-hosted, nothing leaves. The hosted demo is for synthetic charts only.
- Model calls per check
- One criteria map (plus one section pick for a long policy), one check per criterion, a re-check per not-met answer and per exclusion answered met, one form call; with a letter, one draft and one grounding call per sentence: 17 to 20 calls at the median, up to about 50.
- Typical run cost
- About a cent or two per check at the gateway list price (held-out runs); more with long policies and a letter. Each run shows its own measured cost.
- Measured accuracy
- On the latest fresh held-out set (16 synthetic charts, 4 real policies): 14 decisions right; 1 of 11 requests that should have waited called ready. Earlier sets: 24 of 35 and 45 of 60. Measured 28 Sep 2026; see the eval.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
one 48 GB card
The same pipeline on a smaller mixture-of-experts model. Faster and cheaper; accuracy on this task unknown.
- Models
- Gemma 4 26B A4B (instruction-tuned)
- Hardware
- 1x L40S or RTX 6000 Ada 48 GB (not measured)
- Quality evidence
- decision and criteria accuracynot measured yet
- Latency
- not measured yet
- Verification
- Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
- In the hosted demo
Standard
the hosted demo, one 96 GB card
Qwen3.8-27B finds the section, maps and checks the criteria, answers the form and drafts the letter; the quote checks, the recommendation, the clock and the brief are plain code. Every model call receipted.
- Models
- Qwen3.8-27B (NVIDIA NVFP4)
- Hardware
- 1x RTX PRO 6000 Blackwell 96 GB
- Quality evidence
- latest fresh held-out set (16 synthetic charts, 4 real policies, run once after the cold-user fixes): decision right14/16; unsupported requests called ready 1/11; supported called ready 4/5decosa-api docs/evals/prior-auth-check.md, measured on our server 2026-09-28, gateway route
- earlier fresh set (35 charts, 7 policies, after the guards) / first set (60 charts, 12 policies, before them)24/35 (unsupported called ready 4/21) / 45/60 (10/31)decosa-api docs/evals/prior-auth-check.md
- requirement status right (of requirements the model found)91.1%, 91.8%, 93.6%; 81% to 90% of gold requirements founddecosa-api docs/evals/prior-auth-check.md
- undocumented read as not met (the appeal engine's main error, 6 of 80 cases)0 of 111 casesdecosa-api docs/evals/prior-auth-check.md
- letter sentences rated unsupported by a blind reviewer6/375, 2/91, 0/46decosa-api docs/evals/prior-auth-check.md
- Latency
- measured on our server under a shared gateway: seconds per packet, under a minute at the slow end; a packet with a letter makes a couple of dozen model calls
- Verification
- Proof: strongHosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.
Best
DeepSeek-V4-Flash on two more cards
A larger model for long charts and dense payer policies.
- Models
- DeepSeek-V4-Flash (NVIDIA NVFP4)
- Hardware
- 2x RTX PRO 6000 96 GB
- Quality evidence
- decision and criteria accuracynot measured yet
- Latency
- not measured yet
- Verification
- Proof: partialSelf-host onlyNot a hosted model: calls are attested by the box's key only.
- Needs more compute
Wanted: the best setup
two large judges from different families
DeepSeek-V4-Flash and GLM-5.3-Flash each read the file, and a finding stands when they agree; disagreements go to the reviewer. Patient records stay on your own hardware, never on community providers. Not served yet.
- Models
- DeepSeek-V4-Flash (NVIDIA NVFP4)
- GLM-5.3-Flash
- Hardware
- Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash plus 2x 96 GB for GLM-5.3-Flash, or one Mac Studio with 512 GB holding both 4-bit builds (156 + 165 GB, sizes from our Mac; not run together yet). Estimate.
- Quality evidence
- decision and criteria accuracynot measured yet
- Latency
- not measured yet
- Verification
- No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
| After the session | ||||
Picks the criteria section of a long policy, splits the policy into requirements, alternatives and exclusions, checks each against the chart, re-checks every not-met answer (and every exclusion answered met), answers the payer's form questions, drafts the letter of medical necessity when the chart supports every criterion, and judges every letter sentence (the grounding judge)Qwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 57 GBProof: strongIn the hosted demo | Standard | 27.8B · 57 GB | Proof: strongIn the hosted demo | |
| ||||
Lite tier: the same pipeline on a 48 GB cardGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab) 25.2B (3.8B active)No proof yetSelf-host only | Lite | 25.2B (3.8B active) | No proof yetSelf-host only | |
| ||||
Best tier: a larger model for long charts and dense payer policiesDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab) 284B (13B active) · 192 GBNo proof yetSelf-host only | BestWanted | 284B (13B active) · 192 GB | No proof yetSelf-host only | |
| ||||
| Other | ||||
Second judge, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab) 321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only | Wanted | 321B (18B active) · about 170 GB (estimate) | No proof yetSelf-host only | |
| ||||
Does it say "don't send" when it should?
Synthetic charts written blind against 23 real published payer policies and sections (Aetna, Cigna, UnitedHealthcare, CMS LCDs), each labelled twice (author and blind reviewer agreed on all 744 requirements). Prompts were written on 4 other policies. Three sets were each run once: before the guards (60 cases), after them (35), and after fixes from two blind coordinator tests (16).
- Decision right, latest set
- 14 of 16earlier sets: 24 of 35, 45 of 60
- Unsupported requests it called ready
- 1 of 11earlier sets: 4 of 21, 10 of 31
- Supported requests it called ready
- 4 of 5earlier sets: 8 of 14, 26 of 29
- Undocumented read as not met
- 0 of 111 casesthe appeal engine's main error was 6 of 80
- Cost per check
- about $0.014median on the latest set, gateway list price; p95 $0.027
Where it fails
It misses requirements hidden in footnotes, appendices and tables (13% to 19% aren't found), misreads a table lookup (a resection-weight scale by body surface area), and still calls a few unsupported requests ready. Treat ready as "nothing obviously missing", not a guarantee.
What it does not show
Synthetic charts, one author per policy, no labels from working coordinators. Real charts are longer and messier. Check it on your own recent requests before relying on it.
Source: decosa-api docs/evals/prior-auth-check.md, 28 Sep 2026
Tools, services and hardware
Tools
- decosa denial appeal packet (vertical 60) (opens in a new tab)AGPL-3.0-or-later (decosa-api)
The criteria engine: criteria mapping with verbatim policy quotes, the verbatim chart-quote rule, the date-window check and the letter check; subclassed, not copied.
- decosa grounding (vertical 17) (opens in a new tab)AGPL-3.0-or-later (decosa-api)
Judges every letter sentence against the chart, the policy and the denial, and cites the span; imported, not copied.
- decosa typed-judgment (vertical 24) (opens in a new tab)AGPL-3.0-or-later (decosa-api)
Answer parsing and the stated-confidence probability for each criterion check.
- decosa record (vertical 07) and POST /record/verify (opens in a new tab)AGPL-3.0-or-later (decosa-api)
The hash chain and the Ed25519-signed packet; anyone can re-check it.
- pypdfium2 (opens in a new tab)Apache-2.0 or BSD-3-Clause
The policy PDF's text layer, on CPU.
The decision clocks, quoted in /priorauth/info with the date they were read.
- Payer medical policies (Aetna CPB 0004, Cigna 0051, UnitedHealthcare panniculectomy) (opens in a new tab)the payers' published policies, quoted in excerpt with their source for the demo
The demo samples' policies; your own policies are pasted or uploaded.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:0.1.0The policy reader, the quote checks, the recommendation rule, the decision clock, the brief, grounding, signing and the HTTP API (/priorauth/*). No GPU. Binds 127.0.0.1 by default.
- decosa-llm:8000
${DECOSA_REGISTRY}/decosa-llm:0.1.0vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server.
- 1x L40S / RTX 6000 Ada 48 GB
Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite).
- CPU only Fits
The policy reader without a long-policy section pick, the decision clock, the signed record and verification need no GPU; checking the chart needs the model.
Latency per lane
- one check, shared gateway, 4 at a time (16 fresh held-out cases)41.3 s
Measureddecosa-api docs/evals/prior-auth-check.md: p50 41.3 s, p95 105.5 s, 2026-09-28, gateway route
- one check, shared gateway, 4 at a time (35 fresh held-out cases)57.5 s
Measureddecosa-api docs/evals/prior-auth-check.md: p50 57.5 s, p95 109.7 s
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa prior-auth pre-check on this machine
You are setting up a self-hosted prior-authorization pre-check on this Linux machine for a practice's prior-auth
coordinators and medical assistants. Before a request goes to the payer, it reads the payer's own published medical
policy (pasted text or the policy PDF) and the chart, and returns:
- the criteria section for the requested service (code tables removed; licensed criteria sets such as InterQual or MCG
are never included, and a policy that points to them is reported as "can't check");
- each criterion marked met (with a chart quote found word for word), missing (with the record to get), or not met (with
the chart line that contradicts it; every not-met answer is re-checked);
- a recommendation: ready to send, get these records first, or not supported on this chart;
- answers for the payer's form questions with chart quotes, a one-page peer-to-peer brief, and, only when every criterion
is met, a draft letter of medical necessity with every sentence checked against the chart and the policy;
- the federal decision clock for the plan type, and a signed, hash-chained record of what was checked.
Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.
**Before anything else, remind me:**
- Charts are protected health information. Keep everything on this machine: the model route stays local (`direct`), and
nothing goes to a hosted service.
- This is a drafting aid. It is not medical, coding or legal advice and does not predict the payer's decision. A person
checks every answer, the ordering clinician signs the letter, and the practice sends the request. It never logs in to,
fills or submits a payer portal, never suggests codes, and never suggests adding to or changing a note.
Repeat both points in your final summary.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/prior-auth-check.zip (6 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py prior-auth-check` (the api image carries the same bundle under /app/rehearsal/prior-auth-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py prior-auth-check --bundle prior-auth-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the supported request is ready to send", "a letter of medical necessity is drafted and cites the AHI", "the unsupported request is not supported on this chart"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## What you are building
| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
## 1. Check the GPU, driver and Docker
1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
- Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
- Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main`. On a 48 GB card, also
set `LLM_MAX_LEN=32768`. Not measured.
- Under 48 GB: stop and tell me it will not fit.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the
NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories. Then run
`sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 60 GB of free disk.
## 2. Get the images
The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`.
If a pull fails, build from source once the `decosa-api` source is published:
- clone it;
- in the clone, run `docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`;
- run `docker compose build llm` from its compose file.
If neither works, stop and tell me.
## 3. Write the compose file
Create `~/decosa-priorauth/.env`:
```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Clinic prior-auth team>"
```
Create `~/decosa-priorauth/docker-compose.yml` with exactly these services:
```yaml
name: decosa-priorauth
x-health: &health
interval: 15s
timeout: 5s
retries: 5
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
"--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
"--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
"--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
api:
image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
restart: unless-stopped
depends_on: { llm: { condition: service_healthy } }
environment:
DECOSA_LLM_ROUTE: direct # local model only; receipts are signed by this box's key ("attested")
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LOCAL_SIGNING: "on" # Ed25519 key created at /data/attest/ed25519.pem on first start
DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_LLM_TOKENS: "200000" # per session; one check needs up to about 12,000 generated tokens
DECOSA_SESSION_TTL_S: "28800"
DECOSA_PRIORAUTH_MAX_CONCURRENT: "3"
DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
ports: ["127.0.0.1:8445:8445"]
volumes: [decosa-data:/data]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```
Use the named volume `decosa-data` exactly as written. A host bind mount owned by root makes the API fail on
`/data/keys.sqlite`.
Run `docker compose up -d`, then poll `docker compose ps` until both services are healthy. The LLM takes 5-10 minutes
the first time. `curl -s localhost:8445/healthz` should show `"llm": true` (`asr` is false here, which is fine).
## 4. Smoke test
```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"prior-auth-check"}' | jq -r .token)
for s in cpap-ready cpap-not-supported proprietary-criteria; do
curl -s $API/priorauth/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
-d "{\"sample_id\":\"$s\"}" > /tmp/pa-$s.json
jq -c '{decision: .recommendation.decision, letter: (.letter != null), criteria: (.criteria|length), receipts: (.receipts), cost: .usage.cost_usd}' /tmp/pa-$s.json
done
jq -r '.criteria[] | "\(.status)\t\(.text)"' /tmp/pa-cpap-not-supported.json
jq '{record}' /tmp/pa-cpap-ready.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```
Pass if:
- `cpap-ready` is `ready` with a letter of medical necessity (the sample chart has an in-lab sleep study with an AHI of
26.5);
- `cpap-not-supported` is `not_supported` with no letter, and a not-met criterion quotes the home sleep test's index of 3.6;
- `proprietary-criteria` is `proprietary_criteria` (the UnitedHealthcare panniculectomy policy points to InterQual) with no
criteria checked and no letter;
- the record verifies (`ok: true`).
The samples are synthetic charts against excerpts of real published payer policies (Aetna CPB 0004, Cigna 0051,
UnitedHealthcare panniculectomy), quoted with their sources.
## 5. Point the app at the local API
- Base URL: `http://localhost:8445`. For a web app, set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`. Add other origins to
`DECOSA_CORS_ORIGINS`.
- `POST /priorauth/check` takes `{policy: {title?, text | pdf_base64, source?}, records: [{kind?, title?, date?, text}],
service, codes?: {procedure: [...], diagnosis: [...]}, payer?, plan_type?, urgent?, request_date?, form_questions?,
patient_ref?}`. It returns JSON, or streams Server-Sent Events when asked with `Accept: text/event-stream`.
- `POST /priorauth/section` runs the policy reader alone (which section of a long policy or PDF it would use).
- `GET /priorauth/info` lists the decision-clock rules with their citations, links and the date they were read, and the
limits.
- Codes are yours: enter them as numbers; descriptions are not accepted and code-table rows are removed from policies.
- The server stores nothing. Keep each signed record (JSON) with the request. Anyone can re-check it with
`POST /record/verify` against the key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1. For other users on the LAN, put a TLS reverse proxy with authentication in front and set
`DECOSA_TRUSTED_PROXIES`.
## 6. Keep the direct route
`DECOSA_LLM_ROUTE=gateway` would send prompts to the hosted Decosa
API, and those prompts contain the chart. Never use it for
real patients. At most, use it for synthetic training material.
Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- The payer's own policy and the chart in, before you send: each criterion met, missing or not met with quotes, and a plain call on whether to send.
- Who it's for
- Prior-authorization coordinators, medical assistants and physicians at specialty and primary-care practices.
- Where it runs
- Self-host for real patients (hosted demo: synthetic charts only; confidential access can't take patient data yet)
- Key numbers
- 14 of 16 Decision right, latest fresh set (test split, n = 16)
- 1 of 11 Requests that should have waited, called ready (test split, n = 11)
- 24 of 35 Decision right, earlier fresh set (test split, n = 35)
- Models
- Qwen3.8-27B
- Where
- Self-host for real patients (hosted demo: synthetic charts only; confidential access can't take patient data yet)
- Checks
- Receipt per model call; every criterion quote found word for word in the policy, every met or not-met answer quoted from the chart, not-met answers re-checked; every letter sentence grounded; signed hash-chained record
- Industry
- Healthcare
- Runs
- Self-host
- Output
- Notes, reports and drafts · Signed record or verdict
- Data
- Patient data (PHI)
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Self-host
- Built from
- Grounding · Typed judgment · Signed record · Form filling
Questions people ask
Does it always say the request is ready?
No. It says ready only when the chart documents every requirement it found in the policy, after a final check for anything it missed. If the chart contradicts one it says not supported on this chart and drafts no letter; if the chart leaves one out it names the record to get first. On the latest held-out set it still called 1 of 11 requests that should have waited ready, so a person checks every line.
Which payer policies does it use?
The policy you paste or upload: the payer's own published medical policy. It fetches nothing. Licensed criteria such as InterQual or MCG are never included; when a policy points to them, it says it can't check them.
Does it fill in the payer's portal?
No. Payer portals generally forbid automation. It lists each form value with the chart line it came from, and you copy them in yourself.
Will it suggest codes or change the note?
No. Codes are what you enter, and it never suggests adding to or changing a note: a missing criterion names the existing record to get.
Where does patient data go?
Self-hosted, nothing leaves your hardware. The hosted demo takes synthetic charts only; there, the text goes to Qwen3.8-27B through Decosa's gateway, whose receipts hold hashes, not text. Nothing is kept on the server.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Prior-auth pre-check and packet
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…