Skip to content
decosa
PreviewHostedSelf-hostSelf-host first for real data

Read a demand before the deadline

The respond-by date the letter sets and the statute's minimum, every condition with its page, and the specials tied to the bills. No value, no recommendation.

On production2.7 minmedian on production (2026-09-29); slower when the service is busy
List price~$0.020 per demand packagemeasured, at list price

Built on: Grounding, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the deadline rules and the tie-out need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the injury demand reader API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
demand-reader

Use the hosted API

# The Decosa demand reader: use the hosted API

You are wiring the Decosa demand reader into this project. It returns:
- the respond-by date the demand letter sets, computed in code from its own words and the date you received it;
- the statute's minimum for the encoded states (CA, GA, MO, MT, UT; Florida's safe harbour), each with its source;
- every condition of acceptance, quoted with its page, and the state's required elements;
- each bill line tied to its page, summed, and reconciled with what the letter claims;
- providers billed with no records, records with no bill, gaps in care and earlier care;
- notes on the letter's citations (a statute from another state, a case that isn't at its cite).

Every model call has a signed receipt. Use only what is listed below; if you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **Hosted use is for made-up files only.** Real files belong on a self-hosted box (see the self-host prompt) or confidential access. Say so wherever this is wired in.
- It never values a claim or recommends paying, offering or accepting. Deadlines are reminders from the letter's own words; check them against the file and, where it matters, with counsel. States other than CA, GA, MO, MT, UT and FL return "not encoded: check with counsel".

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "demand-reader"}` returns `{"token", "expires_at", "budget"}`. Demo sessions are limited per IP per hour; over a limit you get HTTP 429 with `Retry-After`. A token runs one request at a time (409 otherwise).

## Endpoints
- `POST /demand/read` (token). Body: `{state, received: {date, method}, claim?: {claim_number?, loss_date?}, documents: [{kind: letter|bill|record, title?, provider?, pages: [text, ...]}]}`, or `{"sample_id": "..."}` for a sample from `GET /demand/samples`. Add `"stream": true` (or `Accept: text/event-stream`) for events; the last one is `result`, which carries `usage` (tokens and cost at list price).
- `GET /demand/info`: what is checked, sources, limits, not checked.
- `POST /record/verify` (no auth): re-check the signed record.

Show the result's own words to the user: the verdict first, then each item with its quote and where it is. Never add a value, a recommendation or a "safe" of your own.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# The Decosa demand reader: run it yourself (containers)

You are setting up the Decosa demand reader on this machine, so the files never leave it. It returns:
- the respond-by date the demand letter sets, computed in code from its own words and the date you received it;
- the statute's minimum for the encoded states (CA, GA, MO, MT, UT; Florida's safe harbour), each with its source;
- every condition of acceptance, quoted with its page, and the state's required elements;
- each bill line tied to its page, summed, and reconciled with what the letter claims;
- providers billed with no records, records with no bill, gaps in care and earlier care;
- notes on the letter's citations (a statute from another state, a case that isn't at its cite).

Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with "not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/demand-reader.zip (4 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py demand-reader` (the api image carries the same bundle under /app/rehearsal/demand-reader/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py demand-reader --bundle demand-reader.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "it is a time-limited demand", "respond by 3 Sep 2026, from the letter's 30 days after receipt", "the statute's minimum is the same day"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin (docs.docker.com/engine/install), and the NVIDIA container toolkit; check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service; on `api` set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`; keep its data on a named volume; bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/demand/info` answers; `GET /attest/signing-key` shows this box's public key.
5. Smoke test: get a token with `POST /demo/session {"vertical":"demand-reader"}`, then `POST /demand/read {"sample_id": "ga-policy-limits"}`. Expect:
   - `.deadline.respond_by` is `"2026-09-03"` and `.deadline.statute_minimum.date` is `"2026-09-03"` (30 days from receipt, O.C.G.A. 9-11-67.1);
   - an item says `$19,540.00` of the `$31,450.00` claimed is not supported by any bill;
   - a note says the letter gives 10 days where the statute gives at least 30;
   - no item or the memo values the claim;
   - every receipt has `"status": "attested"` and the record verifies.

Finish with a summary and these reminders:
- Demand packages hold medical records and personal information. Keep everything on this machine: the model route stays local (`direct`).
- It never values a claim or recommends paying, offering or accepting. Deadlines are reminders from the letter's own words; check them against the file and, where it matters, with counsel. States other than CA, GA, MO, MT, UT and FL return "not encoded: check with counsel".
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVIDIA NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns

    The standard tier fits (57.6 of 192 GB). The best tier fits too.

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"demand-reader"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py demand-reader

Download the mock-data bundle (4 KB, 8 checks)expected.json

A synthetic Georgia motor-vehicle demand (two letter pages), two bills and three sets of records, received by certified mail on 4 Aug 2026. The respond-by date must be 3 Sep 2026 (30 days from receipt, O.C.G.A. 9-11-67.1), the $19,540.00 of the $31,450.00 claimed that no bill supports must be found, the letter's '10 days' misstatement and its Florida statute must be noted, and the signed record must verify.

What the rehearsal checks
  • it is a time-limited demand
  • respond by 3 Sep 2026, from the letter's 30 days after receipt
  • the statute's minimum is the same day
  • the bills total $11,910.00
  • $19,540.00 claimed is not supported by any bill
  • the signed record verifies
  • a record with its respond-by date changed no longer verifies
  • every model call has a signed receipt

Licence: Synthetic: the claimant, insured, firm, insurer and providers are invented. Part of decosa-api.

Prompt for your coding agent

# The Decosa demand reader: run it yourself (containers)

You are setting up the Decosa demand reader on this machine, so the files never leave it. It returns:
- the respond-by date the demand letter sets, computed in code from its own words and the date you received it;
- the statute's minimum for the encoded states (CA, GA, MO, MT, UT; Florida's safe harbour), each with its source;
- every condition of acceptance, quoted with its page, and the state's required elements;
- each bill line tied to its page, summed, and reconciled with what the letter claims;
- providers billed with no records, records with no bill, gaps in care and earlier care;
- notes on the letter's citations (a statute from another state, a case that isn't at its cite).

Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with "not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/demand-reader.zip (4 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py demand-reader` (the api image carries the same bundle under /app/rehearsal/demand-reader/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py demand-reader --bundle demand-reader.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "it is a time-limited demand", "respond by 3 Sep 2026, from the letter's 30 days after receipt", "the statute's minimum is the same day"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin (docs.docker.com/engine/install), and the NVIDIA container toolkit; check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service; on `api` set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`; keep its data on a named volume; bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/demand/info` answers; `GET /attest/signing-key` shows this box's public key.
5. Smoke test: get a token with `POST /demo/session {"vertical":"demand-reader"}`, then `POST /demand/read {"sample_id": "ga-policy-limits"}`. Expect:
   - `.deadline.respond_by` is `"2026-09-03"` and `.deadline.statute_minimum.date` is `"2026-09-03"` (30 days from receipt, O.C.G.A. 9-11-67.1);
   - an item says `$19,540.00` of the `$31,450.00` claimed is not supported by any bill;
   - a note says the letter gives 10 days where the statute gives at least 30;
   - no item or the memo values the claim;
   - every receipt has `"status": "attested"` and the record verifies.

Finish with a summary and these reminders:
- Demand packages hold medical records and personal information. Keep everything on this machine: the model route stays local (`direct`).
- It never values a claim or recommends paying, offering or accepting. Deadlines are reminders from the letter's own words; check them against the file and, where it matters, with counsel. States other than CA, GA, MO, MT, UT and FL return "not encoded: check with counsel".

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsInjury demand reader on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

Standard · the hosted demo, one 96 GB card: what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Reads the letter's terms and every condition...: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Injury demand reader, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Injury demand reader on my hardware

Fetch https://decosa.ai/prompts/demand-reader-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=demand-reader)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Reads the letter's terms and every condition...: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/demand-reader-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 29 Sep 2026 · measured 29 Sep 2026: · p50 161 s · p95 208 s (8 runs) · ~$0.020 per run · 23 receipts

Loading the nightly status…

Self-host: not yet verified

Measured cost to run: about $0.020 per demand package (hosted, 29 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Known limits (4)
  • Hosted verification ran on our pre-release server through the production gateway, before these routes reached the production API.
  • Measured on 24 synthetic packages written by one author; real packages are longer, scanned and messier.
  • Typed or pasted page text only in this version; scans go through the document reader first.
  • Six states encoded; the rest return 'not encoded: check with counsel'.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the deadline rules and the tie-out need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the injury demand reader API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.
The open stack

An injury demand package in: the respond-by date the letter sets, every condition with its page, and the specials tied to the bills. Never a value.

For bodily-injury adjusters at carriers, TPAs and self-insureds. It reads the lawyer's letter, the bills and the records and shows whether it is a time-limited demand, the respond-by date the letter sets (computed in code from its words and the day you received it) next to the statute's minimum where the state is encoded, every condition of acceptance with its page, the state's required elements, each bill line tied to its page and reconciled with what the letter claims, providers billed without records, gaps in care, earlier care, and notes on the letter's citations. It never values the claim or recommends paying.

Deployment
Self-host first
Regulatory
Time-limited demand statutes read 28-29 Sep 2026 and encoded with their sources: California CCP 999-999.5 (SB 1155, demands transmitted from 1 Jan 2023: at least 30 days from transmission by email, fax or certified mail, 33 by mail); Georgia O.C.G.A. 9-11-67.1 as amended by SB 83 (2024: at least 30 days from receipt; payment at least 40 days); Missouri RSMo 537.058 (at least 90 days from receipt); Montana MCA 33-18-251 (2023: at least 60 days from receipt, rolled to the next business day); Utah Code 31A-22-323 (S.B. 74, 2026: at least 30 days); Florida 624.155(4) (90-day safe harbour after notice) and 627.4137 (30-day disclosure). Texas (common-law Stowers) and Louisiana were checked and have no such statute; every other state returns "not encoded: check with counsel". Deadlines are reminders from the letter's own words, not legal advice. The NAIC AI model bulletin asks insurers to govern third-party AI: the signed record lists every model call.
Architecture
Text description

The letter, the bills and the records go in as numbered pages with the date received. The open model reads the letter's terms and conditions, each bill page and each record page; the grounding judge checks what it writes. Code keeps only quotes and amounts found on their page, computes the respond-by date and the statute's minimum, sums the bills, matches bills to records and finds gaps. Out come the deadline, the conditions, the specials tie-out, notes on the letter and a signed record.

Architecture

At a glance

What it checks
Time-limited or not; the respond-by and payment dates from the letter's words; the statute's minimum (CA, GA, MO, MT, UT; FL's safe harbour); each condition with its page; the state's required elements; bill lines on their pages, summed and reconciled per provider; records vs bills; gaps; earlier care; citations from another state, cases not at their cite, and misstatements of the law.
What it never does
It never values the claim, recommends paying, offering or accepting, or judges the letter's credibility. A sentence that would do so is removed. The adjuster decides.
Data retention
Nothing kept on the server: the package lives in memory for the request; logs carry counts only. You keep the signed record with the claim file.
What leaves the box (hosted demo)
The pages go to Qwen3.8-27B through the Decosa API, whose receipts hold hashes, not text. Case citations (only the cite) are looked up in the local index and, on a miss, CourtListener. The hosted demo is for made-up packages; real files run on your own box or confidential access.
Model calls per package
About 20-25: the letter (three reads), one per bill page, one per record page, and the grounding checks.
Cost per package
A few cents or less per package on average at the gateway list price (held-out test). Each run shows its own measured cost.
States encoded
California, Georgia, Missouri, Montana, Utah; Florida's safe harbour and disclosure rules. Texas and Louisiana checked with no statute found. Any other state: the letter's own deadline, and 'not encoded: check with counsel'.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 48 GB card

    The same pipeline on a smaller mixture-of-experts model. Faster and cheaper; accuracy on this task unknown.

    Models
    • Gemma 4 26B A4B (instruction-tuned)
    Hardware
    1x L40S or RTX 6000 Ada 48 GB (not measured)
    Quality evidence
    • accuracy on this tasknot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
  • In the hosted demo

    Standard

    the hosted demo, one 96 GB card

    Qwen3.8-27B reads the letter, the bills and the records with quotes; the dates, the statute rules, the sums and the matching are code. About 20-25 receipted calls per package.

    Models
    • Qwen3.8-27B (NVIDIA NVFP4)
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • respond-by date exact, held-out test (8 packages, gateway, run once)7 of 7decosa-api docs/evals/demand-reader.md, measured on our server 2026-09-29, gateway route
    • conditions of acceptance found / billed totals within $1 (test)31 of 31 / 8 of 8decosa-api docs/evals/demand-reader.md, measured on our server 2026-09-29, gateway route
    Latency
    measured: a few minutes per package on the shared gateway under load (test set)
    Verification
    Proof: strongHosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.
  • Best

    DeepSeek-V4-Flash on two more cards

    A larger model for long, multi-claimant files.

    Models
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    Hardware
    2x RTX PRO 6000 96 GB
    Quality evidence
    • accuracy on this tasknot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyNot a hosted model: calls are attested by the box's key only.
  • Needs more compute

    Wanted: the best setup

    two large judges from different families

    DeepSeek-V4-Flash and GLM-5.3-Flash each read the file, and a finding stands when they agree; disagreements go to the reviewer. Claim files stay on your own hardware, never on community providers. Not served yet.

    Models
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    • GLM-5.3-Flash
    Hardware
    Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash plus 2x 96 GB for GLM-5.3-Flash, or one Mac Studio with 512 GB holding both 4-bit builds (156 + 165 GB, sizes from our Mac; not run together yet). Estimate.
    Quality evidence
    • accuracy on this tasknot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
After the session
Reads the letter's terms and every condition with quotes, each bill page's charge lines, and each record page's visits (the medical chronology's page extractor); the grounding judge checks every sentence it writesQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
Lite tier: the same reads on a 48 GB cardGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
25.2B (3.8B active)No proof yetSelf-host only
Best tier: a larger model for packages of several hundred pagesDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active) · 192 GBNo proof yetSelf-host only
Other
Second judge, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    Package intake, quote and amount checks, the deadline rules, the sums, the chronology merge, citation look-ups, signing and the HTTP API (/demand/*). No GPU. Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server.

  • 1x L40S / RTX 6000 Ada 48 GB

    Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite).

  • CPU only Fits

    The date rules, the sums, the header checks, signing and verification need no GPU; reading the text needs the model.

Latency per lane

  • one package (15-23 pages), busy shared gateway161.0 s

    Measureddecosa-api docs/evals/demand-reader.md, measured on our server 2026-09-29, gateway route, test set p50 (p95 208 s)

  • deadline rules and the sums50 ms

    Estimateestimate: no model call

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

demand-reader/assemble-prompt.md156 lines
# Assemble the Decosa demand reader on this machine

You are setting up the Decosa demand reader on this Linux machine for a bodily-injury claims unit (carrier, TPA or self-insured). It returns:
- the respond-by date the demand letter sets, computed in code from its own words and the date you received it;
- the statute's minimum for the encoded states (CA, GA, MO, MT, UT; Florida's safe harbour), each with its source;
- every condition of acceptance, quoted with its page, and the state's required elements;
- each bill line tied to its page, summed, and reconciled with what the letter claims;
- providers billed with no records, records with no bill, gaps in care and earlier care;
- notes on the letter's citations (a statute from another state, a case that isn't at its cite).

Everything is sealed in a signed, hash-chained record anyone can re-check.

Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- Demand packages hold medical records and personal information. Keep everything on this machine: the model route stays local (`direct`).
- It never values a claim or recommends paying, offering or accepting. Deadlines are reminders from the letter's own words; check them against the file and, where it matters, with counsel. States other than CA, GA, MO, MT, UT and FL return "not encoded: check with counsel".

Repeat both points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/demand-reader.zip (4 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py demand-reader` (the api image carries the same bundle under /app/rehearsal/demand-reader/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py demand-reader --bundle demand-reader.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "it is a time-limited demand", "respond by 3 Sep 2026, from the letter's 30 days after receipt", "the statute's minimum is the same day"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
   - Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main`. On a 48 GB card, also set `LLM_MAX_LEN=32768`. Not measured.
   - Under 48 GB: stop and tell me it will not fit.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories, then run `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 60 GB of free disk.

## 2. Get the images

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`.
If a pull fails, build from source once the `decosa-api` source is published: in the clone, `docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`, and `docker compose build llm` from its compose file. If neither works, stop and tell me.

## 3. Write the compose file

Create `~/decosa-demand/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Agency>"
```

Create `~/decosa-demand/docker-compose.yml` with exactly these services:

```yaml
name: decosa-demand
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"                # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "200000"
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
      DECOSA_DEMAND_MAX_CONCURRENT: "3"
      DECOSA_PREFLIGHT_CL_OFF: "1"          # case look-ups: local index or none; no CourtListener calls
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

Use the named volume `decosa-data` exactly as written; a host bind mount owned by root makes the API fail on `/data/keys.sqlite`.
Run `docker compose up -d`, then poll `docker compose ps` until both services are healthy (5-10 minutes the first time). `curl -s localhost:8445/healthz` should show `"llm": true`.

## 4. Smoke test

```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"demand-reader"}' | jq -r .token)
curl -s $API/demand/read -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
  -d '{"sample_id":"ga-policy-limits"}' > /tmp/out.json
jq '{deadline: .deadline.respond_by, statute: .deadline.statute_minimum.date, billed: .specials.billed, claimed: .specials.claimed}' /tmp/out.json
jq -r '.items[] | "\(.severity)\t\(.group)\t\(.title)"' /tmp/out.json
jq '{record}' /tmp/out.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```

The sample `ga-policy-limits` is synthetic. Pass if:
- `.deadline.respond_by` is `"2026-09-03"` and `.deadline.statute_minimum.date` is `"2026-09-03"` (30 days from receipt, O.C.G.A. 9-11-67.1);
- an item says `$19,540.00` of the `$31,450.00` claimed is not supported by any bill;
- a note says the letter gives 10 days where the statute gives at least 30;
- no item or the memo values the claim;
- every receipt has `"status": "attested"` and the record verifies.

## 5. Point the app at the local API

- Base URL: `http://localhost:8445`. For a web app, set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`; add other origins to `DECOSA_CORS_ORIGINS`.
- `POST /demand/read` takes `{state, received: {date, method}, claim?: {claim_number?, loss_date?}, documents: [{kind: letter|bill|record, title?, provider?, pages: [text, ...]}]}`. It returns JSON, or streams Server-Sent Events with `Accept: text/event-stream`.
- `GET /demand/info` lists what is checked, the rules with their sources and dates, the limits and what is not checked.
- The server stores nothing. Keep each signed record (JSON) with your file. Anyone can re-check it with `POST /record/verify` against the key at `GET /attest/signing-key`. Case citations are looked up in the local citation index when `DECOSA_CITEINDEX_PATH` is set; with `DECOSA_PREFLIGHT_CL_OFF=1` nothing else is asked. Only the citation itself would ever leave the box, never the letter.
- Keep the API on 127.0.0.1. For other users on the LAN, put a TLS reverse proxy with authentication in front and set `DECOSA_TRUSTED_PROXIES`.

## 6. Hosted routes (off, and leave them off)

`DECOSA_LLM_ROUTE=gateway` would send prompts, which contain your files, to a hosted gateway. Never use it for real files.

Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A first read of a time-limited demand: it shows the respond-by date the letter sets and the statute's minimum, every condition of acceptance with its page, and the medical specials tied to the bills, and it never values the claim.
Who it's for
Bodily-injury adjusters and claim leads at carriers, TPAs and self-insured claims units.
Where it runs
Self-host or confidential access for real claim files; the hosted demo takes made-up packages only
Key numbers

On a held-out set of 8 synthetic packages it got the respond-by date and the statute's minimum right on all 7 time-limited demands, found all 31 conditions of acceptance and summed the bills to within $1 on all 8; one author wrote the packages and the set is small.

  • 7 of 7 Respond-by date exact (test split, n = 7)
  • 7 of 7 Statute's minimum exact (test split, n = 7)
  • 31 of 31 Conditions of acceptance found (test split, n = 31)
  • 161.0 s Median end-to-end run, hosted (QA sweep 2026-09-29)
All results, datasets and caveats
Models
Qwen3.8-27B (reads the letter, the bills and each record page; the grounding judge checks every sentence it writes)
Where
Self-host or confidential access for real claim files; the hosted demo takes made-up packages only
Checks
Every quote and amount found on the page it cites; dates and totals computed in code; receipt per model call; signed record of page hashes and every item
Output
Notes, reports and drafts · Signed record or verdict
Data
Personal data · Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

How does it work out a time-limited demand's deadline?

From the letter's own words (a date, or N days from receipt, the letter's date or transmission) and the date you received it, in code. Next to it is the statute's minimum for California, Georgia, Missouri, Montana and Utah, with the section quoted. Other states say 'not encoded: check with counsel'.

Will it tell me what the claim is worth?

No. It never values a claim or recommends paying, offering or accepting, and any sentence that would is removed. It shows what the letter asks, what the bills support and what is missing; the adjuster decides.

Does it judge whether a letter was written by AI?

No. It notes things to check: a statute from another state, a case that isn't at its cite, a sentence that misstates the law, hidden characters. Those are notes for the adjuster, not a judgment of the letter or who wrote it.

Can I use it on real claim files?

Not on the hosted demo, which takes made-up packages only. Real files hold medical records: run it on your own hardware or ask for confidential access.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Injury demand reader

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.