Skip to content
decosa
LiveHostedSelf-hostSelf-host first for real data

Review a claim file for conduct

A signed review of each claim file: every deadline worked out from the file's own dates, and each denial reason checked against the policy wording.

Held-out test52 of 54 (precision 0.98, recall 0.96)Planted problems found in test claim files, all checks (held-out test)
On production10 smedian on production (2026-09-26); slower when the service is busy
List price~$0.40 per 100 claim filesmeasured, at list price

Built on: Grounding, Typed judgment, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the deadline rules and the record need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the insurance claims-file conduct pack API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
claims-conduct-pack

Use the hosted API

# Decosa claims-file conduct pack: use the hosted API

You are wiring Decosa's claims-file conduct review into this project. It takes a claim file (first notice of loss,
adjuster notes, letters, estimates) and the policy wording, and returns a conduct checklist with a signed file review.
Every model call has a signed receipt. Use only what is listed below; if you need something else, stop and ask me.

The checklist covers:
- acknowledgement, investigation, decision, status-letter, payment and reply deadlines, computed from the file's own dates
  under the NAIC model (#900/#902), California (10 CCR 2695) or Texas (Ins. Code 542);
- each denial reason, checked for a cited provision and grounded in the policy;
- the Department of Insurance review notice;
- typed judgments on misrepresentation, low offers and missing investigation;
- a record of AI use.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **Hosted use is for synthetic claim files only.** Real claim files belong on a self-hosted box (see the self-host
  prompt). Say so wherever this is wired in.
- This is triage for a claims-quality reviewer. Never label a file "compliant", and never use it to decide coverage.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "claims-conduct-pack"}` returns `{"token", "expires_at", "budget"}`.
   - The demo allows a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz) and a token allowance per session (the `budget` in the session response).
   - Over a limit you get HTTP 429 with `Retry-After`.
   - A demo token runs one review at a time (409 otherwise).
3. A review needs about 4,500 generated tokens left in the budget before it starts (402 otherwise); it usually uses
   fewer, and only what it uses is charged.

## Endpoints
- `POST /claims/check` (token). Body: `{"claim": {"id"?, "line": "auto"|"homeowners"|"other", "state": "NAIC"|"CA"|"TX", "party"?: "first"|"third"}, "documents": [{"kind": "fnol"|"notes"|"letter"|"estimate"|"email"|"other", "title"?, "date"?: "YYYY-MM-DD", "text"}], "policy": {"title"?, "text"}, "timeline"?, "ai_use"?, "title"?}` or `{"sample_id": "..."}`.
  - Limits: 16 documents, 20,000 characters each, 60,000 in total; a policy of 40,000 characters; 512 KB of JSON.
  - `timeline` (optional): `[{event, date, decision?, note?}]` from your claims system. Those dates are used as given,
    instead of the model reading them from the notes. Events: `notice_of_claim`, `acknowledgement`, `forms_sent`,
    `investigation_started`, `proof_of_loss`, `more_time_notice`, `decision` (with `decision`: accept, deny or partial),
    `payment`, `claimant_communication`, `insurer_reply` and `objection`.
  - `ai_use` (optional): `[{step, system, decided_by: "model"|"human"|"model_then_human", note?}]`.
  - The JSON response has:
    - `checklist`: `[{id, group, title, rules, status: "flag"|"review"|"ok"|"na", actor: "rule"|"model"|"model+rule", reason, computed?, evidence?: [{ref, quote, text}], grounding?: {verdict, spans: [{span, quote}]}, answer?, probability?, receipt_ids?}]`;
    - `events` (the dated events used, each with its line), `dropped`, `counts`, `steps`, `ai_use`, `review_md`,
      `record`, `record_check`, `receipts` and `note`.
  - Item ids: `ack`, `investigation_start`, `decision`, `payment`, `replies`, `denial_written`, `reason:N`,
    `review_notice`, `misrepresent`, `lowball`, `investigation`, `explain`, `ai_use:N`.
  - With `Accept: text/event-stream` (or `"stream": true`), the events are:
    - `ready`;
    - a `receipt` per model call, then `events`, then an `item` per checklist item;
    - then `result`, `budget` and `done`.
- `POST /claims/signoff` (token) `{"record", "reviewer": {"name", "role"?}, "decisions": [{"id", "decision": "confirm"|"dismiss"|"escalate", "note"?}]}`
  returns a second signed record pointing at the first. The reviewer's name is as given; identity is not verified.
- `POST /record/verify` (no token) `{"record": {...}}` returns `{ok, summary, checks, first_bad}`.
- `GET /claims/info`, `GET /claims/samples` and `GET /attest/signing-key` need no token.

## Example: review a file and print what needs a person's eyes (Python, `pip install httpx`)
```python
import httpx, json, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {"claim": {"id": "TEST-1", "line": "homeowners", "state": "CA"},
        "documents": json.load(open("claim-documents.json")), "policy": {"text": open("policy.txt").read()}}
r = httpx.post(f"{API}/claims/check", json=body, headers=H, timeout=300)
r.raise_for_status()
js = r.json()
for x in js["checklist"]:
    if x["status"] in ("flag", "review"):
        print(f"{x['status'].upper()}: {x['title']}: {x['reason']}")
open("file-review.md", "w").write(js["review_md"])
json.dump(js["record"], open("file-review.json", "w"))   # keep with the claim file
```

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa claims-file conduct pack: run it yourself (containers)

You are setting up the Decosa claims-file conduct pack on this machine, so claim files never leave it. It reads a claim
file and its policy wording and returns a conduct checklist:
- deadlines computed from the file's own dates (NAIC model, California or Texas);
- denial reasons grounded in the policy;
- the Department of Insurance review notice;
- typed judgments on misrepresentation, low offers and missing investigation;
- a record of any AI use.

It seals a signed file review. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/claims-conduct-pack.zip (4 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py claims-conduct-pack` (the api image carries the same bundle under /app/rehearsal/claims-conduct-pack/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py claims-conduct-pack --bundle claims-conduct-pack.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the late acknowledgement is flagged", "its deadline is 15 calendar days after the first notice", "the denial reason that misstates Exclusion 4 is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - keep its data on a named volume;
   - bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check. The first start
   downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/claims/info` lists every rule with its source and the date it was read.
   `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my file
   reviews.
5. Smoke test: get a token with `POST /demo/session {"vertical":"claims-conduct-pack"}`, then send
   `POST /claims/check {"sample_id": "ca-auto-planted"}`. Expect:
   - `ack`, `reason:1` and `review_notice` as `flag`;
   - the `reason:1` grounding span quotes "organized race or speed contest";
   - every receipt `attested`.

   Then `POST /record/verify {"record": <record>}` should give `ok: true`, and `{"sample_id": "ca-home-clean"}` should
   come back with no flags.
6. Report back: the public key and key id, the smoke-test results, and how long the review took.

Claim files hold personal and often health data. This is triage for a claims-quality reviewer. It never says a file is
compliant, and it is not legal advice.

Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds claim
files. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVIDIA NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns

    The standard tier fits (57.6 of 192 GB). The best tier fits too.

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"claims-conduct-pack"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py claims-conduct-pack

Download the mock-data bundle (4 KB, 8 checks)expected.json

A synthetic California auto claim file (first notice, adjuster log, denial letter) and its policy. Planted: the claim was acknowledged 29 days after notice (the limit is 15), the denial letter says Exclusion 4 excludes loss at excessive speed when the policy excludes organized races and speed contests, and the letter has no Department of Insurance review notice. All three must be flagged, with the deadline computed and the policy line quoted, and the signed file review must verify.

What the rehearsal checks
  • the late acknowledgement is flagged
  • its deadline is 15 calendar days after the first notice
  • the denial reason that misstates Exclusion 4 is flagged
  • the flag quotes the policy line it was checked against
  • the missing Department of Insurance review notice is flagged
  • the signed file review verifies
  • a file review with its flag count changed no longer verifies
  • every model call has a signed receipt

Licence: Synthetic: Harbor Oak Mutual, its policy forms and every person are invented (scripts/claims_cases.py). Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa claims-file conduct pack: run it yourself (containers)

You are setting up the Decosa claims-file conduct pack on this machine, so claim files never leave it. It reads a claim
file and its policy wording and returns a conduct checklist:
- deadlines computed from the file's own dates (NAIC model, California or Texas);
- denial reasons grounded in the policy;
- the Department of Insurance review notice;
- typed judgments on misrepresentation, low offers and missing investigation;
- a record of any AI use.

It seals a signed file review. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/claims-conduct-pack.zip (4 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py claims-conduct-pack` (the api image carries the same bundle under /app/rehearsal/claims-conduct-pack/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py claims-conduct-pack --bundle claims-conduct-pack.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the late acknowledgement is flagged", "its deadline is 15 calendar days after the first notice", "the denial reason that misstates Exclusion 4 is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - keep its data on a named volume;
   - bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check. The first start
   downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/claims/info` lists every rule with its source and the date it was read.
   `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my file
   reviews.
5. Smoke test: get a token with `POST /demo/session {"vertical":"claims-conduct-pack"}`, then send
   `POST /claims/check {"sample_id": "ca-auto-planted"}`. Expect:
   - `ack`, `reason:1` and `review_notice` as `flag`;
   - the `reason:1` grounding span quotes "organized race or speed contest";
   - every receipt `attested`.

   Then `POST /record/verify {"record": <record>}` should give `ok: true`, and `{"sample_id": "ca-home-clean"}` should
   come back with no flags.
6. Report back: the public key and key id, the smoke-test results, and how long the review took.

Claim files hold personal and often health data. This is triage for a claims-quality reviewer. It never says a file is
compliant, and it is not legal advice.

Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds claim
files. If I ask for it later, follow the Provide page instead of improvising.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsInsurance claims-file conduct pack on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

Standard · the hosted demo, one 96 GB card: what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Reads the file: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Insurance claims-file conduct pack, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Insurance claims-file conduct pack on my hardware

Fetch https://decosa.ai/prompts/claims-conduct-pack-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=claims-conduct-pack)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Reads the file: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/claims-conduct-pack-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 10 s · ~$0.004 per run · 5 receipts

Loading the nightly status…

Self-host: verified 26 Sep 2026 · fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

Measured cost to run: about $0.40 per 100 claim files (hosted, 26 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): the planted file flagged ack, reason:1 (quoting 'organized race or speed contest'), review_notice and misrepresent, 5 attested receipts, record verified, 5.0 s; the clean file had 0 flags; sign-off pointed at the first record; the rehearsal bundle passed 8/8. Model-server startup was not re-run.

Known limits (5)
  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.
  • Measured on 84 synthetic, templated files written by the building agent; not on real claim files or with a claims auditor's labels.
  • Three rulepacks only (NAIC model, California, Texas). The NAIC pack is a baseline, not any state's law.
  • Business days skip weekends and US federal holidays; deadlines on weekends are not rolled forward, and one or two days late on such a deadline is marked review.
  • The reviewer's name on a sign-off is as given; identity is not verified.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the deadline rules and the record need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the insurance claims-file conduct pack API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.
The open stack

Every claim file read the way a market-conduct examiner reads it: deadlines from the file's own dates, denial reasons grounded in the policy, a signed file review.

For claims-quality teams, chief compliance officers and TPAs. Give it a claim file (first notice of loss, adjuster notes, letters, estimates) and the policy wording, and pick the rules: the NAIC model (Model #900 and #902), California (10 CCR 2695) or Texas (Insurance Code ch. 542). It computes the acknowledgement, investigation, decision, status-letter, payment and reply deadlines from the dates in the file, each with the line it came from. It checks every denial reason for a cited provision, grounds the letter's statement of the policy in the policy text, and looks for the Department of Insurance review notice where the state requires it. It gives typed yes/no judgments with quotes on misrepresenting the policy, an offer below the insurer's own valuation with no reason in the file, and a denial without a documented investigation. It records any AI system that touched the claim and which steps were a model, a rule or a person, and seals it all in a signed file review; a reviewer's sign-off is a second signed record. Triage for a person, never a compliance finding.

Deployment
Self-host first
Regulatory
Rules read from primary texts on 26 Sep 2026. NAIC Unfair Claims Settlement Practices Act (Model #900, adopted 1990, amended 1991) and Unfair Property/Casualty Claims Settlement Practices Model Regulation (Model #902, 1990, amended 1991): model laws that bind nobody until a state enacts them; #902 asks for acknowledgement within 15 days, a decision within 21 days of proof of loss or a written reason for more time, then a letter every 45 days, payment within 30 days, and a denial that references the provision it relies on. California Fair Claims Settlement Practices Regulations, 10 CCR 2695.5 and 2695.7 (operative 1993; 2695.7 last amended 14 Jul 2021): acknowledge, send forms and begin investigating within 15 calendar days, accept or deny within 40 calendar days of proof of claim or give written notice of the need for more time every 30 days, pay within 30 days, list every basis and the provision in a written denial, and tell the claimant they may have the matter reviewed by the California Department of Insurance, with its address and phone number. Texas Insurance Code 542.055 to 542.058 (eff. 1 Apr 2005): acknowledge within 15 days, accept or reject within 15 business days of receiving all items (or give reasons and decide within 45 days), state the reasons for a rejection, and pay within 5 business days. The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers (adopted 4 Dec 2023) covers claim administration and says the Department may ask for documentation of AI systems used in decisions; it applies in a state once that state issues it. Business days skip weekends and US federal holidays, an assumption. Not legal advice.
Architecture
Text description

A claim file (first notice of loss, adjuster log, letters, estimates) and the policy wording go in, with the rulepack chosen: NAIC model, California or Texas. Qwen3.8-27B extracts dated events, denial reasons, amounts and AI mentions, each with a line reference and a quote; code checks every quote and date against the file and drops the rest. Plain code computes the deadlines (calendar or business days) and checks the review notice. The grounding judge checks each denial reason's statement of the policy against the policy text, and typed judgments answer the conduct questions with quotes. Outputs: a conduct checklist and a signed, hash-chained file review that marks each step as rule, model or human; a reviewer's sign-off is a second signed record. Hosted calls get gateway-signed receipts; a self-hosted box signs with its own key.

Architecture

At a glance

What it checks
Acknowledgement, investigation, decision, status-letter, payment and reply deadlines from the file's own dates; denial reasons (a cited provision that exists and says what the letter says); the Department of Insurance review notice; misrepresentation, low offers and missing investigation; AI use. Rulepacks: NAIC model, California, Texas.
What it does not do
It never says a file is compliant and does not decide coverage. Not checked: total-loss valuation, subrogation, statute-of-limitations notices, fraud-investigation extensions, surplus lines, 'final payment' wording, other states. Scanned letters need OCR first.
Data retention
Nothing kept on the server. Files, policies and records live in memory for the request; logs carry counts only. You keep the signed file review with the claim file.
What leaves the box (hosted demo)
The claim file and policy go to Qwen3.8-27B through the Decosa API, whose receipts hold hashes, not text. Self-hosted, nothing leaves. The hosted demo is for synthetic files only.
Model calls per file
One extraction, one grounding call per denial reason, and up to four typed judgments: 2 to 6 calls.
Typical run cost
A fraction of a cent at the gateway list price for the planted California file (a handful of model calls). Each run shows its own measured cost.
Who did what
The signed record marks each step as rule, model or human; a reviewer's confirm, dismiss or escalate decisions become a second signed record pointing at the first.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 48 GB card

    The same pipeline on a smaller mixture-of-experts model. Faster and cheaper; accuracy on this task unknown.

    Models
    • Gemma 4 26B A4B (instruction-tuned)
    Hardware
    1x L40S or RTX 6000 Ada 48 GB (not measured)
    Quality evidence
    • planted problems flagged / false alarmsnot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
  • In the hosted demo

    Standard

    the hosted demo, one 96 GB card

    Qwen3.8-27B reads the file, grounds the denial and answers the conduct questions; the deadlines are plain code. Every model call receipted.

    Models
    • Qwen3.8-27B (NVIDIA NVFP4)
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • held-out test, 48 synthetic files run once: planted problems flagged / false flags52/54 / 1decosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route; prompts frozen on a separate 24-file dev set; half the test files use phrasing never seen while writing the prompts (26/26 found there)
    • clean files with any flag0/12 test, 0/7 dev, 0/6 replies, 0/2 samplesdecosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route
    • date accuracy on the test set: events with the right date / timeliness items with the right start, act and deadline247/247 / 122/122decosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route
    • per check on the test set (found/planted)acknowledgement 13/13, decision 2/2, payment 4/4, denial reason 11/12, review notice 2/2, misrepresentation 10/11 (1 false), low offer 3/3, investigation 7/7; late replies 6/6 on a 12-file targeted setdecosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route
    • AI use recorded, and whether a person was involved read right31/31decosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route
    • real, de-identified claim files reviewed by a claims-quality auditornot measured yet
    Latency
    measured on our server under a shared gateway: seconds per file, longer at the slow end
    Verification
    Proof: strongHosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.
  • Best

    DeepSeek-V4-Flash on two more cards

    A larger model for long, multi-claimant files.

    Models
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    Hardware
    2x RTX PRO 6000 96 GB
    Quality evidence
    • planted problems flagged / false alarmsnot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyNot a hosted model: calls are attested by the box's key only.
  • Needs more compute

    Wanted: the best setup

    two large judges from different families

    DeepSeek-V4-Flash and GLM-5.3-Flash each read the file, and a finding stands when they agree; disagreements go to the reviewer. Claim files stay on your own hardware, never on community providers. Not served yet.

    Models
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    • GLM-5.3-Flash
    Hardware
    Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash plus 2x 96 GB for GLM-5.3-Flash, or one Mac Studio with 512 GB holding both 4-bit builds (156 + 165 GB, sizes from our Mac; not run together yet). Estimate.
    Quality evidence
    • planted problems flagged / false alarmsnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
After the session
Reads the file (dated events, denial reasons, amounts, AI mentions, each with a quote), grounds each denial reason in the policy, and answers the typed conduct questionsQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
Lite tier: the same extraction, grounding and judgments on a 48 GB cardGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
25.2B (3.8B active)No proof yetSelf-host only
Best tier: a larger model for long, multi-claimant filesDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active) · 192 GBNo proof yetSelf-host only
Other
Second judge, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only

Measured

How well does it do on synthetic claim files?

84 synthetic files from a generator with planted problems (late acknowledgement, a clause misquoted, no review notice, a low offer, no investigation and more), across the NAIC, California and Texas rulepacks. Prompts were written on a 24-file dev set; the 48-file test set was run once.

Planted problems flagged, test set
52 of 541 false flag; the miss came back as review
Clean files with any flag
0 of 27test, dev, replies and demo sets together
Deadlines right end to end
122 of 122start date, act date and deadline, test set
Cost per file
about $0.0045 model calls, 9,351 tokens on the demo file, gateway list price

Where it fails

Both errors are the model reading a letter against the policy: a letter that widened 'livery conveyance' to 'any business purpose' came back as review, not flag, and a letter that shortened the flood exclusion was called a misrepresentation.

What it does not show

The files are templated and synthetic. Real files are longer and messier (scanned letters, email chains, several claimants). Measure it on your own closed files before relying on it.

Source: decosa-api docs/evals/claims-conduct-pack.md, 26 Sep 2026

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    File intake, extraction checks, the deadline rules, grounding, the judgments, signing and the HTTP API (/claims/*). No GPU. Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server.

  • 1x L40S / RTX 6000 Ada 48 GB

    Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite).

  • CPU only Fits

    Date arithmetic, the notice check, the signed record and verification need no GPU; reading the file needs the model.

Latency per lane

  • one claim file (3-6 documents), busy shared gateway10.1 s

    Measureddecosa-api docs/evals/claims-conduct-pack.md, test set p50 on our server 2026-09-26 (p90 19.2 s), gateway route

  • one claim file, self-hosted direct route5.0 s

    Measuredmeasured on our server 2026-09-26, self-host sandbox, ca-auto-planted

  • deadline rules and the signed record50 ms

    Estimateestimate: no model call

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

claims-conduct-pack/assemble-prompt.md199 lines
# Assemble the Decosa claims-file conduct pack on this machine

You are setting up a self-hosted claims-file conduct review on this Linux machine for an insurer's claims-quality team or a
TPA. It reads a claim file (first notice of loss, adjuster notes, letters, estimates) and the policy wording. It returns a
conduct checklist:
- acknowledgement, decision and payment deadlines computed from the file's own dates, under the NAIC model (Model #900 and
  #902), California (10 CCR 2695) or Texas (Insurance Code ch. 542);
- each denial reason checked for a cited provision, and grounded in the policy wording;
- the Department of Insurance review notice, where the state requires it;
- typed judgments on misrepresenting the policy, low offers and missing investigation;
- a record of any AI system that touched the claim.

Everything is sealed in a signed, hash-chained file review. A person's sign-off is a second signed record.

Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- Claim files hold personal, financial and often health data. Keep everything on this machine: the model route stays
  local (`direct`), and nothing goes to a hosted service.
- This is triage for a claims-quality reviewer. It never says a file is compliant, it does not decide coverage, and it is
  not legal advice. Every flag needs a person to read the file.

Repeat both points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/claims-conduct-pack.zip (4 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py claims-conduct-pack` (the api image carries the same bundle under /app/rehearsal/claims-conduct-pack/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py claims-conduct-pack --bundle claims-conduct-pack.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the late acknowledgement is flagged", "its deadline is 15 calendar days after the first notice", "the denial reason that misstates Exclusion 4 is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
   - Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main`. On a 48 GB card, also
     set `LLM_MAX_LEN=32768`. Not measured.
   - Under 48 GB: stop and tell me it will not fit.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the
   NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories. Then run
   `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 60 GB of free disk.

## 2. Get the images

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`.

If a pull fails, build from source once the `decosa-api` source is published:
- clone it;
- in the clone, run `docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`;
- run `docker compose build llm` from its compose file.

If neither works, stop and tell me.

## 3. Write the compose file

Create `~/decosa-claims/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these file reviews, e.g. Example Mutual claims quality>"
```

Create `~/decosa-claims/docker-compose.yml` with exactly these services:

```yaml
name: decosa-claims
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"                # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "200000"        # per session; one file review needs up to about 4,500 generated tokens
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_CLAIMS_MAX_CONCURRENT: "3"
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

Use the named volume `decosa-data` exactly as written. A host bind mount owned by root makes the API fail on
`/data/keys.sqlite`.

Run `docker compose up -d`, then poll `docker compose ps` until both services are healthy. The LLM takes 5-10 minutes
the first time. `curl -s localhost:8445/healthz` should show `"llm": true` (`asr` is false here, which is fine).

## 4. Smoke test

```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"claims-conduct-pack"}' | jq -r .token)
curl -s $API/claims/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
  -d '{"sample_id":"ca-auto-planted"}' > /tmp/claims.json
jq '{counts, receipts: (.receipts|length), statuses: [.receipts[].status] | unique}' /tmp/claims.json
jq -r '.checklist[] | "\(.status)\t\(.id)\t\(.reason)"' /tmp/claims.json
jq '{record}' /tmp/claims.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```

This sample is a synthetic California auto denial (fictional insurer Harbor Oak Mutual). Pass if:
- `ack` is `flag`, and its reason says it was acknowledged 29 days after notice (deadline 2026-06-03);
- `reason:1` is `flag`, and `.grounding.spans[0].quote` contains "organized race or speed contest";
- `review_notice` is `flag`;
- every receipt has `"status": "attested"`;
- the record verifies (`ok: true`).

Then run the clean file:

```bash
curl -s $API/claims/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
  -d '{"sample_id":"ca-home-clean"}' | jq '.counts'
```

Pass if `flag` is 0.

A person's sign-off makes a second signed record:

```bash
jq '{record, reviewer: {name: "QA reviewer", role: "claims quality"}, decisions: [{id: "ack", decision: "confirm"}]}' /tmp/claims.json \
  | curl -s $API/claims/signoff -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @- | jq '.record.statement.review_of'
```

Pass if it prints the first record's root hash.

## 5. Point the app at the local API

- Base URL: `http://localhost:8445`. For a web app, set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`. Add other origins to
  `DECOSA_CORS_ORIGINS`.
- `POST /claims/check` takes `{claim: {id?, line: auto|homeowners|other, state: NAIC|CA|TX}, documents: [{kind, title?,
  date?, text}], policy: {title?, text}, timeline?, ai_use?}`. It returns JSON, or streams Server-Sent Events when asked
  with `Accept: text/event-stream`.
- `GET /claims/info` lists every rule with its source and the date it was read, the deadlines, the conventions and what
  is not checked.
- The server stores nothing. Keep each signed file review (JSON) with the claim file. Anyone can re-check it with
  `POST /record/verify` against the key at `GET /attest/signing-key`.
- If the claims system exports a dated activity log, send it as `timeline`: those dates are used as given, and the model
  does not read them from the notes.
- Keep the API on 127.0.0.1. For other users on the LAN, put a TLS reverse proxy with authentication in front and set
  `DECOSA_TRUSTED_PROXIES`.

## 6. Keep the direct route

`DECOSA_LLM_ROUTE=gateway` would send prompts to the hosted Decosa
API, and those prompts contain the claim file and the policy.
Never use it for real claims. At most, use it for synthetic training material.

Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

5 laws, rules and guidance pages cited; 2 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A claims file review for unfair claims practices rules: it reads a claim file and the policy the way a market-conduct examiner does, works out every deadline from the file's own dates, checks each denial reason against the policy, and seals the result in a signed file review for a person to confirm.
Who it's for
Teams in finance and insurance and compliance and trust.
Where it runs
Self-host for real claim files (hosted demo: synthetic files only)
Key numbers

On 48 synthetic test files, run once after the prompts were frozen, it found 52 of 54 planted problems with one false flag, and all 26 in the files with held-out phrasing. The files are templated and written by the same author as the prompts, so the numbers do not predict accuracy on a carrier's files.

  • 26 / 0 / 0 Held-out phrasing files: found / false / missed (held out, n = 24)
  • 52 of 54 (precision 0.98, recall 0.96) Planted problems found, all checks (test split, n = 54)
  • 1 False flags (test split, n = 48)
  • 10.1 s Median end-to-end run, hosted (QA sweep 2026-09-26)
All results, datasets and caveats
Models
Qwen3.8-27B
Where
Self-host for real claim files (hosted demo: synthetic files only)
Checks
Receipt per model call; every quote and date checked against the file; signed hash-chained file review with each step marked rule, model or human
Output
Signed record or verdict · Structured data
Data
Personal data · Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

Which rules does it check?

Three rulepacks: the NAIC model (Model #900 and #902), California (10 CCR 2695) and Texas (Insurance Code ch. 542). The NAIC pack is a baseline, not any state's law. Rules not checked, such as total-loss valuation and subrogation, are listed.

How long does an insurer have to acknowledge a claim?

Under the rules this checks: 15 days in the NAIC model regulation (#902); 15 calendar days in California (10 CCR 2695.5), which also means sending forms and starting the investigation; and 15 days in Texas (Insurance Code 542.055). The review works out each deadline from the dates in the file and shows the line each date came from. Not legal advice.

How are the deadlines worked out?

From the dates in the file, each with the line it came from. Business days skip weekends and US federal holidays, which is an assumption; deadlines on weekends are not rolled forward, and one or two days late on such a deadline is marked for review. On the synthetic test set, 122 of 122 timeliness items had the right start, act and deadline.

What does it check in a denial letter?

That every reason cites a provision that exists and says what the letter says (the letter's sentence is grounded against the policy text), and that the Department of Insurance review notice is there where the state requires it.

Will it tell me a file is compliant?

No. It is triage for a person and does not decide coverage. A reviewer's confirm, dismiss or escalate decisions become a second signed record pointing at the first.

How accurate is the claims file review?

On 48 synthetic test files it found 52 of 54 planted problems with one false flag, and flagged none of 12 clean files. The files are templated and written by the same author as the prompts, so the numbers do not predict accuracy on a carrier's files.

Where does the claim data go?

Self-hosted, nothing leaves your hardware. The hosted demo is for synthetic files only; there, the file and policy go to Qwen3.8-27B through our gateway, whose receipts hold hashes, not text. Nothing is kept on the server.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Insurance claims-file conduct pack

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.