Skip to content
decosa
LiveHostedSelf-hostMacSelf-host first for real data

Check HCC codes for MEAT evidence

A keep, hold or delete verdict for each submitted code, with the MEAT sentences quoted, and an evidence file for the audit folder.

Held-out test189 / 200Right keep, hold or delete verdict per code (held-out test)
On production24 smedian on production (2026-09-26); slower when the service is busy
List price~$0.52 per 100 membersmeasured, at list price

Built on: Grounding, Typed judgment, Signed record, Document reader

1. Pick a sample

Sample

Real patient data: request confidential access, or run the tool on your own hardware. The demo takes samples or made-up data only.

2. Run it

On production the sample took 24 s (median, 2026-09-26). Slower when the service is busy.

Result

The answer appears here first, then what it found, the draft, and how long it took. Sample: Mixed file: 8 codes, 3 records an audit would not accept (synthetic).

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for Qwen3.8-27B; the V28 map, the record checks and the signed record run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the hcc evidence file and radv defence API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
hcc-evidence-file

Use the hosted API

# Decosa HCC evidence file: use the hosted API (synthetic members only)

You are wiring Decosa's HCC evidence file into this project. For one member and service year it takes the diagnosis
codes the plan submitted and the member's notes, and for each code that maps to a CMS-HCC V28 payment HCC it says
supported (keep), insufficient (hold for a certified coder) or not supported (delete: do not submit), with the MEAT
sentences (monitor, evaluate, assess, treat) quoted from the notes or their absence stated. It checks each note and
addendum as a RADV reviewer would, reports the net effect, and returns a Markdown evidence file and a record signed by
the server, with a signed receipt for every model call. Use only what is listed below. If you need something else, stop
and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes synthetic members only.** Clinical notes are protected health information. Every request must
  carry `"synthetic": true`; anything else gets HTTP 400. Real members belong on a self-hosted box (see the self-host
  prompt). Never send real names, member ids or notes here.
- It checks only the codes you submit. It has no mode that searches notes for codes to add, suggests codes or drafts
  provider queries, and a request with fields such as `mode`, `suggest`, `find`, `opportunities` or `gaps` gets HTTP 400.
  Do not build that feature around it. Show deletions at least as prominently as confirmations, and say that a certified
  coder decides.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "hcc-evidence-file"}` returns `{"token", "expires_at", "budget"}`.
   Sessions per IP are limited; over a limit you get HTTP 429 with `Retry-After`.
3. A member needs about 800 generated tokens of budget per code that maps to an HCC (402 otherwise). One review at a
   time per demo token (409 while one is going).

## Endpoints
- `POST /hcc/review` (token). Body:
  ```json
  {"member": {"id": "SYN-1", "service_year": 2025},
   "codes": ["E11.22", "I50.22"],
   "notes": [{"id": "N1", "date": "2025-03-14", "kind": "physician", "provider": "...", "credential": "MD",
              "specialty": "Internal Medicine", "signed": {"by": "...", "credential": "MD", "date": "2025-03-14"},
              "addenda": [{"date": "2025-04-01", "by": "...", "credential": "MD", "text": "..."}], "text": "..."}],
   "synthetic": true, "stream": false}
  ```
  `kind`: physician, hospital_outpatient, hospital_inpatient, telehealth_video, telehealth_audio, diagnostic_report, lab,
  home_health, hospice, superbill, other. Up to 20 codes and 12 notes (12,000 characters each, 60,000 in total).
  `{"sample": "mixed-file", "synthetic": true}` runs a built-in sample.
- JSON answer: `{status: deletions | holds | all_supported, net_effect: {codes: {submitted, delete, hold, keep,
  not_reviewed}, hccs: {rows: [{hcc, label, outcome, codes}]}, text}, codes: [{code, dotted, description, hccs, verdict:
  not_supported | insufficient | supported, action: delete | hold | keep, reasons, summary, records: [{note, valid,
  issues, quotes: [{id, meat, text, start, end}], typed, grounding}]}] (deletions first), not_reviewed, notes (with the
  RADV record checks), guard, receipts, evidence_file_md, record}`.
- SSE: send `Accept: text/event-stream` (or `"stream": true`): `ready`, `receipt`, `code` (one per code), `report`,
  `budget`, `done`.
- `POST /hcc/codes` (no token, no model): `{codes: [...]}` → V28 payment HCCs and the ones that pay after the hierarchies.
- `POST /hcc/records` (no token, no model): the review body → the RADV record check for every note and addendum.
- `POST /record/verify` (no token): the `record` → `{ok, checks, summary}`.
- `GET /hcc/info`, `GET /hcc/samples` (no token): the rules, the V28 table's source, limits, sources and three samples.

## Errors
400 bad input (the message names the field or limit, says the member is not marked synthetic, or that no code maps to a
V28 HCC, or refuses an add mode), 401/403 token, 402 budget, 409 a review already going on this demo token, 413 body
over 256 KB, 429 busy (`Retry-After`).

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa HCC evidence file: run it yourself (containers)

You are setting up the Decosa HCC evidence file on this machine, so clinical notes never leave it. For each diagnosis
code a plan submitted for a member it says keep, hold or delete, with the MEAT sentences quoted from the notes or their
absence stated, checks each record as a RADV reviewer would, and returns the net effect, a Markdown evidence file and a
signed record. It checks only the submitted codes and never suggests codes to add. Nothing is sent to Decosa's hosted
API. It is a review aid: certified coders decide.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/hcc-evidence-file.zip (3 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py hcc-evidence-file` (the api image carries the same bundle under /app/rehearsal/hcc-evidence-file/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py hcc-evidence-file --bundle hcc-evidence-file.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the 2021 heart attack coded as acute is not supported (delete)", "breast cancer treated in 2016 is not supported (delete)", "morbid obesity is held: its only evidence is an audio-only call"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, with prefix caching on) and the `api` service. For the `api`
   service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
   `DECOSA_HCC_SYNTHETIC_ONLY=0` (so this box accepts real members) and bind every port to 127.0.0.1. Never set the
   gateway route on this box: it would send PHI to the Decosa API.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/hcc/info` shows `synthetic_only: false`, `judgment_method: "logprobs"` and
   the V28 table's source; `GET /attest/signing-key` shows this box's public key. Show me the key: it is what an auditor
   pins to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"hcc-evidence-file"}` and send
   `{"sample": "mixed-file", "synthetic": true}` to `POST /hcc/review`. Expect I21.4 and C50.911 deleted (history),
   E66.01 and J44.9 held with reason `invalid_record` (an audio-only call, a diagnostic radiologist's report), E11.22 and
   I50.22 kept with quotes, and I10 under `not_reviewed`. Send the record to `POST /record/verify`: `ok` must be true.
   Send the same request with `"mode": "opportunities"`: it must be refused with 400.
6. Report back: the public key, the net effect, the held and deleted codes with their reasons, and how long it took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds PHI.
If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/hcc-evidence-file-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMlite tierRuns with a smaller tier

    The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (32 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (32 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"hcc-evidence-file"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py hcc-evidence-file

Download the mock-data bundle (3 KB, 15 checks)expected.json

A synthetic Medicare Advantage member (service year 2025) with nine submitted diagnosis codes (eight map to CMS-HCC V28 payment HCCs) and four notes plus an addendum: an internist's visit, a cardiology visit, an audio-only phone call, a diagnostic radiologist's CT report, and an addendum a coder wrote five months after the visit. The review must delete the 2021 heart attack coded as acute and the breast cancer treated in 2016, hold the codes whose only evidence is in the phone call and the radiology report, keep diabetes with CKD, CKD 3b and heart failure with MEAT quotes, report the net effect, refuse an 'opportunities' mode, and sign a record that verifies and fails when changed.

What the rehearsal checks
  • the 2021 heart attack coded as acute is not supported (delete)
  • breast cancer treated in 2016 is not supported (delete)
  • morbid obesity is held: its only evidence is an audio-only call
  • COPD is held: its only evidence is a diagnostic radiologist's report
  • diabetes with CKD is kept
  • chronic systolic heart failure is kept
  • at least two codes are to be deleted
  • the net effect covers all eight HCC codes
  • the code outside the model (I10) is listed as not reviewed
  • the coder's late addendum is not an acceptable record
  • the evidence file lists deletions first
  • an opportunities mode is refused
  • the signed record verifies
  • a record with its status changed no longer verifies
  • every model call has a signed receipt

Licence: Synthetic: the member, providers and notes are invented (decosa_api/verticals/hcc/synth.py). ICD-10-CM and the CMS-HCC V28 mapping are CMS/CDC public-domain works. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa HCC evidence file: run it yourself (containers)

You are setting up the Decosa HCC evidence file on this machine, so clinical notes never leave it. For each diagnosis
code a plan submitted for a member it says keep, hold or delete, with the MEAT sentences quoted from the notes or their
absence stated, checks each record as a RADV reviewer would, and returns the net effect, a Markdown evidence file and a
signed record. It checks only the submitted codes and never suggests codes to add. Nothing is sent to Decosa's hosted
API. It is a review aid: certified coders decide.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/hcc-evidence-file.zip (3 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py hcc-evidence-file` (the api image carries the same bundle under /app/rehearsal/hcc-evidence-file/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py hcc-evidence-file --bundle hcc-evidence-file.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the 2021 heart attack coded as acute is not supported (delete)", "breast cancer treated in 2016 is not supported (delete)", "morbid obesity is held: its only evidence is an audio-only call"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, with prefix caching on) and the `api` service. For the `api`
   service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
   `DECOSA_HCC_SYNTHETIC_ONLY=0` (so this box accepts real members) and bind every port to 127.0.0.1. Never set the
   gateway route on this box: it would send PHI to the Decosa API.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/hcc/info` shows `synthetic_only: false`, `judgment_method: "logprobs"` and
   the V28 table's source; `GET /attest/signing-key` shows this box's public key. Show me the key: it is what an auditor
   pins to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"hcc-evidence-file"}` and send
   `{"sample": "mixed-file", "synthetic": true}` to `POST /hcc/review`. Expect I21.4 and C50.911 deleted (history),
   E66.01 and J44.9 held with reason `invalid_record` (an audio-only call, a diagnostic radiologist's report), E11.22 and
   I50.22 kept with quotes, and I10 under `not_reviewed`. Send the record to `POST /record/verify`: `ok` must be true.
   Send the same request with `"mode": "opportunities"`: it must be refused with 400.
6. Report back: the public key, the net effect, the held and deleted codes with their reasons, and how long it took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds PHI.
If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/hcc-evidence-file-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsHCC evidence file and RADV defence on GeForce RTX 5090: use the Standard · one GPU for the model (hosted demo) tier

The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB (fits): Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; not run for this use case.

Standard · one GPU for the model (hosted demo): what changesuses estimates

  • Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • V28 map and hierarchies, RADV record checks,...: decosa-api HCC review (decosa_api/verticals/hcc), importing the grounding judge (vertical 17) and the typed-judgment engine (vertical 24). CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Evidence sentences with MEAT tags, the typed...: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for HCC evidence file and RADV defence, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up HCC evidence file and RADV defence on my hardware

Fetch https://decosa.ai/prompts/hcc-evidence-file-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=hcc-evidence-file)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · one GPU for the model (hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- V28 map and hierarchies, RADV record checks,...: decosa-api HCC review (decosa_api/verticals/hcc), importing the grounding judge (vertical 17) and the typed-judgment engine (vertical 24), CPU
- Evidence sentences with MEAT tags, the typed...: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/hcc-evidence-file-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh

Mac prompt for your coding agent

# Decosa HCC evidence file and RADV defence: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa HCC evidence file and RADV defence on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/hcc-evidence-file.zip (3 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py hcc-evidence-file` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the 2021 heart attack coded as acute is not supported (delete)", "breast cancer treated in 2016 is not supported (delete)", "morbid obesity is held: its only evidence is an audio-only call"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| V28 map and hierarchies, RADV record checks, verdict rules, the no-add guard, net effect, evidence file and signed record (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured |
| Evidence sentences with MEAT tags, the typed judgment per record, and the grounding judge | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py hcc-evidence-file`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Built from the same parts as the measured tools (Qwen3.8-27B and CPU code); not run on the Mac yet.
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 24 s · ~$0.003 per run · 15 receipts

Loading the nightly status…

Self-host: verified 26 Sep 2026 · fresh clone, compose up, sample against local model servers

Measured cost to run: about $0.52 per 100 members (hosted, 26 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_HCC_SYNTHETIC_ONLY=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The rehearsal bundle passed 15 of 15 checks in 4.8 s, the smoke module passed in 2.0 s with 9 attested receipts, a member not marked synthetic streamed a full review, and the tampered record failed to verify. Model-server startup itself not re-verified.

Known limits (5)
  • Synthetic only on the hosted demo, and every number here comes from our own synthetic members; not run on real charts or against coder labels.
  • This page takes typed notes: dates, signatures and credentials sent as fields and text. Scanned charts are read only through the API (POST /hcc/read-chart with an API key) or self-hosted; the page has no upload for them yet.
  • V28 for payment year 2026 (2025 dates of service) only; no V24 or blended years.
  • 'Not supported' means not supported by the notes sent; a coder may find another record.
  • It does not compute risk scores or payment amounts, and it does not run the RADV coversheet, attestation or member-identity checks.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for Qwen3.8-27B; the V28 map, the record checks and the signed record run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the hcc evidence file and radv defence API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.
The open stack

Each submitted risk-adjustment code checked against the notes: keep, hold or delete, with the MEAT quotes, and never a suggestion to add one.

For Medicare Advantage risk-adjustment and compliance teams, and provider groups in risk contracts. Send a member's submitted diagnosis codes and the notes for the service year. Code maps each ICD-10-CM code to its CMS-HCC V28 payment HCC (CMS's own 2026 files) and checks every note and addendum the way a RADV reviewer would: date in the year, an acceptable source, face to face, signed with an acceptable credential, and addenda only from the treating provider within 90 days. Then, per code, Qwen3.8-27B picks the sentences about the condition and tags them monitor, evaluate, assess or treat (MEAT); a typed judgment decides whether the record supports the code as coded; and the grounding judge checks that the quoted sentences carry it on their own. The verdict is supported (keep), insufficient (hold: the evidence is only in a record RADV would not accept, or the model was unsure) or not supported (delete: do not submit). Deletions come first and count as much as confirmations. It reports the net effect in codes and paying HCCs, writes a Markdown evidence file for the audit folder, and signs a hash-chained record with no note text. It checks only the codes you submitted: there is no mode that looks for codes to add or drafts provider queries, and the build fails if a suggestion to add a code reaches the output. A review aid: certified coders decide.

Deployment
Self-host first
Regulatory
Checked 26 Sep 2026 against CMS primary sources (links under Tools). CMS said on 21 May 2025 that it will audit all eligible MA contracts for each payment year (about 550, up from about 60) and grow its medical coders from 40 to about 2,000. The PY 2024 RADV Audit Methods and Instructions (28 Aug 2026) define a valid record as a legibly signed record from a face-to-face visit with a provider in the data collection period, accept audio-and-video telehealth as face to face (HPMS memo, 4 May 2022), and list the invalid-record codes this tool uses (INV2 no signature, INV4 no date, INV5 invalid source such as lab only, home health, hospice, superbill or non-face-to-face, INV7 no credential, INV14 outside the period). The Contract-Level RADV Medical Record Reviewer Guidance (10 Jan 2020) excludes diagnostic radiologists and telephone contacts, treats amendments as timely at up to 90 days and only from the treating provider, says a third party may not amend a record or query the provider for additional diagnoses, and lists acceptable specialties (Appendix B, PY 2015; CMS points to its CSSC site for the current list, which we did not check). The ICD-10-CM Official Guidelines FY 2026 (IV.H and IV.J) bar coding uncertain diagnoses in the outpatient setting and conditions that no longer exist. MEAT (monitor, evaluate, assess, treat) is the coding industry's documentation test, not a CMS term. The Kaiser Permanente affiliates' $556M settlement (14 Jan 2026) concerned diagnoses added through addenda; we could not load the DOJ page on 26 Sep 2026, so that detail is from our research notes (unverified here). Clinical notes are PHI under HIPAA: self-host first; the hosted demo takes synthetic members only. The tool never suggests adding codes, and any price will be per member reviewed, never a share of recovered or protected revenue (prices aren't published yet). This is not legal or coding advice and not an audit determination: certified coders decide. Model licence: Apache-2.0 (Qwen3.8-27B). The V28 table is a US government work (public domain).
Architecture
Text description

A member's submitted diagnosis codes and notes go to the HCC review. In code, each code is mapped to its CMS-HCC V28 payment HCC and each note and addendum gets the RADV record checks (date in the year, source, face to face, signature, credential, addendum timing and author). Per code, Qwen3.8-27B picks the sentences about the condition and tags them monitor, evaluate, assess or treat; a typed judgment decides whether a record supports the code as coded; and the grounding judge checks that the quotes carry it. Rules in code give delete, hold or keep, deletions first, with a guard that removes any suggestion to add a code. Outputs: the net effect in codes and paying HCCs, a Markdown evidence file and a signed hash-chained record without note text. On the hosted route, which takes synthetic members only, every model call gets a gateway-signed receipt. Self-hosted, everything stays on your machine.

Architecture

At a glance

Data retention
Nothing stored: the member file lives in memory for the request. The signed record holds hashes, verdicts and receipt ids, never note text; logs carry counts only.
What leaves the box
Hosted: every model call goes through our gateway to the GPU serving Qwen3.8-27B, and only synthetic members are accepted. Self-hosted on the direct route: nothing leaves the box. Scanned charts: the document reader runs on Decosa's hosted service (self-hosted: on yours).
What it will not do
Suggest codes to add, search notes for new codes, or draft provider queries. It checks only the submitted codes, lists deletions first, and a certified coder decides.
Pricing basis
Per member reviewed (or a flat self-host licence). Never a share of recovered, protected or captured revenue.
Input formats
JSON: member id and service year (age and sex optional: they only pick CMS's age and sex edits in the code map), up to 20 submitted codes, up to 12 notes of up to 12,000 characters (60,000 in total) with date, record type, provider, credential, specialty, signature and addenda. Or, through the API only (not on this page), a scanned chart (PDF or image, one record per page, up to 10 pages and 12 MB hosted) through POST /hcc/read-chart, read by the document reader.
Typical run
The clean sample: a handful of model calls and a fraction of a cent at the gateway list price, under a minute on the shared gateway. Each run shows its own measured cost.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • In the hosted demo

    Lite

    map and record checks only, no GPU

    POST /hcc/codes maps codes to V28 payment HCCs after the hierarchies; POST /hcc/records runs the RADV record checks on every note and addendum. No reading of the notes, so no MEAT verdicts.

    Models
    • decosa-api HCC review (decosa_api/verticals/hcc), importing the grounding judge (vertical 17) and the typed-judgment engine (vertical 24)
    Hardware
    Any CPU
    Quality evidence
    • V28 table against CMS's files8,019 codes, 115 HCC labels, 60 hierarchies; unit-tested against CMS's mapping for spot codes and the age edit on C50.911tests/test_hcc.py, 26 Sep 2026
    • Record-check accuracy on real chartsnot measured yetnot measured yet
    Latency
    no model call; milliseconds on CPU (not separately timed)
    Verification
    No proof yetNo model call, so no receipts; the checks are deterministic code.
  • In the hosted demo

    Standard

    one GPU for the model (hosted demo)

    Qwen3.8-27B reads the notes per code (evidence, typed judgment, grounding); the map, record checks, verdict rules and guard are code. This is what the hosted demo runs, on synthetic members only.

    Models
    • decosa-api HCC review (decosa_api/verticals/hcc), importing the grounding judge (vertical 17) and the typed-judgment engine (vertical 24)
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)
    Quality evidence
    • Verdict accuracy, 3 classes (50 held-out synthetic members, 200 codes)189 / 200docs/evals/hcc-evidence-file.md, test split, 26 Sep 2026
    • Supported vs not supported (codes planted as one or the other)96 / 100docs/evals/hcc-evidence-file.md, test split
    • Codes that should be deleted or held but were kept0 / 161docs/evals/hcc-evidence-file.md, test split
    • Supported calls whose quotes include a planted evidence sentence35 / 35docs/evals/hcc-evidence-file.md, test split
    • Add suggestions or unsubmitted codes in the output0 in 50 membersdocs/evals/hcc-evidence-file.md, test split
    Latency
    measured: under a minute per member on the shared gateway; about a minute for the sample; seconds self-hosted on the direct route.
    Verification
    Proof: strongEvery model call is a separate gateway call with a gateway-signed receipt; the signed record lists them all.
  • Best

    scanned charts too (adds the document reader)

    Everything in Standard, plus scanned PDFs and images: the document reader turns each page into a note with its date, record type, provider, credential and signature read as fields tied to boxes, and every MEAT quote is cited to a page and box of the scan.

    Models
    • decosa-api HCC review (decosa_api/verticals/hcc), importing the grounding judge (vertical 17) and the typed-judgment engine (vertical 24)
    • Qwen3.8-27B (NVFP4)
    • Docling 2.130 with the Heron layout model
    • PaddleOCR-VL-1.6 (0.9B)
    Hardware
    1x RTX PRO 6000 96 GB for Qwen3.8, and the document reader on about 6 GB of a GPU (measured 27 Sep: 5.7 GB peak for parser and layout, 1.6 to 2.4 s a page, the GPU shared) or on a Mac Studio (measured: 3.3 s a page, 2.9 GB)
    Quality evidence
    • Verdict accuracy on scanned charts vs the same members as text (50 held-out synthetic members, 200 codes, same run)187 / 200 vs 192 / 200decosa-api docs/evals/document-reader.md, HCC test split, 26 Sep 2026
    • Same verdict, scanned vs text195 / 200docs/evals/document-reader.md, test split
    • Codes that should be deleted or held but were kept, on scans (test / fresh 1 / fresh 2)0 / 1 / 0 of about 160 eachdocs/evals/document-reader.md: fresh 1 found a coder's addendum read into the note; fixed, then fresh 2 (50 new members) had none
    • Record header fields read right from the scans (date, type, provider, credential, signed)181 / 182 eachdocs/evals/document-reader.md, test split
    • Quotes cited to a page and box of the scan315 / 315docs/evals/document-reader.md, test split
    • After both reader fixes, 50 new members: scanned vs text189 / 200 vs 190 / 200docs/evals/document-reader.md, fresh2 split
    Latency
    measured: a scanned member takes several times longer than text (direct route); the scanned sample reads and reviews in under a minute. With the reader on the GPU (one run through the hosted API): read and reviewed in under a minute, the same verdicts.
    Verification
    Proof: partialSelf-host onlyQwen3.8 calls are receipted as in Standard; the page parse is attested by decosa-api's own key until the gateway has model-call receipts (designed, not deployed). The reader service runs next to the hosted API since 27 Sep 2026 (POST /hcc/read-chart with an API key); the console does not offer scanned charts yet.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
V28 map and hierarchies, RADV record checks, verdict rules, the no-add guard, net effect, evidence file and signed record (no model; CPU)decosa-api HCC review (decosa_api/verticals/hcc), importing the grounding judge (vertical 17) and the typed-judgment engine (vertical 24)
0 GBProof: partial
Evidence sentences with MEAT tags, the typed judgment per record, and the grounding judgeQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Scanned charts: finds and orders the regions of each page (layout only)Docling 2.130 with the Heron layout modeldocling-project/docling-layout-heron on Hugging Face (opens in a new tab)
1 GBProof: partialSelf-host only
Scanned charts: reads each region (text, tables as cells); Qwen3.8 re-reads what it is unsure of and reads the header and signature fieldsPaddleOCR-VL-1.6 (0.9B)PaddlePaddle/PaddleOCR-VL-1.6 on Hugging Face (opens in a new tab)
0.9B · 4.4 GBProof: partialSelf-host only

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /hcc/info, /hcc/samples; POST /hcc/review (SSE or JSON), /hcc/codes, /hcc/records; POST /record/verify. Keeps no note text.

  • vLLM (model):8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the hosted demo and the eval ran through the shared gateway; the self-host check ran on the direct route on the same card.

  • 1x RTX 5090 32 GB Fits

    Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; not run for this tool.

  • CPU only Fits

    The lite tier (POST /hcc/codes and /hcc/records: V28 map, hierarchies and record checks) needs no GPU.

Latency per lane

  • one member, 4 codes, hosted gateway route34.3 s

    Measuredmeasured on our server 2026-09-26: mean over the 50 test members, 3 members in parallel, gateway shared with other workloads

  • mixed-file sample, 8 codes, hosted gateway route45.4 s

    Measuredmeasured on our server 2026-09-26: 45.4 s and 70.7 s in two runs under load

  • mixed-file sample, 8 codes, self-hosted direct route4.8 s

    Measuredmeasured on our server 2026-09-26 in the self-host check (fresh clone, compose up): 4.8-5.0 s; typed judgment by log-probabilities

  • codes and record checks only (POST /hcc/codes, /hcc/records)n/a

    Measuredno model call; milliseconds on CPU (not separately timed)

Notes

  • On 50 held-out synthetic members (200 codes) it gave the planted verdict for 189 codes. It kept no code that should have been deleted or held (0 of 161). Its 11 errors all removed or held a code: 9 are one condition, where 'seropositive rheumatoid arthritis' was not read as 'with rheumatoid factor'.
  • Every supported call quoted a planted evidence sentence (35 of 35); 70 of 72 quoted sentences were planted evidence. Quotes are the note's own sentences with character offsets, so they are verbatim by construction.
  • No add suggestion and no unsubmitted code reached the output in any eval run. Tests run every sample with a model that tries to suggest codes and fail the build if one gets through.
  • A code documented only in a record RADV would not accept (an audio-only call, a radiologist's report, the year before, unsigned, or a coder's addendum months later) is held, with the reason, not kept.
  • Everything is measured on synthetic members written by the agent that built the checker. It has not been run on real charts or against certified coders' labels.
  • Scanned charts (POST /hcc/read-chart, the document reader block): on the same 50 held-out synthetic members printed, signed and scanned, 187 of 200 verdicts were right against 192 of 200 for the text version in the same run; 195 of 200 matched the text verdict, and no scan kept a code it should not. A later split found one false keep (a coder's addendum read into the note); after the fix, 50 new members scored 189 of 200 scanned against 190 of 200 as text, with no false keeps. The scans are clean synthetic pages (typed notes, one font, a pen-squiggle signature), so real faxes will read worse.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

hcc-evidence-file/assemble-prompt.md137 lines
# Assemble the Decosa HCC evidence file on this machine

You are setting up a review aid for Medicare Advantage risk-adjustment and compliance teams. For one member and service
year it takes the diagnosis codes the plan submitted and the member's notes, and for each code that maps to a CMS-HCC
V28 payment HCC it says supported (keep), insufficient (hold for a certified coder) or not supported (delete: do not
submit), with the MEAT sentences (monitor, evaluate, assess, treat) quoted from the notes or their absence stated. It
checks each note and addendum the way a RADV reviewer would, reports the net effect, and writes a Markdown evidence file
and a record signed by this box's own key. Work step by step, show me each command before you run anything with `sudo`,
and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/hcc-evidence-file.zip (3 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py hcc-evidence-file` (the api image carries the same bundle under /app/rehearsal/hcc-evidence-file/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py hcc-evidence-file --bundle hcc-evidence-file.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the 2021 heart attack coded as acute is not supported (delete)", "breast cancer treated in 2016 is not supported (delete)", "morbid obesity is held: its only evidence is an audio-only call"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Model: Qwen3.8-27B (Apache-2.0), used to pick the evidence sentences, for the typed judgment per record, and for the
  grounding check. The V28 map (CMS public-domain files), the record checks, the verdict rules and the signed record are
  decosa-api (AGPL-3.0-or-later) and need no GPU of their own.
- Clinical notes are PHI. They stay on this machine. Bind every port to 127.0.0.1. The service keeps no note text:
  nothing is written to disk and logs carry counts only. Keep it that way; do not add request logging.
- It checks only the codes that were submitted. Do not add any feature that searches notes for new codes, suggests
  codes or drafts provider queries: the tests fail if one appears. A certified coder decides every deletion and keep.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; an RTX PRO
   6000 96 GB is what we measured on; an RTX 5090 32 GB should fit but we have not run this tool on one). Driver
   570 or newer. Blackwell cards run NVFP4; on older cards use `Qwen/Qwen3.8-27B-FP8`.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
   from the official Docker and NVIDIA repositories after asking me, then run
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 30 GB free.

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out the newest release tag that
  contains `decosa_api/verticals/hcc/` (`main` until one does), and build `docker/api/Dockerfile`.
- `vllm/vllm-openai:v0.29.0` for the model; weights `nvidia/Qwen3.8-27B-NVFP4` (revision
  `482ca0f3832238542f8f5295dde86b5f22711d80`), or `Qwen/Qwen3.8-27B-FP8` on a card without NVFP4.
- The V28 table ships inside the image (`decosa_api/verticals/hcc/data/v28_2026.json`). To rebuild it from CMS's own
  downloads, see `scripts/build_hcc_map.py`; `GET /hcc/info` shows the source URLs and the zips' SHA-256.

## 3. docker-compose.yml
Write this in `~/decosa/hcc/`:

```yaml
services:
  llm:
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
              "--enable-prefix-caching"]
    ports: ["127.0.0.1:8114:8000"]
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>
    ports: ["127.0.0.1:8445:8445"]
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_HCC_SYNTHETIC_ONLY: "0"
      DECOSA_HCC_MAX_CONCURRENT: "3"
      DECOSA_HCC_WORKERS: "3"
      DECOSA_BUDGET_LLM_TOKENS: "60000"
    volumes: ["decosa-data:/data"]
    depends_on: { llm: { condition: service_healthy } }
    healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/hcc/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
  decosa-data:
```

`DECOSA_HCC_SYNTHETIC_ONLY: "0"` lets this box take real records; the hosted service refuses anything not marked
`"synthetic": true`. The api keeps its state (keys, receipts, this box's signing key) in the named volume `decosa-data`,
not in a host folder: the image runs as an unprivileged user (uid 10001), and a host folder that Docker creates is owned
by root, which stops the api with `PermissionError: [Errno 13] Permission denied: '/data/keys.sqlite'`. Then start
everything: `docker compose up -d`.

On the direct route the typed judgment uses the model's log-probabilities (one call per record, `judgment_method:
"logprobs"` in `/hcc/info`); a record the model is unsure about is held, not kept. Prefix caching matters: every call for
one member sends the same notes first.

On the first start the api creates this box's Ed25519 key in the `decosa-data` volume (`/data/attest/`, mode 0600).
Back it up with `docker compose cp api:/data/attest ./attest-backup` and keep that copy private. Never print it. Every
model call on the direct route gets a receipt signed with that key (status `attested`): an attestation by me, the
operator, not a proof of computation. Never set `DECOSA_LLM_ROUTE=gateway` on this box: that sends PHI to the hosted
Decosa API.

## 4. Smoke test
1. `curl -s localhost:8445/hcc/info | jq '{synthetic_only, method: .model.judgment_method, table: .hcc_model.codes}'`
   shows `synthetic_only: false`, `logprobs` and `8019`.
2. `curl -s -XPOST localhost:8445/hcc/codes -H 'content-type: application/json' -d '{"codes":["E11.22","E11.9","I10"]}' | jq .paying_hccs`
   shows only HCC 37 (E11.9's HCC 38 is dropped by the hierarchy; I10 has no V28 HCC).
3. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"hcc-evidence-file"}' | jq -r .token)`.
4. `curl -s -XPOST localhost:8445/hcc/review -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample":"mixed-file","synthetic":true}' > res.json`.
   Expect `net_effect.codes` = delete 3, hold 2, keep 3 (one of these can shift between runs; read why): I21.4 and
   C50.911 deleted as history, E66.01 and J44.9 held with reason `invalid_record` (an audio-only call and a diagnostic
   radiologist's report), E11.22, N18.32 and I50.22 kept with quotes, and I10 under `not_reviewed`.
5. `jq .record res.json | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @- | jq .ok`
   must print `true`. Change one `code` entry's `action` from `delete` to `keep` and verify again: it must fail.
6. The same request with `"mode": "opportunities"` must return 400: this service has no such mode.
7. If decosa-api's source is at hand, `python scripts/rehearse.py hcc-evidence-file --base-url http://127.0.0.1:8445`
   runs all of this and prints PASS or FAIL per property (15 checks).
8. Time it and tell me what you measure. The mixed file took about 5 s on our RTX PRO 6000 on the direct route, and
   45-70 s through the shared hosted gateway.

## 5. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`, or call `POST /hcc/review` from your
risk-adjustment workflow per member and file `evidence_file_md` and `record` with the member's RADV folder.
`POST /hcc/codes` and `POST /hcc/records` run the V28 map and the record checks alone, with no model call. Contract:
`API_CONTRACT.md`, section "Changes (hcc-evidence-file, 2026-09-26)".

Off by default. Joining serves other people's requests on this GPU; never do it on a box that holds PHI. If I ask for
it, follow the provider guide at `/provide` on the site, and do not enable it without my explicit yes.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

7 laws, rules and guidance pages cited; 5 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Each submitted risk-adjustment code checked against the notes: keep, hold or delete, with the MEAT quotes, and never a suggestion to add one.
Who it's for
Medicare Advantage risk-adjustment and compliance teams, and provider groups in risk contracts, preparing for RADV audits.
Where it runs
Self-host for real records (PHI); the hosted demo takes synthetic members only
Key numbers
  • 189 / 200 Verdict accuracy, 3 classes (test split, n = 200)
  • 96 / 100 Supported vs not supported (test split, n = 100)
  • 0 / 161 Unsupported or held codes called supported (false keep) (test split, n = 161)
  • 23.9 s Median end-to-end run, hosted (QA sweep 2026-09-26)
All results, datasets and caveats
Models
Qwen3.8-27B reads the notes (evidence, typed judgment, grounding); the V28 map, hierarchies and record checks are plain code
Where
Self-host for real records (PHI); the hosted demo takes synthetic members only
Checks
Receipt per model call; every quote is a sentence of the note with offsets; signed hash-chained record with no note text
Output
Signed record or verdict · Structured data
Data
Patient data (PHI)
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

Will it suggest diagnosis codes to add?

No. It checks only the codes you submitted. There is no mode that searches notes for new codes or drafts provider queries, a request asking for one is refused, and the build fails if a suggestion to add a code reaches the output.

What does it check for each code?

Whether a record documents the condition as current, with the sentences that show it was monitored, evaluated, assessed or treated (MEAT) quoted, and whether that record is one RADV accepts: face to face, signed with an acceptable credential, in the service year, not an audio-only call, a radiologist's report or a late addendum.

How accurate is it?

On 50 held-out synthetic members (200 codes) it gave the planted verdict for 189. It kept none of the 161 codes that should have been deleted or held; its errors all removed or held a code. These are synthetic notes written by us, not real charts, and it has not been compared with coders' labels.

Can I send real patient notes to the hosted demo?

No. Clinical notes are protected health information, so it is self-host first and the hosted demo refuses any member not marked synthetic. Self-hosted on the direct route, nothing leaves the box.

Which HCC model does it use?

CMS-HCC V28 for payment year 2026 (2025 dates of service), built from CMS's own 2026 mapping and model software files with their hierarchies. CMS's age and sex edits only change which HCC some codes map to: they are applied when the member file gives age and sex; with no age it assumes 65 and says so on the code, and with no sex it applies no sex edit. It does not otherwise check a diagnosis against the member's age or sex. V24 and blended years are not covered.

Who makes the final call?

A coder does. It is a review aid, not a coding or audit decision. It reports how many codes the review would delete, hold and keep, and any price will be per member reviewed, never a share of recovered revenue (prices aren't published yet).

Can it read scanned charts?

Yes, through the API or self-hosted, not on this page (it takes typed notes). POST /hcc/read-chart, with an API key, runs the document reader: each scanned page becomes a note with its date, record type, provider, credential and signature read as fields, and every quote is cited to a page and box. On synthetic scans its verdicts came close to the same notes sent as text; the Best tier lists the measured figures. Handwritten notes were not tested.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about HCC evidence file and RADV defence

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.