Skip to content
decosa
LiveHostedSelf-hostSelf-host first for real data

Review a Medicare sales call

A yes, no or unclear answer to each CMS marketing question with the quote and its time, and every stated benefit checked against the Summary of Benefits.

Held-out test40/40Call questions answered right (held-out set, never used for tuning)
On production7.5 smedian on production (2026-09-26); slower when the service is busy
List price~$0.72 per 100 callsmeasured, at list price

Built on: Speaker diarization, Typed judgment, Grounding, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the call-sheet items and the record need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the medicare sales-call record API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
medicare-call-record

Use the hosted API

# Decosa Medicare sales-call record: use the hosted API

You are wiring Decosa's Medicare sales-call check into this project. It takes the transcript of a recorded Medicare
Advantage or Part D sales call, a call sheet (when the call was placed and who asked for it, the agency's plan counts, the
Scope of Appointment) and the plan's Summary of Benefits as text.
- **On the call**, it gives typed answers (`yes`, `no`, `unclear` or `na`), each with a quote and its time: the TPMO
  disclaimer, a Scope of Appointment read on the call, "free" for a $0 premium, claiming to be Medicare, an Advantage plan
  called a supplement, gifts, pressure, and for an enrollment the notice, a clear yes and where the Summary of Benefits is.
- **In code**, from the call and the call sheet: whether the disclaimer came before benefits were discussed, its numbers,
  the SOA's date and scope, non-health products, cold calls, and the needs topics and pre-enrollment checklist.
- **Benefit claims**: every premium, copay, allowance or giveback the agent states, judged against the Summary of Benefits.

It returns a signed, hash-chained record with its retention dates (audio for 3 years, audio or a complete and accurate
transcript to year 6, 42 CFR 422.2274(g)(2)(ii)). Every model call has a signed receipt. Use only what is listed below;
if you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **Hosted use is for synthetic material only.** Real sales calls carry Medicare numbers and health conditions (PHI) and
  belong on a self-hosted box (see the self-host prompt). Say so wherever this is wired in.
- **Consent.** CMS requires the recording; some states require every party's consent to it. Collect a consent statement
  in the UI and send it with every check; the API refuses a check without it.
- This is QA triage, not a compliance verdict. Never label a call "compliant".

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "medicare-call-record"}` returns
   `{"token", "expires_at", "budget"}`.
   - The demo allows a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz) and a token allowance per session (the `budget` in the session response), enough for a few checks.
   - Over a limit you get HTTP 429 with `Retry-After`. A demo token runs one check at a time (409 otherwise).
3. A check reserves about 1,800 generated tokens plus 170 per typed question, plus 1,560 with plan facts: about 5,000
   (402 otherwise). A real check of a short call generates far fewer.

## Endpoints
- `POST /medicare/check` (token). Body: `{"transcript": "[00:05] AGENT: ...", "call_sheet": {...}, "plan_facts": {"title": "...", "text": "..."}, "consent": {"recorded_lawfully": true, "method": "all_parties_notified_on_call"}, "speakers"?: {"S01": "Agent"}, "transcript_attestation"?: {"complete_and_accurate": true, "by": "QA lead"}, "title"?: "..."}`.
  - Give exactly one of `transcript` (one utterance per line starting with a time like `[01:23]` and a speaker; WebVTT
    and SRT also work), `segments` (`[{start, end, speaker, text}]` in seconds, as a diarizer returns them), or
    `sample_id` (from `/medicare/samples`, which brings its own call sheet, plan facts and consent).
  - `call_sheet`: `{call: {id?, at: "2026-10-20T10:05:00-05:00", direction: "inbound"|"outbound", contact_basis?: "beneficiary_called"|"beneficiary_requested"|"business_reply_card"|"existing_client"|"plan_business"|"none"|"unknown"}, agency?: {name?, tpmo: true, sells_all: false, organizations: 3, plans: 14}, soa: {status: "documented"|"on_this_call"|"none", method?: "written"|"telephonic"|"electronic", signed_at?: "2026-10-18", products: ["ma", "pdp", "medigap", "dvh", "hospital_indemnity", "cancer_critical", "life", "annuity", "other_non_health"]}, plan?: {name?, type?}}`.
    The call time needs a UTC offset: the call's local date sets the retention dates and which disclaimer rule applies.
  - `plan_facts.text`: the plan's Summary of Benefits (up to 20,000 characters). Claims it does not mention come back
    "not in the plan facts", so give the full summary. Without plan facts, claims are listed for review.
  - `consent.method`: `all_parties_notified_on_call`, `written`, `one_party_state` or `other`. `recorded_lawfully` must be `true`.
  - Limits: 800 transcript lines, 80,000 characters, 12 benefit claims, 512 KB of JSON.
  - JSON response by default:
    - `checklist`: `[{id, rule, title, basis: "call"|"call_sheet", answer?, status: "ok"|"flag"|"review"|"na", reason, line?, at?, t_ms?, quote?, evidence?, probability?, receipt_ids?, missing?, products?}]`
    - `claims`: `[{id, claim, verdict: "supported"|"partial"|"unsupported"|"contradicted", status, reason, at?, quote?, evidence: [{span, text}]}]`
    - also `said`, `counts`, `call`, `timing_rule`, `retention: {audio_until, audio_or_transcript_until, retain_until, enrollment_portion_until?}`, `record`, `record_check`, `receipts`, `note`
  - Typed check ids: `recording_notice`, `tpmo_disclaimer`, `soa_on_call`, `free_misuse`, `medicare_endorsed`,
    `supplement_implied`, `inducement`, `pressure`, `enrollment_notice`, `intent_confirmed`, `sb_location`. Items computed
    in code: `tpmo_disclaimer_timing`, `tpmo_disclaimer_numbers`, `unsolicited`, `soa_documented`, `soa_scope`,
    `non_health`, `needs`, `pecl`.
  - With `Accept: text/event-stream` (or `"stream": true`), the events are `ready`; a `receipt` per model call and a
    `check` per answer as it completes; `claims` then a `claim` per verdict; `sheet`; `result`; `budget`; `done`. If the
    model service does not answer at all you get an `error` event and no record (JSON: 503).
- `POST /record/verify` (no token) `{"record": {...}}` → `{ok, summary, checks, first_bad}`.
- `GET /medicare/info`, `GET /medicare/samples` and `GET /attest/signing-key` need no token.
  `POST /medicare/transcribe` is self-host only (503 here).

## Example: check a call and print what needs a person's eyes (Python, `pip install httpx`)
```python
import httpx, json, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {"transcript": open("call-transcript.txt").read(), "call_sheet": json.load(open("call-sheet.json")),
        "plan_facts": {"title": "Summary of Benefits 2027", "text": open("summary-of-benefits.txt").read()},
        "consent": {"recorded_lawfully": True, "method": "all_parties_notified_on_call"}}
r = httpx.post(f"{API}/medicare/check", json=body, headers=H, timeout=300)
r.raise_for_status()
js = r.json()
for x in js["checklist"] + js["claims"]:
    if x["status"] in ("flag", "review"):
        where = f" (at {x['at']}: \"{x.get('quote', '')[:100]}\")" if x.get("at") else ""
        print(f"{x['status'].upper()}: {x.get('claim') or x['title']} {x.get('answer') or x.get('verdict') or ''}{where}")
json.dump(js["record"], open(f"medicare-call-record-{js['retain_until']}.json", "w"))   # keep with the recording
```

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Medicare sales-call record: run it yourself (containers)

You are setting up the Decosa Medicare sales-call record on this machine, so call recordings, call sheets and plan
documents never leave it. It checks each recorded Medicare Advantage or Part D sales call against the CMS marketing rules
(42 CFR 422/423 subpart V, as amended 1 Jun 2026): typed answers with quotes and times, the call-sheet items in code, and
every benefit claim against the plan's Summary of Benefits. It seals a signed record with the retention dates (audio for 3
years, audio or a complete and accurate transcript to year 6). Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/medicare-call-record.zip (4 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py medicare-call-record` (the api image carries the same bundle under /app/rehearsal/medicare-call-record/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py medicare-call-record --bundle medicare-call-record.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the TPMO disclaimer itself was said (only late)", "benefits were discussed before the TPMO disclaimer", ""free premiums" is flagged with a time"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - keep its data on a named volume;
   - bind every port to 127.0.0.1.

   For recordings, also keep the `diarize` service and set `DECOSA_DIARIZE_URL=http://diarize:8092` and
   `DECOSA_MEDICARE_AUDIO=1` on the api.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check; the first start
   downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/medicare/info` lists the rules with their eCFR sections and the retention
   dates. `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"medicare-call-record"}`, then
   `POST /medicare/check {"sample_id": "tb1-showcase"}`. Expect:
   - `tpmo_disclaimer_timing`, `free_misuse` and `non_health` as `flag`;
   - the "$3,000 dental allowance" claim as `flag` (verdict `contradicted`; the plan facts say $1,500);
   - `retention.audio_until` `2029-11-05` and `retain_until` `2032-11-05`;
   - every receipt `attested`.

   Then `POST /record/verify {"record": <record>}` should give `ok: true`.
6. Report back: the public key and key id, the smoke-test results, and how long the check took.

Sales calls carry PHI: keep the model route local and the API on 127.0.0.1. Some states require every party's consent to
the recording. This is QA triage for a person to review, not a compliance verdict and not legal advice. A machine
transcript counts as complete and accurate only after a person has checked it against the audio.

Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds call
recordings or PHI. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVIDIA NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (61.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns

    The standard tier fits (61.6 of 192 GB). The best tier fits too.

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API, the speech model and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with both models loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "asr": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"medicare-call-record"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py medicare-call-record

Download the mock-data bundle (4 KB, 12 checks)expected.json

A role-played sales call from a fictional agency (Prairie Lantern Benefits) for a fictional HMO, with its call sheet and the plan's Summary of Benefits. The agent pitches "free premiums" and a $3,000 dental allowance before the TPMO disclaimer, then a final-expense life policy. The disclaimer timing, the word free, the non-health pitch and the dental claim (the plan facts say $1,500) must be flagged, and the signed record must verify with the 1 Jun 2026 retention dates.

What the rehearsal checks
  • the TPMO disclaimer itself was said (only late)
  • benefits were discussed before the TPMO disclaimer
  • "free premiums" is flagged with a time
  • the free flag points at a time in the call
  • the final-expense life pitch is flagged
  • the $3,000 dental claim is flagged against the plan facts
  • a check without a consent statement is refused
  • the audio is kept three years after the call
  • audio or a complete transcript is kept six years after the call
  • the signed record verifies
  • a record with its flag count changed no longer verifies
  • every model call has a signed receipt

Licence: Synthetic role-play written from 42 CFR 422/423 subpart V: Prairie Lantern Benefits, Northfield Harbor Health Plan, the plan and every person are fictional. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa Medicare sales-call record: run it yourself (containers)

You are setting up the Decosa Medicare sales-call record on this machine, so call recordings, call sheets and plan
documents never leave it. It checks each recorded Medicare Advantage or Part D sales call against the CMS marketing rules
(42 CFR 422/423 subpart V, as amended 1 Jun 2026): typed answers with quotes and times, the call-sheet items in code, and
every benefit claim against the plan's Summary of Benefits. It seals a signed record with the retention dates (audio for 3
years, audio or a complete and accurate transcript to year 6). Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/medicare-call-record.zip (4 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py medicare-call-record` (the api image carries the same bundle under /app/rehearsal/medicare-call-record/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py medicare-call-record --bundle medicare-call-record.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the TPMO disclaimer itself was said (only late)", "benefits were discussed before the TPMO disclaimer", ""free premiums" is flagged with a time"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - keep its data on a named volume;
   - bind every port to 127.0.0.1.

   For recordings, also keep the `diarize` service and set `DECOSA_DIARIZE_URL=http://diarize:8092` and
   `DECOSA_MEDICARE_AUDIO=1` on the api.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check; the first start
   downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/medicare/info` lists the rules with their eCFR sections and the retention
   dates. `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"medicare-call-record"}`, then
   `POST /medicare/check {"sample_id": "tb1-showcase"}`. Expect:
   - `tpmo_disclaimer_timing`, `free_misuse` and `non_health` as `flag`;
   - the "$3,000 dental allowance" claim as `flag` (verdict `contradicted`; the plan facts say $1,500);
   - `retention.audio_until` `2029-11-05` and `retain_until` `2032-11-05`;
   - every receipt `attested`.

   Then `POST /record/verify {"record": <record>}` should give `ok: true`.
6. Report back: the public key and key id, the smoke-test results, and how long the check took.

Sales calls carry PHI: keep the model route local and the API on 127.0.0.1. Some states require every party's consent to
the recording. This is QA triage for a person to review, not a compliance verdict and not legal advice. A machine
transcript counts as complete and accurate only after a person has checked it against the audio.

Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds call
recordings or PHI. If I ask for it later, follow the Provide page instead of improvising.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsMedicare sales-call record on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

Standard · the hosted demo, one 96 GB card: what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Two extractions: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)
  • Recording to a timed, speaker-labelled transc...: MOSS-Transcribe-Diarize 0.9B. ~4 GB, weights 1.8 GB (estimate). MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate.

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Medicare sales-call record, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Medicare sales-call record on my hardware

Fetch https://decosa.ai/prompts/medicare-call-record-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=medicare-call-record)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 32 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Two extractions: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- Recording to a timed, speaker-labelled transc...: MOSS-Transcribe-Diarize 0.9B (OpenMOSS-Team/MOSS-Transcribe-Diarize), 4 GB

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%), MOSS-Transcribe-Diarize 0.9B ~4 GB (13%); about 0 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/medicare-call-record-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 7.5 s · ~$0.006 per run · 15 receipts

Loading the nightly status…

Self-host: verified 26 Sep 2026 · fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

Measured cost to run: about $0.72 per 100 calls (hosted, 26 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The assembly prompt's smoke tests ran against the already-running local Qwen3.8-27B vLLM (127.0.0.1:8114, network_mode host instead of the compose llm service): tb1-showcase flagged the disclaimer timing, "free", the final-expense pitch and the $3,000 dental claim (contradicted), 15 attested receipts, record verified, retention 2029-11-05 / 2032-11-05, 3.8 s; the pasted cold call flagged Medicare, free, unsolicited, no SOA and the dental claim; no consent gave 400.

Known limits (5)
  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.
  • Measured on 17 short synthetic role-plays written by the building agent, with clean TTS audio. Not measured on real sales calls (20-60 minutes, accents, transfers, Spanish) or with an independent reviewer's labels.
  • Claims the plan facts do not mention come back flagged even when true (2-3 on a clean drug-plan call); give the full Summary of Benefits.
  • ASR errors become claim errors: "eyewear" heard as "in-store" was flagged in one audio run.
  • Audio intake (/medicare/transcribe) is self-host only. The hosted demo's audio samples were transcribed on our server and are bundled.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the call-sheet items and the record need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the medicare sales-call record API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.
The open stack

A recorded Medicare Advantage or Part D sales call checked against the CMS marketing rules, each answer quoted with its time, every benefit claim checked against the plan's Summary of Benefits, sealed in a signed record.

For third-party marketing organizations (agencies, FMOs, call centres), independent agents and plan broker-oversight teams. Give it the sales call (audio on your own box, or a transcript), a call sheet (when the call was placed and who asked for it, your plan counts, the Scope of Appointment) and the plan's Summary of Benefits as text. On the call it answers yes, no or unclear, each with a quote and its time: the TPMO disclaimer, a Scope of Appointment read on the call, "free" for a $0 premium, claiming to be Medicare, calling an Advantage plan a supplement, gifts, pressure, and for an enrollment the enrollment notice, a clear yes and where the Summary of Benefits is. In code: whether the disclaimer came before benefits were discussed (the rule since 1 Jun 2026), its numbers, the SOA's date and scope, non-health products, cold calls, and the needs topics and pre-enrollment checklist before an enrollment. Every premium, copay, allowance or giveback the agent states is judged against the Summary of Benefits. Everything is sealed in a signed, hash-chained record with the retention dates. It is QA triage for a compliance reviewer, not a compliance verdict.

Deployment
Self-host first
Regulatory
Checked 26 Sep 2026 against eCFR (title 42 current to 24 Sep 2026) and the Federal Register. Medicare Advantage and Part D communications and marketing, 42 CFR 422.2260-422.2274 and 423.2260-423.2274, as amended by the CY2027 final rule (91 FR 17384, 6 Apr 2026, effective 1 Jun 2026; CMS's tip sheet 12237-P applies the marketing changes to CY2027 marketing from 1 Oct 2026). Encoded: the TPMO disclaimer, now said before any discussion of benefits (was: within the first minute) and without the SHIP reference (422.2267(e)(41)); the Scope of Appointment agreed and recorded before a personal marketing appointment, no 48-hour wait any more, no products beyond its scope, valid 12 months (422.2264(c)(3), 422.2274(b)(3)); no non-health products such as annuities (422.2263(b)(4); life insurance is treated as non-health here, our reading); no unsolicited calls (422.2264(a)); no "free" for a $0 premium or reduced cost sharing, no claim of Medicare endorsement or misleading use of the Medicare name, no implying a plan is a supplement (422.2262(a)(1)); no cash or non-nominal gifts (422.2263(b)); the needs topics before an enrollment (422.2274(c)(12)), the pre-enrollment checklist reviewed and the Summary of Benefits location given on a telephonic enrollment (422.2267(e)(4), (e)(5)); the telephonic SOA elements from the MCMG (16 Mar 2022). Retention: marketing and sales calls recorded in full and kept 6 years, in audio for years 1-3 and audio or a complete and accurate transcript for years 4-6 (422.2274(g)(2)(ii) as amended; 10 years before 1 Jun 2026); the enrollment portion of a telephonic enrollment is the enrollment form and stays under the 10-year record rule (422.504(d); 91 FR 17463-17464). "Pressure tactics" is not a defined term; it is checked as misleading conduct (422.2262). Recording consent is state law (e.g. California Penal Code 632). Sales calls carry PHI: HIPAA applies to plans and their business associates. QA triage, not legal advice.
Architecture
Text description

A Medicare sales-call recording (self-host only) goes through MOSS-Transcribe-Diarize to a timed, speaker-labelled transcript; a typed, WebVTT or SRT transcript can be given instead. With it come the call sheet (when the call was placed and who asked for it, the agency's plan counts, the Scope of Appointment), the plan's Summary of Benefits as text and a recording-consent statement. Qwen3.8-27B makes two extractions (products, the first discussion of benefits, the enrollment steps; the agent's benefit claims) and answers one typed question per check, each with a quote that must be found in the transcript. Each benefit claim is judged against the Summary of Benefits by the grounding judge. Code computes the disclaimer timing and numbers, the SOA's date and scope, non-health products and cold calls. Outputs: the QA checklist, the benefit claims with verdicts, and a signed hash-chained record with the audio, transcript and enrollment retention dates. Hosted calls get gateway-signed receipts; a self-hosted box signs with its own key.

Architecture

At a glance

What it checks
On each recorded sales call: the TPMO disclaimer and whether it came before benefits, a Scope of Appointment read on the call, "free" for a $0 premium, claiming to be Medicare, an Advantage plan called a supplement, gifts, pressure, and for an enrollment the notice, a clear yes, where the Summary of Benefits is, the needs topics and the pre-enrollment checklist. From the call sheet: the SOA's date and scope, non-health products, cold calls, the disclaimer numbers. Every benefit claim against the Summary of Benefits.
What it does not do
It gives no compliance verdict. It checks benefit claims against the plan facts you give, not the plan's filed benefits: a partial Summary of Benefits produces false flags. It does not check the written SOA form, SMIDs, the Star Ratings document, the plan's outbound enrollment verification, agent licensing, compensation or lead-generation data sharing.
Record-keeping dates
The record carries the dates the rule sets since 1 Jun 2026: audio for 3 years, audio or a complete and accurate transcript to year 6 (42 CFR 422.2274(g)(2)(ii)), and 10 years for the enrollment portion of a telephonic enrollment. A machine transcript counts as complete and accurate only after a person checks it; the record takes that attestation.
Consent
Every run needs a statement that the call was recorded with the consent the law requires (some states need every party's). The statement goes into the signed record, and a check looks for the recording notice on the call.
Data retention
Nothing on the server. Transcripts, call sheets, plan facts and records live in memory for the request, and logs carry counts and timings only. You keep the signed record with the recording until its retention dates.
What leaves the box (hosted demo)
The transcript and the plan facts go to Qwen3.8-27B through our gateway, which is Decosa-operated. The call sheet is processed in code and not sent to the model. The gateway's receipts hold hashes, not text. Self-hosted, nothing leaves.
Model calls per call
Two extractions, 8 or 9 typed checks and one grounding call per benefit claim (up to 12): 13 to 19 calls on the eval calls; the smoke sample used 15 calls and 15,120 tokens, about $0.006.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 48 GB card

    The same checks and claim judge on a smaller mixture-of-experts model. Faster and cheaper; accuracy on this task unknown.

    Models
    • MOSS-Transcribe-Diarize 0.9B
    • Gemma 4 26B A4B (instruction-tuned)
    Hardware
    1x L40S or RTX 6000 Ada 48 GB (not measured)
    Quality evidence
    • typed checks correct / claims graded rightnot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
  • In the hosted demo

    Standard

    the hosted demo, one 96 GB card

    Qwen3.8-27B extracts, answers every typed check and judges every benefit claim; MOSS-Transcribe-Diarize turns recordings into transcripts; the call-sheet items are code. Every model call receipted.

    Models
    • Qwen3.8-27B (NVIDIA NVFP4)
    • MOSS-Transcribe-Diarize 0.9B
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • held-out test B, 4 synthetic calls: typed checks / call-sheet items / claims graded right / false alarms on the 2 clean calls40/40 / 31/32 / 10/10 / 0 of 42decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route, run 1 (run 2 the same); prompts frozen on a 3-call dev split; data, labels and prompts written by the building agent
    • test set, 10 calls: typed checks / planted typed answers / call-sheet items / good claims ok / bad claims flagged101/101 / 9/9 / 79/79 / 25/25 / 4/4decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route, run 4, after one code change made on test run 1 (77/79 call-sheet items before it)
    • false alarms on the 4 clean test calls (flag or review)3 of 87 (claims the plan facts do not mention)decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route, run 4
    • on diarized TTS audio, 8 calls: typed checks / call-sheet items / claims graded right / cited time inside the spoken line81/81 / 63/63 / 20/22 / 28/29 (run 2: 20/22 claims, 29/29)decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26 on audio re-voiced 26 Sep 2026 from macOS voices to Decosa house voices (Kokoro-82M) (the first build: 22/22 and 21/22 claims, 29/29), gateway route, asr runs 1 and 2
    Latency
    measured under shared gateway load: seconds per call, about a minute at the slowest; longer on the re-voiced audio runs, a busier gateway
    Verification
    Proof: strong
  • Best

    two 96 GB cards

    A larger model for long calls and many benefit claims. Not measured on this task.

    Models
    • MOSS-Transcribe-Diarize 0.9B
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    Hardware
    2x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • typed checks correct / claims graded rightnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyNot a hosted model: direct route, calls attested by the box's key.
  • Needs more compute

    Wanted: the best setup

    two large judges from different families

    DeepSeek-V4-Flash and GLM-5.3-Flash each read the file, and a finding stands when they agree; disagreements go to the reviewer. Patient calls stay on your own hardware, never on community providers. Not served yet.

    Models
    • MOSS-Transcribe-Diarize 0.9B
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    • GLM-5.3-Flash
    Hardware
    Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash plus 2x 96 GB for GLM-5.3-Flash, or one Mac Studio with 512 GB holding both 4-bit builds (156 + 165 GB, sizes from our Mac; not run together yet). Estimate.
    Quality evidence
    • typed checks correct / claims graded rightnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
After the session
Two extractions (products, first benefit discussion, enrollment steps; the agent's benefit claims), one typed yes/no/unclear check per question, and one grounding judgment per benefit claim against the plan factsQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
Recording to a timed, speaker-labelled transcript (POST /medicare/transcribe, self-host; the demo's audio samples were transcribed with it)MOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab)
0.9BProof: partialSelf-host only
Lite tier: the same extractions, checks and claim judge on a 48 GB cardGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
25.2B (3.8B active)No proof yetSelf-host only
Best tier: a larger model for long calls and hard casesDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active) · 192 GBNo proof yetSelf-host only
Other
Second judge, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    Transcript, call-sheet and plan-facts intake, the checks, the call-sheet items, signing and the HTTP API (/medicare/*). No GPU. Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.

  • decosa-diarize:8092

    Optional, self-host only: MOSS-Transcribe-Diarize for transcripts made from recordings (DECOSA_MEDICARE_AUDIO=1). No published image yet; built from services/diarize.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server; the diarizer runs on the other card there.

  • 1x L40S / RTX 6000 Ada 48 GB

    Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite).

  • CPU only Fits

    Transcript parsing, the call-sheet items, the signed record and verification need no GPU; the call checks and the claim judge need the model.

Latency per lane

  • check of a 7-19 line sales call with plan facts, quiet shared gateway4.5 s

    Measureddecosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route (test B run 2 median 4.4 s; audio runs medians 6.0-6.2 s)

  • the same, busy shared gateway14.7 s

    Measureddecosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route (test run 4 median 14.7 s, slowest 60.8 s; test B run 1 median 31 s)

  • transcribe a 43-118 s recording (self-host)5.0 s

    Measuredmeasured on our server 2026-09-26, MOSS-Transcribe-Diarize on an RTX PRO 6000 (5.2 s for 64 s of audio)

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

medicare-call-record/assemble-prompt.md215 lines
# Assemble the Decosa Medicare sales-call record on this machine

You are setting up a self-hosted QA checker and record keeper for recorded Medicare Advantage and Part D sales calls, for a
third-party marketing organization (agency, FMO, call centre) or a plan's broker-oversight team. It takes the transcript of
a sales call, a call sheet (when the call was placed and who asked for it, the agency's plan counts, the Scope of
Appointment) and the plan's Summary of Benefits as text.
- **On the call**, each answer is yes, no or unclear, with a quote and a time:
  - the TPMO disclaimer (42 CFR 422.2267(e)(41)), and whether it came before benefits were discussed (the rule since 1 Jun 2026);
  - a Scope of Appointment taken on the call;
  - "free" used for a $0 premium, claiming to be from Medicare, calling an Advantage plan a supplement, gifts, pressure;
  - for an enrollment: the enrollment notice, a clear yes, where the Summary of Benefits is, the needs topics and the pre-enrollment checklist.
- **Benefit claims**: every premium, copay, allowance or giveback the agent states is judged against the Summary of Benefits.
- **From the call sheet**, in code: the SOA's date and scope, non-health products pitched, cold calls, the disclaimer numbers.
- It seals everything in a signed, hash-chained record with the retention dates of 422.2274(g)(2)(ii): audio for 3 years,
  audio or a complete and accurate transcript to year 6, and 10 years for the enrollment portion of a telephonic enrollment.
- With the optional diarizer, it turns a recording into a timed, speaker-labelled transcript.

Work step by step. Show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- Sales calls carry Medicare numbers, dates of birth and health conditions (PHI). Keep everything on this machine: the
  model route stays local (`direct`), and nothing goes to a hosted service.
- CMS requires the recording; some states require every party's consent to it (for example California Penal Code 632).
  Each check needs a consent statement, and the statement goes into the record.
- This is QA triage for a person to review. It is not a compliance verdict and not legal advice. A machine transcript is
  not a "complete and accurate transcript" for years 4-6 until a person has checked it.

Repeat these points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/medicare-call-record.zip (4 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py medicare-call-record` (the api image carries the same bundle under /app/rehearsal/medicare-call-record/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py medicare-call-record --bundle medicare-call-record.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the TPMO disclaimer itself was said (only late)", "benefits were discussed before the TPMO disclaimer", ""free premiums" is flagged with a time"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
| `diarize` (optional, recordings only) | built from the decosa-api source (`services/diarize`) | `OpenMOSS-Team/MOSS-Transcribe-Diarize` @ `704aa4a9c304e8520be88901e0d1960158ef5b15`, Apache-2.0 | internal 8092 |

Typed, WebVTT or SRT transcripts need no speech model. The call-sheet items (disclaimer timing and numbers, SOA, scope, cold calls) are computed in code.

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
   - Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main`; on 48 GB also
     set `LLM_MAX_LEN=32768`. Not measured.
   - Under 48 GB: stop and tell me it will not fit.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories (`sudo nvidia-ctk runtime configure --runtime=docker`, then restart Docker).
3. Confirm about 60 GB of free disk.

## 2. Get the images

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file.
1. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`.
2. If a pull fails, build from source once the `decosa-api` source is published. Clone it, then run
   `docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .` in it, and
   `docker compose build llm` from its compose file.
3. If neither works, stop and tell me.

## 3. Write the compose file

Create `~/decosa-mc/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Benefits compliance>"
```

Create `~/decosa-mc/docker-compose.yml` with exactly these services:

```yaml
name: decosa-mc
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"                # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "200000"        # per session; a call needs about 6,000-8,000 generated tokens
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_MEDICARE_AUDIO: "0"             # "1" only with the diarize service
      DECOSA_DIARIZE_URL: ""
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

Use the named volume `decosa-data` as written: a host bind mount owned by root makes the API fail on
`/data/keys.sqlite`.
1. Run `docker compose up -d`.
2. Poll `docker compose ps` until both services are healthy. The LLM takes 5-10 minutes the first time.
3. `curl -s localhost:8445/healthz` should show `"llm": true`. `asr` is false here, and that is fine.

## 4. Optional: recordings (diarize)

Do this only if I want transcripts made from call recordings.
1. Build `services/diarize` from the decosa-api source on a CUDA PyTorch base image. Give it the env
   `DIARIZE_HOST=0.0.0.0 DIARIZE_PORT=8092 DIARIZE_DEVICE=cuda:0`, the `hf-cache` volume, the same GPU, and a health
   check on `GET /health`.
2. Lower `LLM_GPU_UTIL` to 0.80. On the api, set `DECOSA_DIARIZE_URL: http://diarize:8092` and
   `DECOSA_MEDICARE_AUDIO: "1"`.
3. `POST /medicare/transcribe` then takes a 16 kHz mono 16-bit WAV
   (`ffmpeg -i in.m4a -ac 1 -ar 16000 -sample_fmt s16 out.wav`) and returns timed segments with a speech receipt.
4. Pass the segments to `/medicare/check` as `segments`, with `speakers` to name S01 and S02.

The fit beside the LLM is an estimate, not measured.

## 5. Smoke test

```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"medicare-call-record"}' | jq -r .token)
curl -s $API/medicare/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
  -d '{"sample_id":"tb1-showcase"}' > /tmp/mc.json
jq '{counts, retention, receipts: (.receipts|length), statuses: [.receipts[].status] | unique}' /tmp/mc.json
jq -r '.checklist[] | "\(.status)\t\(.answer // "-")\t\(.at // "")\t\(.id)"' /tmp/mc.json
jq -r '.claims[] | "\(.status)\t\(.verdict)\t\(.at // "")\t\(.claim)"' /tmp/mc.json
jq '{record}' /tmp/mc.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```

In this sample a fictional agency pitches a fictional HMO: "free premiums" and a "$3,000 dental allowance" before the
disclaimer, then a final-expense life policy. Pass if:
- `tpmo_disclaimer_timing` is `flag`, `free_misuse` is `flag` with a time, and `non_health` is `flag`;
- the dental claim is `flag` with verdict `contradicted` (the plan facts say $1,500 a year);
- every receipt has `"status": "attested"`;
- the record verifies (`ok: true`); `retention.audio_until` is `2029-11-05` and `retain_until` is `2032-11-05`.

Then a pasted call with your own call sheet, plan facts and consent:

```bash
curl -s $API/medicare/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{
  "transcript": "[00:00] AGENT: Hi, I am calling from Medicare about your benefits. This call is recorded.\n[00:06] BENEFICIARY: Oh, okay.\n[00:08] AGENT: The Example HMO is free, no premium, with a two thousand dollar dental allowance.",
  "call_sheet": {"call": {"at": "2026-11-02T18:00:00-05:00", "direction": "outbound", "contact_basis": "none"},
                 "agency": {"tpmo": true, "organizations": 3, "plans": 14}, "soa": {"status": "none"}},
  "plan_facts": {"title": "Example SB", "text": "Monthly plan premium: $0.\nComprehensive dental allowance of $1,000 per year."},
  "consent": {"recorded_lawfully": true, "method": "all_parties_notified_on_call"}}' | jq -r '.checklist[], .claims[] | "\(.status)\t\(.id)"'
```

Pass if `medicare_endorsed`, `free_misuse`, `unsolicited`, `soa_documented` and the dental claim are `flag`. Without
`consent`, the API answers 400.

## 6. Point the app at the local API

- The base URL is `http://localhost:8445` (web app: `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`). Add other origins to
  `DECOSA_CORS_ORIGINS`.
- `POST /medicare/check` returns JSON, or streams Server-Sent Events with `Accept: text/event-stream`.
  `GET /medicare/info` lists the call-sheet fields, product types and check ids.
- Build the call sheet from your CRM or dialer: the call time with a UTC offset; `inbound` or `outbound` and who asked
  for the call; your organization and product counts for the caller's area; the SOA status, date and product types.
- Paste the plan's Summary of Benefits as `plan_facts.text` (up to 20,000 characters). Claims it does not mention come back
  "not in the plan facts", so give the full summary.
- The server stores nothing. Keep each signed record (JSON) with the recording until its retention dates. Once a person has
  checked the transcript against the audio, send `transcript_attestation: {"complete_and_accurate": true, "by": "..."}` so the
  record says so. Anyone can re-check a record with `POST /record/verify` against the key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1. For other users on the LAN, put a TLS reverse proxy with authentication in front and set
  `DECOSA_TRUSTED_PROXIES`.

## 7. Keep the direct route

`DECOSA_LLM_ROUTE=gateway` would send prompts, which contain the call transcript and PHI, to the hosted Decosa
API. Never use it
for real calls; at most, use it for synthetic training material.

Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

18 laws, rules and guidance pages cited; 17 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A recorded Medicare Advantage or Part D sales call checked against the CMS marketing rules, each answer quoted with its time, every benefit claim checked against the plan's Summary of Benefits, sealed in a signed record.
Who it's for
TPMO and FMO compliance teams, independent Medicare agents, and plan broker-oversight teams.
Where it runs
Self-host for real calls (hosted demo: synthetic role-plays only)
Key numbers
  • 40/40 Per-rule accuracy, test B run 1 (never used to change anything) (held out, n = 40)
  • 101/101 Per-rule accuracy, test run 4 (test split, n = 101)
  • 9/9 Planted typed answers found, test run 4 (test split, n = 9)
  • 7.5 s Median end-to-end run, hosted (QA sweep 2026-09-26)
All results, datasets and caveats
Models
MOSS-Transcribe-Diarize · Qwen3.8-27B
Where
Self-host for real calls (hosted demo: synthetic role-plays only)
Checks
Receipt per model call; every yes backed by a quote found in the transcript; each benefit claim judged against the plan facts with the span cited; signed hash-chained record with retention dates
Output
Signed record or verdict · Structured data
Data
Patient data (PHI) · Personal data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

What does it check on a recorded sales call?

Yes, no or unclear, each with a quote and its time: the TPMO disclaimer and whether it came before benefits were discussed, a Scope of Appointment read on the call, "free" for a $0 premium, claiming to be Medicare, an Advantage plan called a supplement, gifts, pressure, and for an enrollment the notice, a clear yes and where the Summary of Benefits is.

When must the TPMO disclaimer be said?

Since 1 Jun 2026 (the CY2027 rule, 91 FR 17384), before any discussion of benefits; before that, within the first minute of the call. The check picks the rule by the call's date, and for calls between 1 Jun and 30 Sep 2026 it shows both tests.

How are benefit claims checked?

Every premium, copay, allowance or giveback the agent states is judged against the plan's Summary of Benefits that you supply, with the supporting line. A claim the Summary does not mention comes back flagged even when it is true, so give the full Summary.

How long must sales calls be kept?

Since 1 Jun 2026, 6 years: audio for the first 3, then audio or a complete and accurate transcript for years 4 to 6 (42 CFR 422.2274(g)(2)(ii)). The enrollment portion of a phone enrollment stays under the 10-year record rule. The signed record carries these dates.

Does it say a call is compliant?

No. It is QA triage for a compliance reviewer, not a verdict and not legal advice. It does not check the written SOA form, SMIDs, agent licensing or compensation.

How accurate is it?

On 17 synthetic scripted calls: 101 of 101 per-rule answers on the test set and 40 of 40 on a held-out set, every benefit claim graded right on text transcripts, and 3 false alarms in 87 items on the clean test calls. Real sales calls are not measured.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Medicare sales-call record

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.