Skip to content
decosa
LiveHostedSelf-hostSelf-host first for real data

Triage a device complaint for MDR

A triage memo answering the 803.50(a) questions with quotes from the complaint, a suggested call, the MDR deadline worked out, and a draft 3500A narrative.

Held-out test6 of 96 (6.3%)Reportable complaints wrongly called not reportable (held-out test)
On production12 smedian on production (2026-09-26); slower when the service is busy
List price~$0.19 per 100 complaintsmeasured, at list price

Built on: Typed judgment, Grounding, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) for Qwen3.8-27B; the clocks, the file check and the signed records run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the device complaint mdr triage API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
device-mdr-triage

Use the hosted API

# Decosa device complaint MDR triage: use the hosted API

You are wiring Decosa's MDR triage into this project (a complaint-handling system or a quality dashboard). It takes a
medical device complaint and the product's labeling and returns a triage for FDA Medical Device Reporting (21 CFR 803),
with a signed record. Every model call has a signed receipt. Use only what is listed below; if you need something else,
stop and ask me.

The triage has:
- the 803.50(a) questions answered one by one (`outcome`: death, serious_injury, injury_not_serious, no_injury or
  not_stated; `cause`, `malfunction` and `likely`: yes, no or unclear), each with a quote found word for word;
- a suggestion: `reportable`, `not_reportable` or `needs_investigation`, with the reasoning chain;
- the clock (30 calendar days, 5 work days, or 10 work days for user facilities) from the awareness date, with the rule
  and the arithmetic;
- the 820.35(a) complaint-record check and the Form 3500A items still unknown;
- a draft 3500A event description (unless not reportable), every sentence checked against the complaint.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **Hosted use is for public MAUDE narratives and synthetic complaints only.** Complaints can hold patient health
  information: real ones belong on a self-hosted box (see the self-host prompt). Say so wherever this is wired in.
- This is a triage aid. A qualified person decides and signs (`POST /mdr/decide`). Never auto-file, never auto-close a
  complaint on the suggestion, and never label a firm or a file "compliant".

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "device-mdr-triage"}` returns `{"token", "expires_at", "budget"}`.
   - Sessions per IP are limited; over a limit you get HTTP 429 with `Retry-After`.
   - A demo token runs one triage at a time (409 otherwise).
3. A triage needs about 5,000 generated tokens left in the budget before it starts (402 otherwise); it usually uses
   far fewer, and only what it uses is charged.

## Endpoints
- `POST /mdr/triage` (token). Body: `{"complaint": {"narrative", "id"?, "received_date"?, "aware_date"?, "event_date"?, "device"?: {"name", "model", "lot", "serial", "udi", "product_code"}, "complainant"?: {"name", "address", "phone"}, "patient"?: {"identifier", "age", "sex", "weight"}, "correction"?, "reply"?, "investigation"?, "device_returned"?}, "product"?: {"name", "common_name", "labeling", "implant"?, "life_supporting"?, "class"?}, "role"?: "manufacturer"|"importer"|"user_facility", "fda_5day_request"?, "remedial_action"?: {"needed", "aware_date"?}, "new_info_date"?, "presumption"?, "draft_narrative"?, "title"?, "today"?}` or `{"sample_id": "..."}`.
  - `aware_date` is the first day any employee had the information (21 CFR 803.3(b)(2)); without it, `received_date`
    is used, and a date the complaint names for when an employee was told moves the start earlier.
  - Limits: a narrative of 12,000 characters (at least 5 words), labeling of 4,000, 128 KB of JSON. Paste the labeling's
    intended use and performance claims: URLs are not fetched.
  - The JSON response has `decision` (`{decision, headline, why, paths, event_type}`), `judgments`
    (`[{id, question, answer, evidence: [{ref, quote}], reason, probability, note?}]`), `clock`
    (`{applies, primary: {step, deadline, arithmetic, start, start_is, status, days_left, rules: [{cite, url}]}, rows, notes}`),
    `complaint_file`, `form_3500a`, `narrative`, `narrative_sentences`, `failure_mode`, `triage_md`, `record`,
    `record_check`, `receipts`, `steps` and `note`.
  - With `Accept: text/event-stream` (or `"stream": true`), the events are `ready`, a `receipt` per model call, a
    `judgment` per question, `facts`, `decision`, `clock`, `file_check`, a `sentence` per narrative sentence,
    `narrative`, then `result`, `budget` and `done`.
- `POST /mdr/decide` (token, no model) `{"triage_record", "decision": "reportable"|"not_reportable", "decider": {"name", "role", "id"?}, "rationale", "report_type"?}`
  returns the named person's signed decision record. A not-reportable decision against a reportable suggestion needs a
  physician, nurse, risk manager or biomedical engineer as the decider (21 CFR 803.20(c)(2)).
- `POST /mdr/trend` (token, one model call) `{"complaints": [{"id", "device", "failure_mode", "date"?, "outcome"?, "malfunction"?, "triage"?}]}`
  groups complaints by failure mode and flags FDA's two-year presumption and trends to review.
- `POST /mdr/clock` (no token, no model) `{"aware_date" | "received_date", "role"?, "fda_5day_request"?, "remedial_action"?}` returns the clock alone.
- `POST /record/verify` (no token) `{"record": {...}}` returns `{ok, summary, checks, first_bad}`.
- `GET /mdr/info`, `GET /mdr/samples` and `GET /attest/signing-key` need no token.

## Example: triage a complaint and route it (Python, `pip install httpx`)
```python
import httpx, json, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {"complaint": {"narrative": open("complaint.txt").read(), "received_date": "2026-09-14", "device": {"name": "Example pump"}},
        "product": {"name": "Example pump", "labeling": open("ifu-claims.txt").read()}, "role": "manufacturer"}
r = httpx.post(f"{API}/mdr/triage", json=body, headers=H, timeout=600)
r.raise_for_status()
js = r.json()
print(js["decision"]["headline"], "| decide by", js["clock"]["primary"]["deadline"])
for j in js["judgments"]:
    print(f"{j['id']:>12}  {j['answer']:<20} {j['evidence'][0]['quote'] if j['evidence'] else j.get('note', '')}")
open("mdr-triage.md", "w").write(js["triage_md"])
json.dump(js["record"], open("mdr-triage-record.json", "w"))   # keep in the MDR event file; a person decides next
```

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa device complaint MDR triage: run it yourself (containers)

You are setting up the Decosa MDR triage on this machine, so complaint files never leave it. It reads a medical device
complaint and the product's labeling, and returns:
- the 21 CFR 803.50(a) questions answered one by one, each with a quote found word for word in the complaint;
- a suggestion (reportable, not reportable, needs investigation) with the reasoning chain;
- the MDR clock (30 calendar days, 5 work days, 10 work days for user facilities) from the awareness date, with the rule;
- the 820.35(a) complaint-record check and a draft 3500A event description with every sentence checked.

It seals a signed triage record, and signs the named person's decision. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/device-mdr-triage.zip (5 KB, 16 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py device-mdr-triage` (the api image carries the same bundle under /app/rehearsal/device-mdr-triage/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py device-mdr-triage --bundle device-mdr-triage.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "a 5-day report requested by FDA is due 5 work days after 2 Sep 2026, skipping Labor Day (no model)", "the catheter complaint looks reportable", "the outcome is a serious injury, with a quote from the complaint"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - keep its data on a named volume;
   - bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check. The first start
   downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/mdr/info` lists every rule it cites with its citation, link and the date it
   was read. `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my
   records.
5. Smoke test:
   - `POST /mdr/clock {"aware_date":"2026-09-02","fda_5day_request":true}` needs no model and should give `2026-09-10`
     (Labor Day skipped); without `fda_5day_request`, `2026-10-02`.
   - Get a token with `POST /demo/session {"vertical":"device-mdr-triage"}`, then send
     `POST /mdr/triage {"sample_id": "rep-told-earlier"}`. Expect `decision.decision` `reportable`, the outcome
     `serious_injury` with a quote, the clock starting `2026-09-02` and due `2026-10-02`, and every receipt `attested`.
   - `{"sample_id": "cpap-lid-cosmetic"}` should come back `not_reportable` with `narrative: null`.
   - `POST /record/verify {"record": <record>}` should give `ok: true`.
6. Report back: the public key and key id, the smoke-test results, and how long each triage took.

Complaints can hold patient health information. This is a triage aid, not a reportability determination and not legal
or regulatory advice: a qualified person decides and signs, and it files nothing with FDA.

Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds
complaint files. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVIDIA NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns

    The standard tier fits (57.6 of 192 GB). The best tier fits too.

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"device-mdr-triage"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py device-mdr-triage

Download the mock-data bundle (5 KB, 16 checks)expected.json

Two synthetic device complaints and a synthetic trend. The first: a PICC catheter tip separated on removal and the fragment was snared out in interventional radiology; the sales rep was told on 2 Sep 2026 and the complaint unit logged it on 14 Sep. The triage must say reportable (a serious injury: an intervention to prevent permanent damage), quote it, and start the 30-day clock on 2 Sep (due 2 Oct), not 14 Sep. A named person then signs the decision. The second: a hairline crack in a CPAP humidifier lid's cosmetic cover, device working, no injury: not reportable, and no narrative drafted. The trend: seven complaints about one fictional pump, four with an occlusion alarm that did not sound, one of them a serious injury; the malfunction-only ones of that mode are flagged to re-triage under FDA's two-year presumption. Both signed records verify, and fail once changed.

What the rehearsal checks
  • a 5-day report requested by FDA is due 5 work days after 2 Sep 2026, skipping Labor Day (no model)
  • the catheter complaint looks reportable
  • the outcome is a serious injury, with a quote from the complaint
  • the 30-day clock starts on 2 Sep, the day the sales rep was told, not the 14 Sep received date
  • and is due 2 Oct 2026
  • a draft event description is written, attributed to the report
  • the signed triage record verifies
  • a record whose suggestion was changed no longer verifies
  • the named decision is signed and verifies
  • the decision record names the decider
  • the cosmetic crack does not look reportable
  • and no narrative is drafted for it
  • the complaint record check finds the complainant's address missing
  • the occlusion-alarm complaints are grouped together
  • the malfunction-only complaint C-0719 is flagged to re-triage under the presumption
  • every model call has a signed receipt

Licence: Synthetic (CC0): the complaints, devices and companies are invented (decosa_api/verticals/mdr/data/samples.json). Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa device complaint MDR triage: run it yourself (containers)

You are setting up the Decosa MDR triage on this machine, so complaint files never leave it. It reads a medical device
complaint and the product's labeling, and returns:
- the 21 CFR 803.50(a) questions answered one by one, each with a quote found word for word in the complaint;
- a suggestion (reportable, not reportable, needs investigation) with the reasoning chain;
- the MDR clock (30 calendar days, 5 work days, 10 work days for user facilities) from the awareness date, with the rule;
- the 820.35(a) complaint-record check and a draft 3500A event description with every sentence checked.

It seals a signed triage record, and signs the named person's decision. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/device-mdr-triage.zip (5 KB, 16 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py device-mdr-triage` (the api image carries the same bundle under /app/rehearsal/device-mdr-triage/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py device-mdr-triage --bundle device-mdr-triage.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "a 5-day report requested by FDA is due 5 work days after 2 Sep 2026, skipping Labor Day (no model)", "the catheter complaint looks reportable", "the outcome is a serious injury, with a quote from the complaint"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - keep its data on a named volume;
   - bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check. The first start
   downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/mdr/info` lists every rule it cites with its citation, link and the date it
   was read. `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my
   records.
5. Smoke test:
   - `POST /mdr/clock {"aware_date":"2026-09-02","fda_5day_request":true}` needs no model and should give `2026-09-10`
     (Labor Day skipped); without `fda_5day_request`, `2026-10-02`.
   - Get a token with `POST /demo/session {"vertical":"device-mdr-triage"}`, then send
     `POST /mdr/triage {"sample_id": "rep-told-earlier"}`. Expect `decision.decision` `reportable`, the outcome
     `serious_injury` with a quote, the clock starting `2026-09-02` and due `2026-10-02`, and every receipt `attested`.
   - `{"sample_id": "cpap-lid-cosmetic"}` should come back `not_reportable` with `narrative: null`.
   - `POST /record/verify {"record": <record>}` should give `ok: true`.
6. Report back: the public key and key id, the smoke-test results, and how long each triage took.

Complaints can hold patient health information. This is a triage aid, not a reportability determination and not legal
or regulatory advice: a qualified person decides and signs, and it files nothing with FDA.

Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds
complaint files. If I ask for it later, follow the Provide page instead of improvising.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsDevice complaint MDR triage on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

Standard · the hosted demo, one 96 GB card: what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Reads the complaint: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Device complaint MDR triage, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Device complaint MDR triage on my hardware

Fetch https://decosa.ai/prompts/device-mdr-triage-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=device-mdr-triage)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Reads the complaint: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/device-mdr-triage-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 12 s · ~$0.002 per run · 5 receipts

Loading the nightly status…

Self-host: verified 26 Sep 2026 · fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

Measured cost to run: about $0.19 per 100 complaints (hosted, 26 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): clocks 2026-10-02 and 2026-09-10; rep-told-earlier reportable, clock from 2026-09-02, due 2026-10-02, with a narrative in 4.2 s (10 attested receipts); cpap-lid-cosmetic not reportable; meter-reads-high needs investigation; triage and decision records verified; the rehearsal bundle passed 16/16. Model-server startup was not re-run.

Known limits (5)
  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges. p50 is the eval's test median without the narrative; receipts and cost are from the smoke run (5 calls).
  • False not-reportable rate 6 of 96 on the test set: it can clear a reportable complaint, so a qualified person reviews every suggestion.
  • Measured on public MAUDE narratives (filer labels) and synthetic complaints (one labeller's labels); not on real complaint files or with an RA specialist's labels.
  • US FDA rules only (21 CFR 803, 820.35). No event problem codes, eMDR filing, remedial-action decisions or EU MDR vigilance.
  • The trend grouping and the two-year presumption were checked on the 7-complaint demo trend only.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) for Qwen3.8-27B; the clocks, the file check and the signed records run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the device complaint mdr triage API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.
The open stack

A device complaint in; the 21 CFR 803.50 questions answered with quotes, the MDR clock from the right day, and a signed record of who decided.

For regulatory affairs and quality teams at medical device makers, and the services that handle complaints for them. Give it a complaint (narrative, dates, device) and the product's labeling. Qwen3.8-27B answers the 803.50(a) questions one at a time: the worst patient outcome, whether the device may have caused or contributed, whether it malfunctioned, and whether a recurrence would be likely to cause a death or serious injury. Every answer that settles anything needs a quote the code finds word for word, or it counts as unclear. A rule turns the answers into reportable, not reportable or needs investigation; not reportable needs evidence on every path, and a death or serious injury is never cleared by the tool. The clock is code: 30 calendar days, 5 work days (FDA's request or remedial action) or 10 work days for user facilities, starting on the day any employee heard, which the complaint itself may name. It checks the complaint record against 820.35(a), drafts the Form 3500A event description with every sentence grounded in the complaint, seals a signed triage record, and signs the named person's decision. A trend view groups complaints by failure mode and flags FDA's two-year presumption. A triage aid: a qualified person decides.

Deployment
Self-host first
Regulatory
Rules read on 26 Sep 2026 from eCFR (title 21 current to 24 Sep 2026). Manufacturers report deaths, serious injuries and malfunctions likely to cause or contribute to one if they recurred no later than 30 calendar days after they become aware (21 CFR 803.50(a)); awareness is when any employee has the information (803.3(b)(2)); 5-work-day reports when a reportable event needs remedial action to prevent an unreasonable risk of substantial harm, or FDA asks in writing (803.53); work days exclude federal holidays (803.3(y)); supplemental reports within 30 calendar days of new information (803.56); importers 30 calendar days and user facilities 10 work days (803.10). A qualified person (physician, nurse, risk manager or biomedical engineer) may conclude an event is not reportable, and the MDR event file keeps that information and the deliberations (803.20(c)(2), 803.18). The Quality Management System Regulation took effect on 2 Feb 2026: part 820 incorporates ISO 13485:2016 by reference, 820.35(a) lists what a complaint record must hold, and FDA now inspects under compliance programme 7382.850 instead of QSIT (FDA QMSR page, content current 2 Feb 2026). The malfunction factors, the two-year presumption and the start of the remedial-action clock are from FDA's guidance Medical Device Reporting for Manufacturers (8 Nov 2016), which is not binding. Complaints can hold patient health information: run it on your own hardware; the hosted demo takes public MAUDE narratives and synthetic complaints only. A triage aid, not a reportability determination or legal advice; it never says a firm or file is compliant and files nothing with FDA.
Architecture
Text description

A device complaint (narrative, dates, device fields) and the product's labeling go into decosa-api. Qwen3.8-27B reads the dates an employee was told and the device problem, and answers the 21 CFR 803.50(a) questions as typed answers: outcome and malfunction first, then caused or contributed and likely if it recurred. Code keeps an answer only when its quote is found word for word. A rule turns the answers into reportable, not reportable or needs investigation, and code computes the 30-day, 5-work-day or 10-work-day clock from the awareness date and checks the complaint record against 820.35(a). When the triage does not say not reportable, the model drafts the 3500A event description and the grounding judge checks each sentence. Outputs: the suggestion with quotes, the clock, the file check, the draft, a signed triage record, and the named person's signed decision. A trend view groups complaints by failure mode. Hosted calls get gateway-signed receipts; a self-hosted box signs with its own key.

Architecture

At a glance

What it gives you
The 803.50(a) questions answered with quotes, a suggestion (reportable, not reportable, needs investigation) with the reasoning, the clock with its CFR rule and arithmetic, the 820.35(a) complaint-record check and missing 3500A items, a grounded draft event description, a Markdown memo, a signed triage record and a signed decision record naming the decider.
What it does not do
It does not decide or file: no eMDR submission, no event problem codes, no remedial-action or recall decision (21 CFR 806), no EU MDR vigilance. It never says a firm or file is compliant. It reads the text you send; scanned forms need OCR first.
Data retention
Nothing kept on the server. Complaints and records live in memory for the request; logs carry counts and decisions only. You keep the signed records in the MDR event file.
What leaves the box (hosted demo)
The complaint and labeling go to Qwen3.8-27B through the Decosa API, whose receipts hold hashes, not text. Self-hosted, nothing leaves. The hosted demo is for public MAUDE narratives and synthetic complaints only.
Model calls per triage
One facts read and 2 to 4 questions; with the narrative, one draft and one grounding call per sentence: about 3 to 11 calls on the demo complaints.
Typical run cost
A fraction of a cent at the gateway list price for the catheter complaint with the narrative; less for a cosmetic complaint that does not look reportable. Each run shows its own measured cost.
Clocks
30 calendar days, 5 work days (FDA's written request, or remedial action from the day a supervisor knew), 10 work days for user facilities, and 30 days for supplements, from the day any employee had the information. Deadlines are not moved off weekends or holidays.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 48 GB card

    The same pipeline on a smaller mixture-of-experts model. Faster and cheaper; accuracy on this task unknown.

    Models
    • Gemma 4 26B A4B (instruction-tuned)
    Hardware
    1x L40S or RTX 6000 Ada 48 GB (not measured)
    Quality evidence
    • false not-reportable rate and clock accuracynot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
  • In the hosted demo

    Standard

    the hosted demo, one 96 GB card

    Qwen3.8-27B answers the questions and drafts the narrative; the decision rule, the quote checks, the clocks and the file check are plain code. Every model call receipted.

    Models
    • Qwen3.8-27B (NVIDIA NVFP4)
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • test set (154 complaints, run once): false "not reportable"6 of 96 gold-reportable (6.3%); MAUDE 5 of 60, synthetic 1 of 36decosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26, gateway route; prompts frozen after one change on a 34-complaint dev set
    • false "reportable" / not-reportable complaints cleared0 of 58 / 53 of 58 (5 went to needs investigation)decosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26; complaints written and labelled by a separate agent from written rules
    • reportable caught (MAUDE / synthetic)43 of 60 / 34 of 36; the rest went to needs investigation except the 6 abovedecosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26
    • clock deadline right36 of 36 synthetic cases (earlier employee awareness, FDA 5-day requests, remedial action, user facilities, holidays); 60 of 60 MAUDE date checksdecosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26; gold dates hand-computed and script-checked by the labeller
    • narrative sentences kept by the grounding check43 of 43 on 12 reportable test complaints (its own verdicts; no separate reader)decosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26
    • real complaint files labelled by an RA specialistnot measured yet
    Latency
    measured on our server: seconds per triage without the narrative, a few at a time on the shared gateway, up to a minute as the gateway's load changed with the narrative; seconds self-hosted on the direct route
    Verification
    Proof: strongHosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.
  • Best

    DeepSeek-V4-Flash on two more cards

    A larger model for long, messy complaint files.

    Models
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    Hardware
    2x RTX PRO 6000 96 GB
    Quality evidence
    • false not-reportable rate and clock accuracynot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyNot a hosted model: calls are attested by the box's key only.
  • Needs more compute

    Wanted: the best setup

    two large judges from different families

    DeepSeek-V4-Flash and GLM-5.3-Flash each read the file, and a finding stands when they agree; disagreements go to the reviewer. Patient complaints stay on your own hardware, never on community providers. Not served yet.

    Models
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    • GLM-5.3-Flash
    Hardware
    Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash plus 2x 96 GB for GLM-5.3-Flash, or one Mac Studio with 512 GB holding both 4-bit builds (156 + 165 GB, sizes from our Mac; not run together yet). Estimate.
    Quality evidence
    • false not-reportable rate and clock accuracynot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
After the session
Reads the complaint (dates an employee was told, the event date, the device problem), answers the 803.50(a) questions one at a time (outcome, caused or contributed, malfunction, likely if it recurred), drafts the 3500A event description, judges every narrative sentence (the grounding judge), and groups complaints by failure mode for the trend viewQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
Lite tier: the same pipeline on a 48 GB cardGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
25.2B (3.8B active)No proof yetSelf-host only
Best tier: a larger model for long, messy complaint filesDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active) · 192 GBNo proof yetSelf-host only
Other
Second judge, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only

Measured

How often does it clear a complaint that should be reported?

154 complaints run once: 60 public MAUDE event narratives (reported to FDA, so labelled reportable by the filer's decision) and 94 synthetic complaints written and labelled by a separate agent from written rules, without seeing the prompts. Prompts were changed once on a 34-complaint dev set, then frozen.

Reportable complaints called not reportable
6 of 966.3%; 2 clear errors, 4 conservative MAUDE filings
Not-reportable complaints called reportable
0 of 5853 cleared, 5 sent to needs investigation
Clock deadlines right
36 of 36earlier employee awareness, 5-day and 10-work-day clocks, holidays
Cost per triage with the narrative
about $0.00411 model calls, 9,900 tokens on the catheter demo, gateway list price

Where it fails

Judging whether a malfunction with no injury would be likely to cause serious harm if it recurred: it cleared an infusion pump whose anti-free-flow clamp failed with no harm. It also missed that a planned hip revision is surgery. Terse MAUDE narratives often go to needs investigation (12 of 60).

What it does not show

MAUDE labels are the filers' decisions, which lean towards reporting, and the synthetic labels are one labeller's, not an RA specialist's. Real complaint files are longer and messier. Measure it on your own closed complaints before relying on it.

Source: decosa-api docs/evals/device-mdr-triage.md, 26 Sep 2026

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    Intake, the codes, the deadline rules, the quote checks, the recommendation rule, grounding, signing and the HTTP API (/appeal/*). No GPU. Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server, and the self-host sandbox used the same server.

  • 1x L40S / RTX 6000 Ada 48 GB

    Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite).

  • CPU only Fits

    The clocks (POST /mdr/clock), the decision record, the signed records and verification need no GPU; reading the complaint needs the model.

Latency per lane

  • one triage without the narrative, busy shared gateway11.9 s

    Measureddecosa-api docs/evals/device-mdr-triage.md, test median on our server 2026-09-26 (90th percentile 30.7 s), gateway route, 3 at a time

  • one triage with the narrative, self-hosted direct route4.2 s

    Measuredmeasured on our server 2026-09-26, self-host sandbox, rep-told-earlier (10 model calls)

  • clock rules only50 ms

    Estimateestimate: no model call

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

device-mdr-triage/assemble-prompt.md207 lines
# Assemble the Decosa device complaint MDR triage on this machine

You are setting up a self-hosted MDR reportability triage on this Linux machine for a medical device maker's regulatory
affairs and quality team. It reads a complaint (narrative, dates, device) and the product's labeling, and returns:
- the 21 CFR 803.50(a) questions answered one by one (death or serious injury? could the device have caused or
  contributed? a malfunction? likely to cause or contribute to a death or serious injury if it recurred?), each with the
  words from the complaint that decide it, found word for word by code;
- a suggestion: reportable, not reportable, or needs investigation, with the reasoning chain;
- the clock (30 calendar days, 5 work days, or 10 work days for user facilities) from the awareness date, with the rule
  and the arithmetic;
- the complaint-record check against 21 CFR 820.35(a) and the Form 3500A items still unknown;
- a draft 3500A event description with every sentence checked against the complaint;
- a signed triage record, a signed decision record for the named person who decides, and a failure-mode trend view.

Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- Complaints can hold patient health information. Keep everything on this machine: the model route stays local
  (`direct`), and nothing goes to a hosted service.
- This is a triage aid, not a reportability determination and not legal or regulatory advice. A qualified person decides
  and signs; it never says a firm or a file is compliant, and it files nothing with FDA.

Repeat both points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/device-mdr-triage.zip (5 KB, 16 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py device-mdr-triage` (the api image carries the same bundle under /app/rehearsal/device-mdr-triage/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py device-mdr-triage --bundle device-mdr-triage.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "a 5-day report requested by FDA is due 5 work days after 2 Sep 2026, skipping Labor Day (no model)", "the catheter complaint looks reportable", "the outcome is a serious injury, with a quote from the complaint"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
   - Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main`. On a 48 GB card, also
     set `LLM_MAX_LEN=32768`. Not measured.
   - Under 48 GB: stop and tell me it will not fit.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the
   NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories. Then run
   `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 60 GB of free disk.

## 2. Get the images

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`.

If a pull fails, build from source once the `decosa-api` source is published:
- clone it;
- in the clone, run `docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`;
- run `docker compose build llm` from its compose file.

If neither works, stop and tell me.

## 3. Write the compose file

Create `~/decosa-mdr/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Medical Devices RA/QA>"
```

Create `~/decosa-mdr/docker-compose.yml` with exactly these services:

```yaml
name: decosa-mdr
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"                # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "200000"        # per session; one triage needs up to about 5,000 generated tokens
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_MDR_MAX_CONCURRENT: "3"
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

Use the named volume `decosa-data` exactly as written. A host bind mount owned by root makes the API fail on
`/data/keys.sqlite`.

Run `docker compose up -d`, then poll `docker compose ps` until both services are healthy. The LLM takes 5-10 minutes
the first time. `curl -s localhost:8445/healthz` should show `"llm": true` (`asr` is false here, which is fine).

## 4. Smoke test

The clock rules need no model:

```bash
API=localhost:8445
curl -s $API/mdr/clock -H 'content-type: application/json' -d '{"aware_date":"2026-09-02"}' | jq '.primary | {step, deadline, arithmetic}'
curl -s $API/mdr/clock -H 'content-type: application/json' -d '{"aware_date":"2026-09-02","fda_5day_request":true}' | jq -r .primary.deadline
```

Pass if the 30-day deadline is `2026-10-02` and the 5-day one is `2026-09-10` (Labor Day, 7 Sep, is not a work day).

Then the triage (demo complaints: public MAUDE narratives and synthetic ones):

```bash
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"device-mdr-triage"}' | jq -r .token)
for s in rep-told-earlier cpap-lid-cosmetic meter-reads-high; do
  curl -s $API/mdr/triage -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
    -d "{\"sample_id\":\"$s\"}" > /tmp/mdr-$s.json
  jq -c '{decision: .decision.decision, start: .clock.primary.start, due: .clock.primary.deadline, narrative: (.narrative != null), receipts: (.receipts|length), statuses: ([.receipts[].status]|unique)}' /tmp/mdr-$s.json
done
jq -r '.judgments[] | "\(.id)\t\(.answer)\t\(.evidence[0].quote // "-")"' /tmp/mdr-rep-told-earlier.json
jq '{record}' /tmp/mdr-rep-told-earlier.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```

Pass if:
- `rep-told-earlier` is `reportable`, the outcome is `serious_injury` with a quote, and the clock starts `2026-09-02` (the
  day the sales rep was told, read from the complaint) and is due `2026-10-02`, with a narrative;
- `cpap-lid-cosmetic` is `not_reportable` with no narrative;
- `meter-reads-high` is `needs_investigation`;
- every receipt has `"status": "attested"`, and the record verifies (`ok: true`).

Then record a decision and check it:

```bash
jq '{triage_record: .record, decision: "reportable", report_type: "30_day", decider: {name: "Test Reviewer", role: "RA specialist"}, rationale: "Snare retrieval of a retained fragment is an intervention to prevent permanent damage."}' /tmp/mdr-rep-told-earlier.json \
  | curl -s $API/mdr/decide -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @- > /tmp/mdr-decision.json
jq '{record}' /tmp/mdr-decision.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```

Pass if the decision record verifies.

## 5. Point your complaint system at the local API

- Base URL: `http://localhost:8445`. For a web app, set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`. Add other origins to
  `DECOSA_CORS_ORIGINS`.
- `POST /mdr/triage` takes `{complaint: {narrative, id?, received_date?, aware_date?, event_date?, device?: {name, model,
  lot, serial, udi, product_code}, complainant?, patient?, correction?, reply?, investigation?}, product?: {name,
  common_name, labeling, implant?, life_supporting?, class?}, role?: manufacturer|importer|user_facility,
  fda_5day_request?, remedial_action?: {needed, aware_date?}, presumption?, draft_narrative?}`. It returns JSON, or
  streams Server-Sent Events with `Accept: text/event-stream`.
- Give `aware_date`: the first day any employee (sales, service, call centre) had the information. The received date is
  used only when it is missing. Paste the labeling's intended use and performance claims: the malfunction question is
  judged against them, and URLs are not fetched.
- `POST /mdr/decide` records the named person's decision; `POST /mdr/trend` groups complaints by failure mode (send the
  `failure_mode` each triage returns). `GET /mdr/info` lists every rule quoted with its citation, link and the date read.
- The server stores nothing. Keep the signed triage and decision records in the MDR event file (21 CFR 803.18). Anyone
  can re-check them with `POST /record/verify` against the key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1. For other users on the LAN, put a TLS reverse proxy with authentication in front and set
  `DECOSA_TRUSTED_PROXIES`.

## 6. Keep the direct route

`DECOSA_LLM_ROUTE=gateway` would send prompts to the hosted Decosa
API, and those prompts contain the complaint text. Never
use it for real complaints. At most, use it for public MAUDE narratives or synthetic training material.

Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

17 laws, rules and guidance pages cited; 17 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A device complaint in; the 21 CFR 803.50 questions answered with quotes, the MDR clock from the right day, and a signed record of who decided.
Who it's for
Regulatory affairs and quality teams at medical device makers, and complaint-handling service providers.
Where it runs
Self-host for real complaints (they can hold PHI); the hosted demo takes public MAUDE narratives and synthetic complaints only
Key numbers
  • 6 of 96 (6.3%) Reportable complaints called not reportable (test split, n = 96)
  • 0 of 58 Not-reportable complaints called reportable (test split, n = 58)
  • 53 of 58 Not-reportable complaints cleared (test split, n = 58)
  • 11.9 s Median end-to-end run, hosted (QA sweep 2026-09-26)
All results, datasets and caveats
Models
Qwen3.8-27B (typed answers, the narrative and its grounding check); the decision rule, the clocks and the file check are plain code
Where
Self-host for real complaints (they can hold PHI); the hosted demo takes public MAUDE narratives and synthetic complaints only
Checks
Receipt per model call; every answer that settles anything has a quote found word for word in the complaint; narrative sentences grounded; signed triage and decision records
Output
Signed record or verdict · Structured data
Data
Patient data (PHI)
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

Does it decide whether a complaint is reportable?

No. It suggests reportable, not reportable or needs investigation, with a quote for every answer, and a named person decides and signs the decision record. A death or serious injury is never cleared by the tool; the regulation leaves that conclusion to a qualified person (21 CFR 803.20(c)(2)).

How often does it miss a reportable complaint?

On a test set run once, it called 6 of 96 reportable complaints not reportable (6.3%). Two were clear errors; four were conservative MAUDE filings an RA specialist could have closed. It called none of 58 non-reportable complaints reportable.

When does the MDR clock start?

On the day any employee had the information, not the day the complaint unit logged it (21 CFR 803.3(b)(2)). If the complaint says a sales rep or field engineer was told earlier, the clock starts then, and the quote is shown. It computes 30 calendar days, 5 work days or 10 work days with the arithmetic.

Can the draft 3500A narrative add facts?

Each sentence is checked against the complaint by the grounding judge and a number guard, and unsupported sentences are removed and listed. It is a draft: the RA specialist edits it and files.

Where does complaint data go?

Self-hosted, nothing leaves your hardware. The hosted demo takes public MAUDE narratives and synthetic complaints only; there, the text goes to Qwen3.8-27B through our gateway, whose receipts hold hashes, not text. Nothing is kept on the server.

Will it tell me my complaint files are compliant?

No. It checks one complaint against the questions of 21 CFR 803.50(a) and the items 820.35(a) asks a complaint record to hold. It never says a firm or a file is compliant, and it files nothing with FDA.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Device complaint MDR triage

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.