Skip to content
decosa

64 · Healthcare · Compliance and trust · live

Device complaint MDR triage

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)

  • Reportable complaints called not reportable6 of 96 (6.3%)test splitn = 96The risky direction. Wilson 95% interval 2.9% to 13.0%. MAUDE 5 of 60, synthetic 1 of 36; 2 clear errors, 4 conservative filings
  • Not-reportable complaints called reportable0 of 58test splitn = 58Complaints written and labelled by a separate agent from written rules
  • Not-reportable complaints cleared53 of 58test splitn = 58The other 5 went to needs investigation (3 were another maker's device with an injury)
  • MAUDE reported events called reportable43 of 60test splitn = 6012 needs investigation, 5 not reportable; labels are the filers' decisions
  • Synthetic reportable complaints called reportable34 of 36test splitn = 361 needs investigation, 1 not reportable
  • Accuracy on decisive answers130 of 136 (95.6%)test splitn = 136Reportable or not reportable answers only
  • Clock deadline right36 of 36test splitn = 36Synthetic clock and positive cases; gold dates hand-computed and script-checked by the labeller. MAUDE date checks 60 of 60
  • Narrative sentences kept by the grounding check43 of 43test splitn = 4312 reportable test complaints; the check's own verdicts, no separate reader

Dataset

154 test and 34 dev complaints: public openFDA MAUDE event narratives from 2025 (CC0; shortened, maker names replaced; labelled reportable because they were reported) and synthetic complaints about fictional devices written and labelled by a separate agent from rules drawn from 21 CFR 803.

Caveats

  • MAUDE labels are the filers' decisions, which lean towards reporting; some misses are conservative filings an RA specialist could have closed.
  • The synthetic labels are one labeller's (an agent working from written rules), not an RA specialist's; no real complaint files were used.
  • One prompt change was made on the dev set before the test run; the test set was run once. One test item hit a date-range bug, fixed and rerun on its own.
  • The narrative check was measured on 12 complaints with the check's own verdicts only; the trend grouping on the demo trend only.
  • US FDA rules only; latency depends on the shared gateway's load.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
26 Sep 2026
Latency, this run
n/a
p50 over passed runs
12 s
Receipts
5
Model calls
n/a
Tokens
n/a
Cost per run
$0.002

Self-host verification

Verified on 26 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): clocks 2026-10-02 and 2026-09-10; rep-told-earlier reportable, clock from 2026-09-02, due 2026-10-02, with a narrative in 4.2 s (10 attested receipts); cpap-lid-cosmetic not reportable; meter-reads-high needs investigation; triage and decision records verified; the rehearsal bundle passed 16/16. Model-server startup was not re-run.

Rehearsal bundle: device-mdr-triage.zip (5 KB, 16 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges. p50 is the eval's test median without the narrative; receipts and cost are from the smoke run (5 calls).
  • False not-reportable rate 6 of 96 on the test set: it can clear a reportable complaint, so a qualified person reviews every suggestion.
  • Measured on public MAUDE narratives (filer labels) and synthetic complaints (one labeller's labels); not on real complaint files or with an RA specialist's labels.
  • US FDA rules only (21 CFR 803, 820.35). No event problem codes, eMDR filing, remedial-action decisions or EU MDR vigilance.
  • The trend grouping and the two-year presumption were checked on the 7-complaint demo trend only.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Reads the complaint (dates an employee was told, the event date, the device problem), answers the 803.50(a) questions one at a time (outcome, caused or contributed, malfunction, likely if it recurred), drafts the 3500A event description, judges every narrative sentence (the grounding judge), and groups complaints by failure mode for the trend viewQwen3.8-27B (NVIDIA NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

How often does it clear a complaint that should be reported?

  • Reportable complaints called not reportable: 6 of 96 (6.3%; 2 clear errors, 4 conservative MAUDE filings)
  • Not-reportable complaints called reportable: 0 of 58 (53 cleared, 5 sent to needs investigation)
  • Clock deadlines right: 36 of 36 (earlier employee awareness, 5-day and 10-work-day clocks, holidays)
  • Cost per triage with the narrative: about $0.004 (11 model calls, 9,900 tokens on the catheter demo, gateway list price)

Source: decosa-api docs/evals/device-mdr-triage.md, 26 Sep 2026

Lite · one 48 GB card (1)
  • false not-reportable rate and clock accuracy: not measured yet
Standard · the hosted demo, one 96 GB card (6)
  • test set (154 complaints, run once): false "not reportable": 6 of 96 gold-reportable (6.3%); MAUDE 5 of 60, synthetic 1 of 36decosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26, gateway route; prompts frozen after one change on a 34-complaint dev set
  • false "reportable" / not-reportable complaints cleared: 0 of 58 / 53 of 58 (5 went to needs investigation)decosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26; complaints written and labelled by a separate agent from written rules
  • reportable caught (MAUDE / synthetic): 43 of 60 / 34 of 36; the rest went to needs investigation except the 6 abovedecosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26
  • clock deadline right: 36 of 36 synthetic cases (earlier employee awareness, FDA 5-day requests, remedial action, user facilities, holidays); 60 of 60 MAUDE date checksdecosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26; gold dates hand-computed and script-checked by the labeller
  • narrative sentences kept by the grounding check: 43 of 43 on 12 reportable test complaints (its own verdicts; no separate reader)decosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26
  • real complaint files labelled by an RA specialist: not measured yet
Best · DeepSeek-V4-Flash on two more cards (1)
  • false not-reportable rate and clock accuracy: not measured yet
Wanted · two large judges from different families (1)
  • false not-reportable rate and clock accuracy: not measured yet

How we measure · All tools