Eval: device complaint -> MDR reportability triage (64)
Run 26 Sep 2026 on the pre-release server (decosa-api the pre-release branch, 127.0.0.1:8471 on our server) against the hosted
Qwen3.8-27B through the model gateway, every model call receipted. The gateway was shared with other workloads' runs, so
latencies are as measured under load. Prompts were frozen before the test run (prompt hashes in test-results.json).
Files: docs/evals/device-mdr-triage/ (sets, results, the labeller's files, the MAUDE pool); scripts scripts/mdr_cases.py
(sets) and scripts/mdr_eval.py (runner and scoring).
What was measured
A labelled set of device complaints went through POST /mdr/triage (without the narrative, except on 6 dev and 12 test
items). The triage says reportable, not reportable or needs investigation. We score:
- the false "not reportable" rate: gold reportable, triage "not reportable". This is the risky direction (a missed MDR), so it comes first;
- the false "reportable" rate, how often each class lands in "needs investigation", and accuracy on decisive answers;
- clock accuracy: the primary deadline against the gold date;
- narrative sentences removed or flagged by the grounding check.
The data, and who labelled it
- MAUDE (public, CC0). 756 openFDA device adverse event reports received in 2025, 12 product codes (infusion and insulin pumps, ventilators, oximeters, glucose meters, pacemakers and ICDs, spinal, cervical and knee implants, electrodes, syringes), keeping only the event description (3500A B5). Sentences that mention MDR, 803, FDA, MedWatch or "reportable" were dropped so the text does not announce its label, the maker's name was replaced with "[manufacturer]", and half the narratives were put into sentence case. Label: reportable, because the filer reported it. That is the filer's decision, not ours, and filers lean towards reporting: some of these are conservative malfunction reports (a wrong model shipped, a meter memory fault) that an RA specialist could reasonably have closed. The awareness date is the report's "date received by manufacturer" (G4).
- Synthetic (CC0), written and labelled by a separate agent. It worked from written rules only
(
label-rules.md, drawn from 21 CFR 803 and FDA's 2016 guidance) and never saw the prompts or the code: 70 non-reportable complaints (N1 no malfunction and no serious injury 20, N2 non-serious injury 12, N3 low-risk malfunction 16, N4 erroneous or another maker's device 10, N5 user error only 12), 20 hard reportable ones (serious injury mentioned late, plain-words interventions, malfunction on an implant or ventilator, user error with serious injury) and 24 clock cases (an earlier employee-awareness date written in the narrative, FDA 5-day requests, remedial action, user facilities, due dates on weekends and holidays). Every gold date was computed by hand and checked by the labeller's own script (44 of 44 agree). About half are in MAUDE's capitals style. - Split (seeded): dev 34 (14 MAUDE, 12 negatives, 4 positives, 4 clock cases); test 154 (60 MAUDE, 58 negatives, 16 positives, 20 clock cases). The prompts were written on the samples and changed once on dev (below); the test set was run once.
Results
Test set (154 complaints, run once)
| n | reportable | needs investigation | not reportable | |
|---|---|---|---|---|
| Gold reportable, MAUDE | 60 | 43 | 12 | 5 |
| Gold reportable, synthetic positives | 16 | 15 | 1 | 0 |
| Gold reportable, synthetic clock cases | 20 | 19 | 0 | 1 |
| Gold not reportable, synthetic | 58 | 0 | 5 | 53 |
(One MAUDE item was refused with a 400 in the test run: its event date, 1 Jan 1995, was outside the date range. That was
a bug; the range was fixed and the item rerun on its own: reportable. It is counted above;
test-rerun-maude-23067450.json.)
- False "not reportable": 6 of 96 gold-reportable complaints, 6.3% (Wilson 95% interval 2.9% to 13.0%). On MAUDE 5 of 60 (8.3%); on the synthetic reportable cases 1 of 36.
- False "reportable": 0 of 58 (Wilson 95% interval 0% to 6.2%).
- Reportable caught: 77 of 96 (80%). Not reportable cleared: 53 of 58 (91%).
- Needs investigation: 13 of 96 reportable, 5 of 58 not reportable.
- Accuracy on decisive answers: 130 of 136 (95.6%).
- Clock: 96 of 96 deadlines right. The 36 synthetic clock checks are the real test (30-day, 5-work-day FDA request and remedial action, 10-work-day user facility, earlier employee awareness read from the narrative, holidays and weekends): 36 of 36. The 60 MAUDE checks (awareness + 30 days) only check the date plumbing.
- Narratives drafted on 12 reportable test items: 43 sentences, all kept by the grounding check (none removed or flagged).
- Latency: median 11.9 s, 90th percentile 30.7 s per triage without the narrative, 3 at a time on the shared gateway; about 4 model calls. Single runs measured the same evening ranged from 5 s to 71 s with the narrative (11 calls) as the gateway's load changed.
The six misses (read by the building agent)
| Item | Gold | What the triage said | Reading |
|---|---|---|---|
| syn-clk-009 | R3 (labeller) | malfunction yes, not likely to cause serious harm | A real miss. An infusion pump's anti-free-flow clamp failed; less than 5 mL reached the patient. A free-flow failure on an infusion pump is not a remote risk. |
| maude-22478667 | R2 (filer) | no injury, no malfunction | A real miss on the outcome. A planned hip revision for "normal and expected" liner wear: the revision is surgery, which the outcome question should have counted. |
| maude-21365969 | R2 (filer) | injury not serious, malfunction not likely | Meter error code and a headache, no medical attention. The filer reported; closing it is arguable. |
| maude-23072336 | R2 (filer) | injury not serious, no malfunction | Skin blisters under electrodes, treated with antiseptic soap. The filer reported as an injury; arguable. |
| maude-21057257 | R3 (filer) | malfunction, not likely | Glucose meter did not store readings; no adverse event. A conservative filing. |
| maude-21051200 | R3 (filer) | no malfunction | The wrong pump model was shipped; no patient involvement. A conservative filing. |
Two of the six are clear errors; four are filings an RA specialist could reasonably have closed. We count all six.
Dev set (34 complaints)
- First run (v1): false "not reportable" 1 of 22, false "reportable" 1 of 12 (a competitor's catheter, where the model answered "the device may have contributed"). One change was made on dev: the "caused or contributed" and "malfunction" questions now say the device in question is the one named in the product information, so another maker's device answers no. (A death or serious injury is still never cleared by the tool; those land in needs investigation.)
- Second run (v2): false "not reportable" 1 of 22, false "reportable" 0 of 12; decisive accuracy 29 of 30; clock 22 of 22. The prompts were then frozen.
What it does well, and where it fails
- The asymmetry holds. No complaint the labeller called not reportable was called reportable, and "not reportable" only comes with a quote for every closed path. Deaths never come out "not reportable": the 7 deaths (6 MAUDE, 1 synthetic) the triage did not call reportable all went to needs investigation (literature reviews and deaths the narrative does not tie to the device).
- The clock is right whenever it is computed (it is code), including the rule that awareness starts when any employee heard, read from the narrative with its quote.
- Where it fails: judging "likely to cause serious harm if it recurred" for a malfunction without injury (3 of the 6 misses), and counting a planned revision surgery as an intervention. These are judgement calls, and the RA person makes the call; the false "not reportable" rate is why the tool never files or clears anything by itself.
- Needs investigation is common for MAUDE (12 of 60): terse narratives that do not say what happened to the patient. That costs the RA person time, not safety.
Limits
- MAUDE labels are the filers' decisions, which lean towards reporting; the synthetic labels are one labeller's (an agent working from written rules), not an RA specialist's. No real complaint files, which are confidential, were used.
- The narratives' grounding check was measured on 12 test items only, and only its own verdicts (no separate reader).
- The two-year presumption and the trend grouping were checked on the demo trend only (7 synthetic complaints).
- US FDA rules only. Latency depends on the shared gateway's load.
Expected properties of the sample run (the rehearsal bundle checks these)
rep-told-earlieris reportable, the outcome answer is "serious injury" with a quote from the complaint.- Its 30-day clock starts on 2026-09-02 (the day the sales rep was told, read from the narrative), not the 2026-09-14 received date, and is due 2026-10-02.
cpap-lid-cosmeticis not reportable, with no narrative, and the complaint-record check finds the complainant's address missing.- A 5-day FDA-request clock from 2026-09-02 is due 2026-09-10 (Labor Day skipped).
- The signed triage record verifies at POST /record/verify, and fails once the suggestion is changed; the named decision record verifies.
- The trend view groups the occlusion-alarm complaints and flags C-0719 to re-triage under the presumption.