HCC evidence file (62): eval
Date: 2026-09-26. Model: Qwen3.8-27B (NVFP4) through the model gateway, typed judgments by the samples method (the
temperature-0 answer plus 2 votes), pre-release server on our server under the usual multi-agent load.
Script: scripts/eval_hcc.py. Results: docs/evals/hcc-evidence-file/{dev,test}.json (every row, with the verdict and reason).
What was measured
Synthetic members only (decosa_api/verticals/hcc/synth.py; invented people, providers and notes, no PHI). Each member has
four submitted codes drawn from 15 conditions that map to CMS-HCC V28 payment HCCs, and two to five notes with filler
(vitals, other stable problems). Every code is planted as one case:
| Case | Planted | Expected |
|---|---|---|
| supported | MEAT sentences in a signed face-to-face visit in 2025 | supported (keep) |
| history | named only as past or resolved ("Remote NSTEMI in 2021 ...") | not supported (delete) |
| uncertain | only as suspected, rule-out or pending | not supported |
| less_specific | present, but without the detail the code states (CKD stage 2 for a stage 3b code) | not supported |
| absent | not in any note | not supported |
| audio_only / radiology / outside_year / unsigned / late_addendum | the MEAT sentences only in a record RADV does not accept | insufficient (hold) |
Splits: dev = seeds 1-12 (12 members, 48 codes), used while writing the prompts and rules. test = seeds 101-150 (50 members, 200 codes), run once after the dev work was frozen; nothing was changed because of it.
Results
| Metric | dev | test |
|---|---|---|
| Verdict accuracy (3 classes) | 46 / 48 | 189 / 200 |
| Supported vs not supported (codes planted as one or the other) | 22 / 22 | 96 / 100 |
| Unsupported or held codes called supported (false keep) | 0 / 42 | 0 / 161 |
| Supported codes called not supported (false delete) | 0 / 6 | 4 / 39 |
| Supported calls whose quotes include a planted evidence sentence | 6 / 6 | 35 / 35 |
| Quoted sentences that are planted evidence | 12 / 13 | 70 / 72 |
| Add suggestions in any output string or the evidence file | 0 (12 members) | 0 (50 members) |
| ICD-10-CM codes in the output that were not submitted | 0 | 0 |
| Net-effect report consistent with the per-code verdicts | 12 / 12 | 50 / 50 |
| Receipts per member (all model calls) | 17.2 | 17.1 |
| Seconds per member (4 codes, 3 members in parallel, shared gateway) | 21.3 | 34.3 |
Test confusion: supported -> supported 35, supported -> not supported 4; not supported -> not supported 61; insufficient -> insufficient 93, insufficient -> not supported 7. Per case (test): uncertain 13/13, history 14/14, absent 22/22, less_specific 12/12, outside_year 18/18, late_addendum 17/18, unsigned 35/38, audio_only 19/21, radiology 4/5, supported 35/39.
The first dev run scored 42 / 48: the typed judgment read "acute ischemic stroke of the left MCA territory" as less specific than I63.9 (unspecified). The "less specific" option now says that a record more specific than the code still supports it. That was the only prompt change made on dev results.
Where it fails
- 9 of the 11 test errors are one condition: M05.79 (rheumatoid arthritis with rheumatoid factor, multiple sites). Our synthetic notes say "seropositive rheumatoid arthritis"; the model reads that as not stating rheumatoid factor and answers "less specific", so the code is deleted when it should be kept (3) or deleted when it should be held (6). A coder would likely accept "seropositive" here. Not fixed (it showed up on test).
- One acute NSTEMI planted as supported ("Admitted with an NSTEMI ... LAD stented; discharged on ...") in an office visit was read as history. ICD-10-CM allows I21 for four weeks after the infarct; the note gives no date, so this is arguable.
- One heart-failure radiology report was called less specific rather than held.
- Every error errs toward removing a code. None keeps an unsupported one.
Safeguards tested (build-failing)
tests/test_hcc.py:
test_no_add_suggestion_reaches_the_outputruns every sample with a model that writes "Consider adding E11.9 and query the provider" and "You could also code I10" into its reasons; the run must remove them, count them, and passguard.assert_cleanover the whole result and evidence file.test_no_add_guard_catches_what_it_should: add, capture, query, opportunity, gap-closure, "could also code", unsubmitted codes; and our own fixed copy ("Delete: do not submit") passes.test_no_add_mode_or_route_exists:mode,suggest,find,opportunities,gaps,suspects,addin a request are refused; no route or function for suggestions exists;/hcc/*has exactly five routes.review.finalcallsguard.assert_cleanbefore any result leaves the server: a violation raises instead of returning.
Caveats
- Synthetic notes written by the same author as the checker, from sentence templates; real notes are longer, messier and copy-forward heavy. There are no public gold MEAT labels; a 200-chart set labelled by certified coders is the next step.
- MIMIC-IV was not used: its credentialed licence does not fit a hosted demo, and the planted cases need known answers.
- The record checks (date, source, signature, credential, addenda) are code over the metadata sent, not read from scanned pages. The acceptable-credential list is our reading of Appendix B of CMS's January 2020 reviewer guidance (PY 2015).
- V28 only, payment year 2026 (2025 dates of service). No V24 or blended years.
- Samples are not eval data: after the eval, the clean sample's rheumatoid arthritis line was written as "rheumatoid factor positive", and the mixed sample's expected verdict for the coder's addendum ("Alzheimer's disease, on donepezil") is "not supported" (a bare listing; the model gives no MEAT for it).
Rehearsal: checkable properties of the sample run (mixed-file)
rehearsal/hcc-evidence-file/expected.json runs the mixed-file member. It must:
- delete I21.4 (a 2021 NSTEMI coded as acute) and C50.911 (breast cancer treated in 2016);
- hold E66.01 and J44.9 with reason
invalid_record(only in an audio-only call and a diagnostic radiologist's report); - keep E11.22 and I50.22 with quotes;
- list I10 as not reviewed (no V28 HCC) and mark the coder's addendum N1-A1 as not an acceptable record;
- refuse the same request with
"mode": "opportunities"(400); - sign a decosa.record.v1 that verifies at
/record/verifyand fails once its status is changed.