Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

HCC evidence file (62): eval

Date: 2026-09-26. Model: Qwen3.8-27B (NVFP4) through the model gateway, typed judgments by the samples method (the temperature-0 answer plus 2 votes), pre-release server on our server under the usual multi-agent load. Script: scripts/eval_hcc.py. Results: docs/evals/hcc-evidence-file/{dev,test}.json (every row, with the verdict and reason).

What was measured

Synthetic members only (decosa_api/verticals/hcc/synth.py; invented people, providers and notes, no PHI). Each member has four submitted codes drawn from 15 conditions that map to CMS-HCC V28 payment HCCs, and two to five notes with filler (vitals, other stable problems). Every code is planted as one case:

Case Planted Expected
supported MEAT sentences in a signed face-to-face visit in 2025 supported (keep)
history named only as past or resolved ("Remote NSTEMI in 2021 ...") not supported (delete)
uncertain only as suspected, rule-out or pending not supported
less_specific present, but without the detail the code states (CKD stage 2 for a stage 3b code) not supported
absent not in any note not supported
audio_only / radiology / outside_year / unsigned / late_addendum the MEAT sentences only in a record RADV does not accept insufficient (hold)

Splits: dev = seeds 1-12 (12 members, 48 codes), used while writing the prompts and rules. test = seeds 101-150 (50 members, 200 codes), run once after the dev work was frozen; nothing was changed because of it.

Results

Metric dev test
Verdict accuracy (3 classes) 46 / 48 189 / 200
Supported vs not supported (codes planted as one or the other) 22 / 22 96 / 100
Unsupported or held codes called supported (false keep) 0 / 42 0 / 161
Supported codes called not supported (false delete) 0 / 6 4 / 39
Supported calls whose quotes include a planted evidence sentence 6 / 6 35 / 35
Quoted sentences that are planted evidence 12 / 13 70 / 72
Add suggestions in any output string or the evidence file 0 (12 members) 0 (50 members)
ICD-10-CM codes in the output that were not submitted 0 0
Net-effect report consistent with the per-code verdicts 12 / 12 50 / 50
Receipts per member (all model calls) 17.2 17.1
Seconds per member (4 codes, 3 members in parallel, shared gateway) 21.3 34.3

Test confusion: supported -> supported 35, supported -> not supported 4; not supported -> not supported 61; insufficient -> insufficient 93, insufficient -> not supported 7. Per case (test): uncertain 13/13, history 14/14, absent 22/22, less_specific 12/12, outside_year 18/18, late_addendum 17/18, unsigned 35/38, audio_only 19/21, radiology 4/5, supported 35/39.

The first dev run scored 42 / 48: the typed judgment read "acute ischemic stroke of the left MCA territory" as less specific than I63.9 (unspecified). The "less specific" option now says that a record more specific than the code still supports it. That was the only prompt change made on dev results.

Where it fails

  • 9 of the 11 test errors are one condition: M05.79 (rheumatoid arthritis with rheumatoid factor, multiple sites). Our synthetic notes say "seropositive rheumatoid arthritis"; the model reads that as not stating rheumatoid factor and answers "less specific", so the code is deleted when it should be kept (3) or deleted when it should be held (6). A coder would likely accept "seropositive" here. Not fixed (it showed up on test).
  • One acute NSTEMI planted as supported ("Admitted with an NSTEMI ... LAD stented; discharged on ...") in an office visit was read as history. ICD-10-CM allows I21 for four weeks after the infarct; the note gives no date, so this is arguable.
  • One heart-failure radiology report was called less specific rather than held.
  • Every error errs toward removing a code. None keeps an unsupported one.

Safeguards tested (build-failing)

tests/test_hcc.py:

  • test_no_add_suggestion_reaches_the_output runs every sample with a model that writes "Consider adding E11.9 and query the provider" and "You could also code I10" into its reasons; the run must remove them, count them, and pass guard.assert_clean over the whole result and evidence file.
  • test_no_add_guard_catches_what_it_should: add, capture, query, opportunity, gap-closure, "could also code", unsubmitted codes; and our own fixed copy ("Delete: do not submit") passes.
  • test_no_add_mode_or_route_exists: mode, suggest, find, opportunities, gaps, suspects, add in a request are refused; no route or function for suggestions exists; /hcc/* has exactly five routes.
  • review.final calls guard.assert_clean before any result leaves the server: a violation raises instead of returning.

Caveats

  • Synthetic notes written by the same author as the checker, from sentence templates; real notes are longer, messier and copy-forward heavy. There are no public gold MEAT labels; a 200-chart set labelled by certified coders is the next step.
  • MIMIC-IV was not used: its credentialed licence does not fit a hosted demo, and the planted cases need known answers.
  • The record checks (date, source, signature, credential, addenda) are code over the metadata sent, not read from scanned pages. The acceptable-credential list is our reading of Appendix B of CMS's January 2020 reviewer guidance (PY 2015).
  • V28 only, payment year 2026 (2025 dates of service). No V24 or blended years.
  • Samples are not eval data: after the eval, the clean sample's rheumatoid arthritis line was written as "rheumatoid factor positive", and the mixed sample's expected verdict for the coder's addendum ("Alzheimer's disease, on donepezil") is "not supported" (a bare listing; the model gives no MEAT for it).

Rehearsal: checkable properties of the sample run (mixed-file)

rehearsal/hcc-evidence-file/expected.json runs the mixed-file member. It must:

  1. delete I21.4 (a 2021 NSTEMI coded as acute) and C50.911 (breast cancer treated in 2016);
  2. hold E66.01 and J44.9 with reason invalid_record (only in an audio-only call and a diagnostic radiologist's report);
  3. keep E11.22 and I50.22 with quotes;
  4. list I10 as not reviewed (no V28 HCC) and mark the coder's addendum N1-A1 as not an acceptable record;
  5. refuse the same request with "mode": "opportunities" (400);
  6. sign a decosa.record.v1 that verifies at /record/verify and fails once its status is changed.