Skip to content
decosa

62 · Healthcare · Finance and insurance · live

HCC evidence file and RADV defence

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)

  • Verdict accuracy, 3 classes189 / 200test splitn = 20050 synthetic members; supported, not supported, insufficient.
  • Supported vs not supported96 / 100test splitn = 100Codes planted as supported or not supported.
  • Unsupported or held codes called supported (false keep)0 / 161test splitn = 161
  • Supported codes called not supported (false delete)4 / 39test splitn = 393 are 'seropositive rheumatoid arthritis' read as not stating rheumatoid factor.
  • Supported calls whose quotes include a planted evidence sentence35 / 35test splitn = 35
  • Quoted sentences that are planted evidence70 / 72test splitn = 72
  • Add suggestions in any output string or the evidence file0 (50 members)test splitn = 50
  • Verdict accuracy, 3 classes (dev)46 / 48dev (tuned on)n = 4842 / 48 before the one prompt change made on dev.
  • Verdict accuracy, scanned charts (document reader)187 / 200test splitn = 200Same members as text in the same run: 192 / 200. Synthetic scans.
  • Same verdict, scanned vs text195 / 200test splitn = 200All 5 differences deleted or held. A later split found one false keep (fixed; 50 new members after: none).

Dataset

Synthetic members from the vertical's own generator (invented people, providers and notes): 15 conditions that map to V28 payment HCCs, each code planted as supported, history, uncertain, less specific, absent, or evidence only in a record RADV does not accept. Dev 12 members (48 codes), test 50 members (200 codes), run once after the dev work was frozen.

Caveats

  • Everything is synthetic, from sentence templates written by the same agent that wrote the checker and prompts: evidence the mechanisms work, not accuracy on real charts.
  • No certified-coder labels: a 200-chart coder-labelled set is the next step.
  • Scanned charts are read by the document reader first; its scans are clean synthetic pages (typed notes, no handwriting in the body), so real faxes will read worse. Text input is unchanged.
  • 9 of the 11 test errors are one condition's wording ('seropositive' not read as 'with rheumatoid factor'); not fixed, since it showed up on test.
  • V28, payment year 2026 only.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
26 Sep 2026
Latency, this run
n/a
p50 over passed runs
24 s
Receipts
15
Model calls
n/a
Tokens
n/a
Cost per run
$0.003

Self-host verification

Verified on 26 Sep 2026: fresh clone, compose up, sample against local model servers

A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_HCC_SYNTHETIC_ONLY=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The rehearsal bundle passed 15 of 15 checks in 4.8 s, the smoke module passed in 2.0 s with 9 attested receipts, a member not marked synthetic streamed a full review, and the tampered record failed to verify. Model-server startup itself not re-verified.

Rehearsal bundle: hcc-evidence-file.zip (3 KB, 15 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Synthetic only on the hosted demo, and every number here comes from our own synthetic members; not run on real charts or against coder labels.
  • This page takes typed notes: dates, signatures and credentials sent as fields and text. Scanned charts are read only through the API (POST /hcc/read-chart with an API key) or self-hosted; the page has no upload for them yet.
  • V28 for payment year 2026 (2025 dates of service) only; no V24 or blended years.
  • 'Not supported' means not supported by the notes sent; a coder may find another record.
  • It does not compute risk scores or payment amounts, and it does not run the RADV coversheet, attestation or member-identity checks.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • V28 map and hierarchies, RADV record checks, verdict rules, the no-add guard, net effect, evidence file and signed record (no model; CPU)decosa-api HCC review (decosa_api/verticals/hcc), importing the grounding judge (vertical 17) and the typed-judgment engine (vertical 24)AGPL-3.0-or-later
  • Evidence sentences with MEAT tags, the typed judgment per record, and the grounding judgeQwen3.8-27B (NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · map and record checks only, no GPU (2)
  • V28 table against CMS's files: 8,019 codes, 115 HCC labels, 60 hierarchies; unit-tested against CMS's mapping for spot codes and the age edit on C50.911tests/test_hcc.py, 26 Sep 2026
  • Record-check accuracy on real charts: not measured yetnot measured yet
Standard · one GPU for the model (hosted demo) (5)
  • Verdict accuracy, 3 classes (50 held-out synthetic members, 200 codes): 189 / 200docs/evals/hcc-evidence-file.md, test split, 26 Sep 2026
  • Supported vs not supported (codes planted as one or the other): 96 / 100docs/evals/hcc-evidence-file.md, test split
  • Codes that should be deleted or held but were kept: 0 / 161docs/evals/hcc-evidence-file.md, test split
  • Supported calls whose quotes include a planted evidence sentence: 35 / 35docs/evals/hcc-evidence-file.md, test split
  • Add suggestions or unsubmitted codes in the output: 0 in 50 membersdocs/evals/hcc-evidence-file.md, test split
Best · scanned charts too (adds the document reader) (6)
  • Verdict accuracy on scanned charts vs the same members as text (50 held-out synthetic members, 200 codes, same run): 187 / 200 vs 192 / 200decosa-api docs/evals/document-reader.md, HCC test split, 26 Sep 2026
  • Same verdict, scanned vs text: 195 / 200docs/evals/document-reader.md, test split
  • Codes that should be deleted or held but were kept, on scans (test / fresh 1 / fresh 2): 0 / 1 / 0 of about 160 eachdocs/evals/document-reader.md: fresh 1 found a coder's addendum read into the note; fixed, then fresh 2 (50 new members) had none
  • Record header fields read right from the scans (date, type, provider, credential, signed): 181 / 182 eachdocs/evals/document-reader.md, test split
  • Quotes cited to a page and box of the scan: 315 / 315docs/evals/document-reader.md, test split
  • After both reader fixes, 50 new members: scanned vs text: 189 / 200 vs 190 / 200docs/evals/document-reader.md, fresh2 split

How we measure · All tools