Skip to content
decosa

101 · Healthcare · Compliance and trust · live

Payer audit response

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 29 Sep 2026

  • Weak claims flagged, new blind BCBSM letter (29 Sep)9 / 9held outn = 922 claims; 6 lacked an objective tool (only the client's own ratings such as SUDS). Clean flagged weak 3 / 13, all from 24-hour times written without colons; fixed after (0 / 13 on a rerun, not held out).
  • Weak claims flagged, new blind Optum letter (29 Sep)7 / 9held outn = 9A control with no measure requirement: clean flagged weak 0 / 12.
  • Weak claims flagged, held-out B rerun (29 Sep)23 / 26held outn = 26Clean flagged weak 0 / 35, requirements 423 / 433; the extra miss was weak on an immediate rerun (run-to-run variance).
  • Weak claims flagged, held-out B24 / 26held outn = 263 published policies (Evernorth BH, NC Medicaid telehealth, CMS therapy plan certification), blind-written; both misses: two-signer plans of care
  • Clean claims flagged weak, held-out B0 / 35held outn = 35pack mode
  • Requirements marked as labelled, held-out B423 / 433held outn = 433found or missing per claim per requirement
  • Respond-by date right3 / 3held outn = 3plus 5/5 on dev letters
  • Unsupported cover-letter sentences0 / 26held outn = 26blind Claude Code judge, 6 letters
  • Pasted policy instead of a pack: clean flagged weak22 / 35held outn = 35not reliable; weak 21/26
  • Weak claims flagged, dev (after fixes)38 / 39dev (tuned on)n = 39sets A and Centene; set A was held out for v1, then used to fix mechanisms

Dataset

Blind-written synthetic audits: dev = 5 sets (97 claims) on BCBSM, Centene, CMS 220.3, DME order and NC Medicaid 8C policies plus Optum; held-out B = 3 sets (61 claims) opened only after the engine was frozen.

Caveats

  • Synthetic cases written by Claude agents; real charts are longer and messier.
  • v1 was run once on set A (33/34 weak but 27/48 clean flagged weak); set A then became dev, which is disclosed.
  • Labels are the writers' own; small n per policy.
  • No human auditor or practice manager rated the output yet.
  • Cross-claim checks (overlap, copied notes) came after the held-out design; the final engine re-run on held-out B gave the same weak and false-weak counts and no cross-claim flags.
  • Objective tools: BCBS Michigan does not define the term; reading it as a scored, standardised instrument's result (not a SUDS or 0-10 rating) is ours.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
28 Sep 2026
Latency, this run
n/a
p50 over passed runs
60 s
Receipts
18
Model calls
n/a
Tokens
n/a
Cost per run
$0.018

Self-host verification

Verified on 28 Sep 2026: fresh clone of the branch into a clean directory on our server, api image built from docker/api/Dockerfile, run with a named data volume, direct route to the already-running local Qwen3.8-27B, local signing; torn down after

Rehearsal bundle 10/10 in 25.3 s (4 weak claims listed first, respond by 2026-10-01, record verifies, every receipt attested); smoke ok in 21.5 s with 18/18 attested receipts and a PDF packet. Model-server startup was not re-run.

Rehearsal bundle: payer-audit.zip (7 KB, 11 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted timing: 5 runs of the 12-claim sample on the pre-release server over the shared gateway (47.8-81.2 s; p95 is the slowest of 5). Production is re-measured after the merge.
  • Synthetic cases only, written by other workloads from real published policies; no practice manager or auditor has rated the output.
  • One signature per note: documents signed by two people (a therapist and a certifying physician) are read as one.
  • Text charts only: scanned PDFs are not read yet.
  • A pasted policy is read by the model and is not reliable (see the eval); a reviewed pack is.
  • Checks across claims (overlapping sessions by the same clinician, near-identical notes) were added after the blind cold-user test; they mark claims to check by hand and are not validated on real charts (the 92% wording threshold was set on the demo data).

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Reads the auditor's letter (who, dates, reference, policy named), reads a pasted policy's requirements when no pack is given, and for each claim points at the note's times, signature and addenda and judges each content requirement found or missing with the exact words; drafts the cover letter body and judges each of its sentences (the grounding judge)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Does it flag the weak claims?

  • Weak claims flagged (held-out B): 24 of 26 (3 policies, 61 claims; both misses were plans of care with two signers)
  • Clean claims flagged weak (held-out B): 0 of 35 (no false alarms in pack mode)
  • Requirements found/missing as labelled: 423 of 433 (held-out B, pack mode)
  • Cost per audit of about 20 claims: about $0.03 (held-out B median $0.028, 80 s, gateway list price)
Lite · one 48 GB card (1)
  • recommendation and criteria accuracy: not measured yet
Standard · the hosted demo, one 96 GB card (5)
  • weak claims flagged, held-out set B (pack mode): 24/26; 0/35 clean claims flagged weak; both misses were plans of care with two signers (the engine reads one signature per note)decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route
  • requirements marked as labelled, held-out set B: 423/433 (97.7%); the found words were on the labelled line 355/377decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route
  • respond-by date right: 3/3 held-out letters (and 5/5 dev letters)decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route
  • kept cover-letter sentences a blind judge found unsupported: 0/26 (6 letters); 0 argued, advised or promiseddecosa-api docs/evals/payer-audit.md; blind judge: Claude Code (Opus 5.5) sub-agent that saw only the sources and the sentences
  • pasted policy text instead of a pack (the model reads the requirements): weak 21/26 but 22/35 clean claims flagged weak: not reliable; review the requirements it read, or use a packdecosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route

How we measure · All tools