{"schema_version":"1","site":"https://decosa.ai","id":"payer-audit","num":"101","name":"Payer audit response","tool_name":"Answer a payer audit","short":"Payer audit","blurb":"For the owner or manager of a small practice, or its billing service, when a payer or its review contractor asks for records or audits paid claims. Give it the auditor's letter and the charts. It works out the respond-by date from the letter in code, reads the claim list, matches each claim to its note, and checks every documentation requirement of the payer's published policy: found, with the chart, page and line, or missing. Weak claims are listed first, as plainly as the supported ones. It drafts a cover letter whose sentences are each checked against the sources, builds an indexed, bookmarked response packet, and signs a record of what was checked. It never suggests changing a record, and it never touches a payer portal. The same check runs as a self-audit before a payer asks.","status":"live","labels":{"industry":["healthcare","compliance-trust"],"job":["review","draft"],"input":["text","files"],"deploy":["selfhost"],"status":"live","output":["text","record"],"data":["phi"],"hardware":"gpu-96","licence":"permissive"},"industries":["healthcare","compliance-trust"],"runs_in":["selfhost"],"part_of":[],"built_from":["grounding","signed-record"],"models":"Qwen3.8-27B","where":"Self-host for real records (hosted demo: synthetic cases only; confidential access can't take patient data yet)","hardware":"1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the deadline, the matching, the packet and the signed record need no GPU","final_artifact":"The respond-by date, a per-claim table of the policy's requirements with cites, weak claims first, a draft cover letter, a bookmarked response packet, a worksheet, a CSV and a signed record.","self_host_first":true,"verification":{"receipt_coverage":"full","summary":"Receipt per model call; every found element's words located in the note by code; times, dates and signatures read in code; every cover-letter sentence grounded; signed hash-chained record","manual_qa":{"hosted":{"date":"2026-09-28","result":"pass","p50_ms":59900,"p95_ms":81200,"runs":5,"receipts_per_run":18,"cost_per_run_usd":0.0176},"selfhost":{"date":"2026-09-28","result":"pass","method":"fresh clone of the branch into a clean directory on our server, api image built from docker/api/Dockerfile, run with a named data volume, direct route to the already-running local Qwen3.8-27B, local signing; torn down after","notes":"Rehearsal bundle 10/10 in 25.3 s (4 weak claims listed first, respond by 2026-10-01, record verifies, every receipt attested); smoke ok in 21.5 s with 18/18 attested receipts and a PDF packet. Model-server startup was not re-run."},"known_limits":["Hosted timing: 5 runs of the 12-claim sample on the pre-release server over the shared gateway (47.8-81.2 s; p95 is the slowest of 5). Production is re-measured after the merge.","Synthetic cases only, written by other workloads from real published policies; no practice manager or auditor has rated the output.","One signature per note: documents signed by two people (a therapist and a certifying physician) are read as one.","Text charts only: scanned PDFs are not read yet.","A pasted policy is read by the model and is not reliable (see the eval); a reviewed pack is.","Checks across claims (overlapping sessions by the same clinician, near-identical notes) were added after the blind cold-user test; they mark claims to check by hand and are not validated on real charts (the 92% wording threshold was set on the demo data)."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Weak claims flagged, new blind BCBSM letter (29 Sep)","value":"9 / 9","unit":null,"n":9,"split":"heldout","note":"22 claims; 6 lacked an objective tool (only the client's own ratings such as SUDS). Clean flagged weak 3 / 13, all from 24-hour times written without colons; fixed after (0 / 13 on a rerun, not held out)."},{"name":"Weak claims flagged, new blind Optum letter (29 Sep)","value":"7 / 9","unit":null,"n":9,"split":"heldout","note":"A control with no measure requirement: clean flagged weak 0 / 12."},{"name":"Weak claims flagged, held-out B rerun (29 Sep)","value":"23 / 26","unit":null,"n":26,"split":"heldout","note":"Clean flagged weak 0 / 35, requirements 423 / 433; the extra miss was weak on an immediate rerun (run-to-run variance)."},{"name":"Weak claims flagged, held-out B","value":"24 / 26","unit":null,"n":26,"split":"heldout","note":"3 published policies (Evernorth BH, NC Medicaid telehealth, CMS therapy plan certification), blind-written; both misses: two-signer plans of care"},{"name":"Clean claims flagged weak, held-out B","value":"0 / 35","unit":null,"n":35,"split":"heldout","note":"pack mode"},{"name":"Requirements marked as labelled, held-out B","value":"423 / 433","unit":null,"n":433,"split":"heldout","note":"found or missing per claim per requirement"},{"name":"Respond-by date right","value":"3 / 3","unit":null,"n":3,"split":"heldout","note":"plus 5/5 on dev letters"},{"name":"Unsupported cover-letter sentences","value":"0 / 26","unit":null,"n":26,"split":"heldout","note":"blind Claude Code judge, 6 letters"},{"name":"Pasted policy instead of a pack: clean flagged weak","value":"22 / 35","unit":null,"n":35,"split":"heldout","note":"not reliable; weak 21/26"},{"name":"Weak claims flagged, dev (after fixes)","value":"38 / 39","unit":null,"n":39,"split":"dev","note":"sets A and Centene; set A was held out for v1, then used to fix mechanisms"}],"dataset":"Blind-written synthetic audits: dev = 5 sets (97 claims) on BCBSM, Centene, CMS 220.3, DME order and NC Medicaid 8C policies plus Optum; held-out B = 3 sets (61 claims) opened only after the engine was frozen.","held_out":true,"caveats":["Synthetic cases written by Claude agents; real charts are longer and messier.","v1 was run once on set A (33/34 weak but 27/48 clean flagged weak); set A then became dev, which is disclosed.","Labels are the writers' own; small n per policy.","No human auditor or practice manager rated the output yet.","Cross-claim checks (overlap, copied notes) came after the held-out design; the final engine re-run on held-out B gave the same weak and false-weak counts and no cross-claim flags.","Objective tools: BCBS Michigan does not define the term; reading it as a scored, standardised instrument's result (not a SUDS or 0-10 rating) is ours."],"date":"2026-09-29","doc_url":null},"stack":{"summary":"For owners and managers of small practices, and the billing services that work for them, when a payer or its review contractor asks for records or audits paid claims. Give it the letter, the charts and the payer's published documentation policy. It works out the respond-by date from the letter in code, reads the claim list, matches each claim to its note, and marks each of the policy's requirements found (with chart, page and line) or missing. Weak claims come first, as plainly as the supported ones. It drafts a cover letter whose sentences are each checked against the sources, builds a bookmarked response packet with an index, lists the values an upload form asks for, and signs a record of what was checked. It never suggests changing a record and never touches a payer portal. The same check runs as a self-audit.","tagline":"An auditor's records request and the charts in; the respond-by date, each claim checked against the payer's policy with page:line cites, weak claims first, and a packet.","deployment":"self-host-first","regulatory_note":"Charts are protected health information under HIPAA: run it on the practice's own hardware (confidential access can't take patient data until a business associate agreement is in place); the hosted demo takes synthetic cases only. Not legal, coding or compliance advice: a person, with their compliance lead or counsel, decides what to send and signs. It never suggests adding to, changing or back-dating documentation (a code guard enforces this) and it flags entries dated after the request instead of counting them. It holds no CPT descriptors or AMA guideline text: the practice supplies its codes, and the check uses the payer's own published policy text, quoted with the source. Payer and auditor portals are not automated (their terms forbid it). In a self-audit, rules on identified overpayments may apply (for Medicare, 42 CFR 401.305, text read on eCFR 28 Sep 2026); ask counsel.","components":[{"id":"llm","role":"Reads the auditor's letter (who, dates, reference, policy named), reads a pasted policy's requirements when no pack is given, and for each claim points at the note's times, signature and addenda and judges each content requirement found or missing with the exact words; drafts the cover letter body and judges each of its sentences (the grounding judge)","name":"Qwen3.8-27B (NVIDIA NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP) + FP8 (attention/GDN), FP8 KV cache, MTP speculative decoding k=3","vram_gb":57,"memory_gb_estimate":null,"engine":"vLLM 0.29.0","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"llm-lite","role":"Lite tier: the same pipeline on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","hf_repo":"google/gemma-4-26B-A4B-it","license":"Apache-2.0 (model card also links the Gemma 4 licence page)","params":"25.2B","quant":"BF16 weights; FP8 at load time (vLLM --quantization fp8) to fit a 48 GB card","vram_gb":null,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 (Gemma4ForConditionalGeneration is in its model registry)","receipt_coverage":"none","in_hosted_demo":false,"tiers":["lite"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · one 48 GB card","summary":"The same pipeline on a smaller mixture-of-experts model. Accuracy on this task not measured.","components":["llm-lite"],"hardware":"1x L40S or RTX 6000 Ada 48 GB (not measured)","quality_evidence":[{"metric":"recommendation and criteria accuracy","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"Direct route: calls are attested by the box's key; no gateway receipts.","hosting":null},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","summary":"Qwen3.8-27B reads the letter and each note; the respond-by date, the matching, times, signatures, dates, the claim status and the packet are code. Every model call receipted.","components":["llm"],"hardware":"1x RTX PRO 6000 Blackwell 96 GB","quality_evidence":[{"metric":"weak claims flagged, held-out set B (pack mode)","value":"24/26; 0/35 clean claims flagged weak; both misses were plans of care with two signers (the engine reads one signature per note)","source":"decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route"},{"metric":"requirements marked as labelled, held-out set B","value":"423/433 (97.7%); the found words were on the labelled line 355/377","source":"decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route"},{"metric":"respond-by date right","value":"3/3 held-out letters (and 5/5 dev letters)","source":"decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route"},{"metric":"kept cover-letter sentences a blind judge found unsupported","value":"0/26 (6 letters); 0 argued, advised or promised","source":"decosa-api docs/evals/payer-audit.md; blind judge: Claude Code (Opus 5.5) sub-agent that saw only the sources and the sentences"},{"metric":"pasted policy text instead of a pack (the model reads the requirements)","value":"weak 21/26 but 22/35 clean claims flagged weak: not reliable; review the requirements it read, or use a pack","source":"decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route"}],"latency_note":"measured over the shared gateway: about a minute for the small sample; a few minutes for a full audit depending on load; seconds per claim","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Gateway route: every model call has a gateway-signed receipt.","hosting":null}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:0.1.0","purpose":"Letter and claim-list reading, the respond-by date, note matching, the checks in code, the cover letter check, the PDF packet, signing and the HTTP API (/payer-audit/*). No GPU. Binds 127.0.0.1 by default."},{"name":"decosa-llm","port":8000,"image":"${DECOSA_REGISTRY}/decosa-llm:0.1.0","purpose":"vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network."}],"tools":[{"name":"decosa grounding (vertical 17)","url":"https://decosa.ai/apps/grounding","license":"AGPL-3.0-or-later (decosa-api)","purpose":"Judges every cover-letter sentence against the request letter, the facts code wrote and the index; imported, not copied."},{"name":"claims dates module (vertical 54)","url":"https://decosa.ai/apps/claims-conduct-pack","license":"AGPL-3.0-or-later (decosa-api)","purpose":"Date parsing and calendar/business-day arithmetic for the respond-by date; imported."},{"name":"decosa record (vertical 07) and POST /record/verify","url":"https://decosa.ai/apps/record","license":"AGPL-3.0-or-later (decosa-api)","purpose":"The hash chain and the Ed25519-signed record (hashes only, no chart text); anyone can re-check it."},{"name":"BCBSM behavioral health medical record documentation requirements (non-ABA), revised September 2025","url":"https://authorizations.bcbsm.com/docs/bh-documentation-rqumts-nonaba.pdf","license":"Commercial payer's published policy; the built-in pack stores the short requirement lines, quoted with the source","purpose":"The demo's policy: the individual-therapy progress-note requirements."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server."},{"tier":"1x L40S / RTX 6000 Ada 48 GB","fits":null,"notes":"Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite)."},{"tier":"CPU only","fits":true,"notes":"The respond-by date (POST /payer-audit/deadline), the claim list, the packet PDF and record verification need no GPU; reading the notes needs the model."}],"latency":[{"lane":"12-claim sample, busy shared gateway","typical_ms":59900,"source":"5 runs on our server 2026-09-28, pre-release branch, gateway route: 58.6, 81.2, 64.3, 59.9, 47.8 s"},{"lane":"12-claim sample, self-hosted direct route","typical_ms":21500,"source":"self-host sandbox on our server 2026-09-28 (smoke run)"},{"lane":"140-claim sample, shared gateway","typical_ms":157900,"source":"recorded run on our server 2026-09-28 (final engine), 145 model calls, $0.1513; an earlier run under heavier load took 273.5 s"},{"lane":"respond-by date only","typical_ms":50,"source":"estimate: no model call"}],"benchmark":{"title":"Does it flag the weak claims?","intro":"Blind-written synthetic audits against real published payer policies: another agent wrote each policy's requirement list, the letters, the charts with planted defects and the labels, without seeing the tool. The engine was frozen before held-out set B was opened, then run once.","rows":[{"label":"Weak claims flagged (held-out B)","value":"24 of 26","detail":"3 policies, 61 claims; both misses were plans of care with two signers"},{"label":"Clean claims flagged weak (held-out B)","value":"0 of 35","detail":"no false alarms in pack mode"},{"label":"Requirements found/missing as labelled","value":"423 of 433","detail":"held-out B, pack mode"},{"label":"Cost per audit of about 20 claims","value":"about $0.03","detail":"held-out B median $0.028, 80 s, gateway list price"}],"points":[{"heading":"Where it fails","text":"A document with two signers (a therapist's plan of care certified by a physician) is read as one signature, so a late or uncredentialed certifying signature was missed twice. Pasting the policy text instead of using a reviewed pack makes the model read the requirements, and that flagged 22 of 35 clean claims: use a pack, or review what it read."},{"heading":"What it will not do","text":"Suggest adding to, changing or back-dating a record; argue the case; judge medical necessity or coding; send anything or touch a portal."},{"heading":"What a practice manager said","text":"A blind test user playing a practice manager put a 140-chart request at about 25 hours by hand (her estimate); the tool took about 4 minutes and $0.15. She would pay $150-300 per audit once she can put her own charts in, and would still call counsel when a lot of money is at stake or the letter mentions fraud."}],"source":null},"notes":[]},"buyer_facts":[{"label":"What it gives you","value":"The respond-by date with its arithmetic; each claim's note and each policy requirement found (chart, page, line and the words) or missing; weak claims first; a draft cover letter; a bookmarked PDF packet with an index; the values an upload form asks for with their sources; a worksheet, a CSV and a signed record."},{"label":"Objective measures","value":"BCBS Michigan asks for objective tools to monitor progress: a scored instrument's result (PHQ-9, GAD-7, PCL-5) counts, the client's own rating (SUDS, a 0-10 score) does not (our reading; BCBSM does not define the term). Optum and Evernorth ask for no instrument in a progress note, and the packs say so with their sources."},{"label":"Checks across claims","value":"Two sessions by the same clinician that overlap on the same day, and a note whose wording is nearly identical to another claim's note, are flagged to check by hand, with the other claim named."},{"label":"What it does not do","value":"It does not send, fax or upload anything, or log in to a portal. It does not judge medical necessity or whether a code was right, know your contract or the auditor's sampling, or read scanned charts (text and pages only). It never suggests changing a record."},{"label":"Data retention","value":"Nothing written to disk. The letter, charts and result live in memory; a finished run is kept one hour so its downloads work, then dropped. Logs carry counts only. The signed record holds hashes, not chart text."},{"label":"What leaves the box (hosted demo)","value":"The letter and notes go to Qwen3.8-27B through Decosa's gateway, whose receipts hold hashes, not text. Self-hosted, nothing leaves. The hosted demo takes synthetic cases only; real records go through self-host (confidential access can't take patient data until a business associate agreement is in place)."},{"label":"Model calls per audit","value":"One letter read, one call per claim, one cover-letter draft and one grounding call per letter sentence: 18 calls for 12 claims, about 145 for 140."},{"label":"Typical run cost","value":"A few cents for the small sample and more for a full audit at the gateway list price; a fraction of a cent per claim on the held-out sets. Each run shows its own measured cost."},{"label":"Codes","value":"You supply the billed codes. The tool holds no code descriptors or code-set guidelines; it checks the chart against the payer's own published policy text, quoted with its source."}],"data_handling":{"page":"/data#payer-audit","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":true,"summary":"Hosted demo on sample or public data only; self-host for real data.","gpus":"operator-contracted","third_parties":[],"retention":"Nothing written to disk. The letter, charts and result live in memory; a finished run is kept one hour so its downloads work, then dropped. Logs carry counts only. The signed record holds hashes, not chart text.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/clinics/payer-audit","input":"runner","lanes":[{"id":"deadline","title":"Respond by","kind":"list"},{"id":"table","title":"Claims x requirements","kind":"list"},{"id":"letter","title":"Draft cover letter","kind":"markdown"},{"id":"packet","title":"Response packet","kind":"list"},{"id":"record","title":"Signed record","kind":"json"}],"samples":[{"n":1,"id":"therapist-140","title":"Therapist 140","deep_link":"/clinics/payer-audit?sample=1&autorun=0"},{"n":2,"id":"therapist-12","title":"Therapist 12","deep_link":"/clinics/payer-audit?sample=2&autorun=0"},{"n":3,"id":"self-audit-60","title":"Self audit 60","deep_link":"/clinics/payer-audit?sample=3&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/payer-audit-hosted.md","selfhost":"/prompts/payer-audit-selfhost.md","assemble":"/prompts/payer-audit-assemble.md","mac":null},"rehearsal":{"bundle":"/samples/payer-audit.zip","bundle_url":"https://decosa.ai/samples/payer-audit.zip","folder":"/samples/payer-audit/","expected":"/samples/payer-audit/expected.json","files":["/samples/payer-audit/expected.json","/samples/payer-audit/inputs/charts.json","/samples/payer-audit/inputs/letter.txt"],"bytes":6952,"checks":["the respond-by date is the letter date + 10 calendar days (no model)","the run gives the same date","12 claims are read from the letter","exactly the four planted claims are weak","four weak claims","and they are listed first","the two sessions that overlap (same clinician, same day) are marked check, not supported","claim 6's times are only in an addendum dated after the request","a draft cover letter names the request reference","the signed record verifies","every model call has a signed receipt"],"licence":"Synthetic (CC0): practice, clinician, clients, IDs, plan and auditor are invented (scripts/payer_audit_cases.py). Policy: BCBSM's published documentation requirements, quoted short with the source. Part of decosa-api.","about":"A synthetic records request from a made-up review contractor, dated 21 Sep 2026, with ten calendar days to respond, for 12 psychotherapy claims of a made-up solo practice, and the practice's progress notes. Checked against BCBSM's published individual-therapy documentation requirements (the built-in pack). Four claims are planted weak: two notes with no start and stop times, one with an empty interventions field, and one whose times appear only in an addendum dated after the request. Two sessions on the same day overlap, which an auditor would ask about. The run must mark exactly the four weak, mark the overlapping pair to check by hand and list them first, cite every found requirement by chart, page and line, give the respond-by date 2026-10-01, and seal a record that verifies.","run":{"containers":"docker compose exec api python scripts/rehearse.py payer-audit","checkout":"python scripts/rehearse.py payer-audit --bundle payer-audit.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py payer-audit"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=payer-audit","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":40,"basis":"estimate","unknown":[]},{"id":"standard","gpu_gb":57.6,"basis":"stack","unknown":[]}],"mac":null},"links":{"page":"/clinics/payer-audit","json":"/use-cases/payer-audit.json","metrics":"/metrics/payer-audit","console":"/clinics/payer-audit","console_sample":"/clinics/payer-audit?sample=1&autorun=0","stack":"/clinics/payer-audit#stack","try_live":"/clinics/payer-audit","watch":"/clinics/payer-audit","build":"/clinics/payer-audit#build","self_host":"/clinics/payer-audit#self-host","prompts":{"hosted":"/prompts/payer-audit-hosted.md","selfhost":"/prompts/payer-audit-selfhost.md","assemble":"/prompts/payer-audit-assemble.md","mac":null}}}