54 · Finance and insurance · Compliance and trust · live
Insurance claims-file conduct pack
Eval results
Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)
- Planted problems found, all checks52 of 54 (precision 0.98, recall 0.96)test splitn = 5448 test files, run once after the prompts were frozen
- False flags1test splitn = 48A shortened flood exclusion called a misrepresentation
- Clean files with any flag0 of 12test splitn = 12Dev 0 of 7, replies 0 of 6, samples 0 of 2
- Held-out phrasing files: found / false / missed26 / 0 / 0held outn = 24Sentences never seen while the prompts were written; denial-reason wordings are not held out
- Timeliness items with the right start, act and deadline122 of 122test splitn = 122Dev: 64 of 65
- Late replies found (targeted replies set)6 of 6, 0 false flagstest splitn = 12
Dataset
Synthetic claim files for a fictional insurer from scripts/claims_cases.py: dev 24 files, test 48 (half with held-out phrasing), a targeted replies set of 12, and 6 demo samples; problems planted on a structured truth.
Caveats
- Synthetic, templated files: real claim files are longer and messier (scanned letters, email chains, several claimants). These numbers do not predict accuracy on a carrier's files.
- The same author wrote the generator, the planted problems and the prompts.
- Small positive counts for some checks: decision (2), review notice (2) and lowball (3) on the test set.
- Denial-reason wordings are the same six in every set, so misquote detection is not held out on wording.
- Rules are a subset (total-loss valuation, subrogation and others are not checked); no OCR in this build.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 26 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 10 s
- Receipts
- 5
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.004
Self-host verification
Verified on 26 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after
The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): the planted file flagged ack, reason:1 (quoting 'organized race or speed contest'), review_notice and misrepresent, 5 attested receipts, record verified, 5.0 s; the clean file had 0 flags; sign-off pointed at the first record; the rehearsal bundle passed 8/8. Model-server startup was not re-run.
Rehearsal bundle: claims-conduct-pack.zip (4 KB, 8 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.
- Measured on 84 synthetic, templated files written by the building agent; not on real claim files or with a claims auditor's labels.
- Three rulepacks only (NAIC model, California, Texas). The NAIC pack is a baseline, not any state's law.
- Business days skip weekends and US federal holidays; deadlines on weekends are not rolled forward, and one or two days late on such a deadline is marked review.
- The reviewer's name on a sign-off is as given; identity is not verified.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Reads the file (dated events, denial reasons, amounts, AI mentions, each with a quote), grounds each denial reason in the policy, and answers the typed conduct questionsQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
How well does it do on synthetic claim files?
- Planted problems flagged, test set: 52 of 54 (1 false flag; the miss came back as review)
- Clean files with any flag: 0 of 27 (test, dev, replies and demo sets together)
- Deadlines right end to end: 122 of 122 (start date, act date and deadline, test set)
- Cost per file: about $0.004 (5 model calls, 9,351 tokens on the demo file, gateway list price)
Source: decosa-api docs/evals/claims-conduct-pack.md, 26 Sep 2026
Lite · one 48 GB card (1)
- planted problems flagged / false alarms: not measured yet
Standard · the hosted demo, one 96 GB card (6)
- held-out test, 48 synthetic files run once: planted problems flagged / false flags: 52/54 / 1decosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route; prompts frozen on a separate 24-file dev set; half the test files use phrasing never seen while writing the prompts (26/26 found there)
- clean files with any flag: 0/12 test, 0/7 dev, 0/6 replies, 0/2 samplesdecosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route
- date accuracy on the test set: events with the right date / timeliness items with the right start, act and deadline: 247/247 / 122/122decosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route
- per check on the test set (found/planted): acknowledgement 13/13, decision 2/2, payment 4/4, denial reason 11/12, review notice 2/2, misrepresentation 10/11 (1 false), low offer 3/3, investigation 7/7; late replies 6/6 on a 12-file targeted setdecosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route
- AI use recorded, and whether a person was involved read right: 31/31decosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route
- real, de-identified claim files reviewed by a claims-quality auditor: not measured yet
Best · DeepSeek-V4-Flash on two more cards (1)
- planted problems flagged / false alarms: not measured yet
Wanted · two large judges from different families (1)
- planted problems flagged / false alarms: not measured yet