Eval: insurance claims-file conduct pack (54)
Run 26 Sep 2026 on the pre-release server (127.0.0.1:8471) against the hosted Qwen3.8-27B through the model gateway, with every call receipted. The gateway was shared with other workloads' work, so the latencies were measured under load.
What was measured
The input is a synthetic claim file and its policy. Everything comes from scripts/claims_cases.py:
- a fictional insurer (Harbor Oak Mutual) and invented people;
- two invented policy forms (personal auto and homeowners);
- per file, a first notice of loss, an adjuster's activity log, and letters (acknowledgement, status, decision or denial, and payment).
Each file comes from a structured truth: the dates, the outcome, the reason, the amounts and any AI use. Problems are
planted on top of that truth. The generator also writes down which checklist items should come back flag.
The model reads the documents, not the truth. The date format (5 styles), the log layout (one line, or a date line then the text) and the letter date style vary by seed.
The sets:
| Set | Files | Planted flags | Clean files | Use |
|---|---|---|---|---|
| dev (seed 1) | 24 | 25 | 7 | all prompt and rule work happened on this set |
| test (seed 2) | 48 | 54 | 12 | run once, after the prompts were frozen (commit 4381b0a) |
| replies (seed 3) | 12 | 6 | 6 | targeted: late replies to the claimant (no positives survived in the random sets) |
| samples | 6 | 8 | 2 | the demo files shown on the site |
Held-out phrasing. Half of the test files (the odd-numbered ones, "B") use sentences that never appeared while the prompts were written, for example "Field appraiser looked at the car ... 28 photos uploaded", "Check for $2,490.00 cut and mailed", and "After reviewing claim X, we are unable to pay it". The other half ("A") use the dev set's templates with new dates, amounts and combinations. The denial reasons (the policy statements in the letters) are the same six wordings in every set, so misquote detection on the test set is not held out on wording.
Changes made after looking at dev results (all before the test run):
- The investigation question now says that the claimant's own estimate is not an investigation.
- Quotes stitched from several lines are accepted when each piece is found.
- A "partial" grounding verdict is flagged only when the misrepresentation judgment points at the same sentence.
- The misrepresentation item also flags when a denial reason's policy statement is graded unsupported or contradicted.
- A claimant message that disputes a denial counts as an objection.
- The extraction prompt now says how to tell "model" from "model then human".
Two generator bugs were fixed as well: a partial claim's staff estimate now covers only the accepted part, and a file planted with no investigation no longer gets a photo-estimate model (a photo estimate is itself an investigation step).
Results on the test set (48 files, run once)
Detection of planted problems. A positive is an item that comes back flag. Every other file is a negative for that
check.
| Check | Planted | Found | False flags | Missed | Precision | Recall | Recall counting "review" |
|---|---|---|---|---|---|---|---|
| Acknowledgement late or missing | 13 | 13 | 0 | 0 | 1.00 | 1.00 | 1.00 |
| Decision late (no notice, or notices stopped) | 2 | 2 | 0 | 0 | 1.00 | 1.00 | 1.00 |
| Payment late | 4 | 4 | 0 | 0 | 1.00 | 1.00 | 1.00 |
| Denial reason: no clause, or misstated clause | 12 | 11 | 0 | 1 | 1.00 | 0.92 | 1.00 |
| Department of Insurance review notice missing | 2 | 2 | 0 | 0 | 1.00 | 1.00 | 1.00 |
| Misrepresenting policy provisions | 11 | 10 | 1 | 1 | 0.91 | 0.91 | 1.00 |
| Offer below own valuation, no basis | 3 | 3 | 0 | 0 | 1.00 | 1.00 | 1.00 |
| No investigation before denial | 7 | 7 | 0 | 0 | 1.00 | 1.00 | 1.00 |
| All checks | 54 | 52 | 1 | 2 | 0.98 | 0.96 |
The replies set (12 files) found 6 of 6 late replies, with 0 false flags.
By phrasing:
- held-out phrasing (B, 24 files): 26 found, 0 false, 0 missed;
- dev templates (A, 24 files): 26 found, 1 false, 2 missed.
All three errors are in the model's reading of the letter against the policy:
- Missed (test-016, Texas auto). The letter says Exclusion 1 excludes use "for any business purpose". The policy says "public
or livery conveyance". Grounding graded it partial, and the file-level judgment answered unclear even though its reason
named the mismatch. Both items came back as
review, notflag. A reviewer would see it; the strict count misses it. - False flag (test-042, California homeowners). The letter shortens the flood exclusion ("flood, surface water or overflow of a body of water"), leaving out "waves, tidal water ... spray". The judgment called the omission a misrepresentation. The grounding check on the same sentence said supported.
False alarms on clean files:
- test: 0 of 12 clean files had any flag;
- dev: 0 of 7;
- replies: 0 of 6;
- samples: 0 of 2.
About 5% of all checklist items came back review. These are mostly the "explains how each provision applies" judgment
and model-only AI use, which are review-by-design.
Date computation:
- Dated events: 247 of 247 notice, acknowledgement, status-letter, proof-of-loss, decision and payment dates in the test files were used with the right date and kind (dev: 125 of 126).
- Timeliness items: 122 of 122 had the right start date, act date and deadline, computed end to end from the documents (dev: 64 of 65). Texas decisions use 15 business days (US federal holidays skipped) and Texas payments 5 business days.
- The arithmetic alone is unit-tested against hand-computed dates in
tests/test_claims.py, including Thanksgiving, an observed 4 July, and a deadline on a weekend.
AI use: 31 of 31 test files that mention an AI system recorded it. In all 31, whether a person was involved (model alone vs. a person deciding) was read correctly.
Cost and time:
- Model calls: 2 to 6 per file (median about 4): one extraction, one grounding call per denial reason, and up to four typed judgments.
- The smoke run of the planted California file took 5 calls, 9,351 tokens and 7.6 s. At the gateway's list price that is $0.004.
- Test-set latency under shared load: p50 10.1 s, p90 19.2 s, max 26.4 s.
Honest limits
- Synthetic, templated files. Real claim files are longer and messier: scanned letters, many more notes, email chains, reserve changes, several claimants. The generator's letters state one reason each, in one of six wordings. These numbers show that the pipeline and the rules work. They do not predict accuracy on a carrier's files, so measure on your own closed files before relying on it.
- Small positive counts for some checks: decision (2), review notice (2) and lowball (3) on the test set.
- Rules are a subset. The item list is in
/claims/info, and what is left out is undernot_checked: total-loss valuation, subrogation, statute-of-limitations notices, fraud-extension periods, surplus lines, and "final payment" wording. - NAIC 902 s.7B's cadence ("forty-five (45) days from the initial notification and every forty-five (45) days
thereafter") is read as at most 45 days between notices. That is an interpretation, and
/claims/infostates it. - No roll-forward of deadlines that fall on weekends or holidays. Items one or two days late on such a deadline are
marked
review.
Expected properties of the sample run (for the rehearse-first kit)
ca-auto-planted:ack,reason:1,review_noticeandmisrepresentcome backflag;ack.computed.deadlineis 15 calendar days after the notice date (2026-06-03);reason:1.grounding.spans[0].quotecontains "organized race or speed contest".
ca-home-cleanandtx-auto-cleancome back withcounts.flag == 0.tx-home-lowball:decisionisflagwithcomputed.unit == "business";lowballisflagwithcomputed.offerbelowcomputed.threshold;- there is no
review_noticeitem (Texas ch. 542 has none).
naic-auto-noclause:reason:1isflag("names no policy provision"), and anai_useitem hasdecided_by: "model"and statusreview.ca-home-notices:decisionisflag, and its reason says the next status letter was due 30 days after the first.- Every run returns a record for which
POST /record/verifygivesok: true, andstepslists extract (model), timeliness (rule), denial (model+rule), conduct (model) and human_review (human, pending).
Verdict
A buyer would pay for it as a triage queue and an exam-prep file, not as a compliance product.
- On 48 held-out synthetic files it found 52 of 54 planted problems, with 1 false flag. It raised no false alarm on any of the 12 clean files, and every date it used was right.
- What makes it worth paying for: the timeliness arithmetic is exact and shows its dates, and the misquoted-clause flag quotes the policy line.
- Before a carrier would buy it:
- accuracy measured on real, de-identified closed files, including scanned letters (OCR is not in this build);
- more state packs (Florida 627.70131, New York Reg. 64, and others);
- claims-system connectors (Guidewire, Duck Creek) so every file runs, not just the ones pasted in;
- an examiner-facing export.
- The skeptic's point stands: carriers buy through their claims-system vendor. The likely first buyers are TPAs and claims-audit consultancies.