Eval: certificate-check (Check a certificate request against the contract and the policy)
29 Sep 2026, build-insurance-finance-opus. Runner: scripts/insfin_eval/certcheck_eval.py. Data and raw results:
docs/evals/certificate-check/.
Data
- 27 synthetic triples (contract insurance article or exhibit, the holder's request email, 1-4 policies each with
declarations, a forms schedule and paraphrased endorsement wording), 354 labelled requirements, written blind by
a separate author (a sub-agent that never saw the tool). Fictional companies and insurers, made-up form numbers, no ISO
or ACORD text;
check_quotes.pyconfirms every contract, email and policy quote is verbatim. 17 states. - Labels per requirement: kind, line, required value, section, quote, and a status: met 276, needs an endorsement 45, a certificate can't do this 15, the papers don't say 18. Planted: limit shortfalls (with and without an umbrella the contract lets count), completed-ops AI missing or time-limited, AI scheduled to the wrong party, waivers missing on one line, primary/non-contributory missing, per-project aggregate missing, umbrella not following form, notice asks, holder-wording asks, emails asking the certificate itself to add a lender or "confirm no exclusions", policies not sent.
- Split: stratified (all-met vs with gaps), every third case by id held out: dev 18 cases (230 requirements), test 9 cases (124 requirements). Prompts and code rules were changed on dev only (four rounds); the test set was run on the gateway route after they were frozen (three runs; see Results).
What was measured
- Extraction: a labelled requirement counts as found when the tool lists the same kind with overlapping quote words (or the same section); precision is found / listed.
- Status accuracy on found requirements (4 classes), with a case-level bootstrap CI.
- "Met" precision (a wrong "met" is the agency's E&O risk) and "met" calls with no policy quote found (must be 0).
Results
The test set was run three times on the gateway, and this page reports all three:
- gw1, before the cold-user fixes (held out, run once): the first frozen version.
- gw2, after the fixes a blind agency account manager asked for on two dev samples (every email ask as its own row, the parties each requirement protects, one carrier letter per carrier, paste text only for what the policy gives). One of those fixes was a judge instruction ("a party named only in the email isn't covered by a blanket endorsement") that over-fired on blanket endorsements: 13 "met" came back "needs an endorsement".
- gw3, with that rule moved from the prompt into code and scoped to parties the contract doesn't name (checked on dev first: 204 / 212). Because gw2's result on the test set prompted this change, gw3 is not fully held out.
| dev 18 (direct) | test gw1 (held out) | test gw2 | test gw3 (shipped code) | |
|---|---|---|---|---|
| Requirements found (recall) | 212 / 230 = 0.92 | 119 / 124 = 0.96 | 119 / 124 | 120 / 124 = 0.97 (95% CI 0.93-1.00) |
| Listed requirements that are labelled (precision) | 212 / 248 = 0.85 | 119 / 139 = 0.86 | 119 / 142 | 120 / 144 = 0.83 (0.80-0.87) |
| Status right on found requirements | 204 / 212 = 0.96 | 108 / 119 = 0.91 | 100 / 119 = 0.84 | 109 / 120 = 0.91 (0.88-0.95) |
| "Met" that was right | 175 / 176 | 103 / 105 | 93 / 94 | 102 / 103 = 0.99 (0.95-1.00) |
| "Met" with no policy quote found | 0 | 0 | 0 | 0 |
| Cost per request, gateway list price | $0.017 | $0.018 | $0.020 | $0.019 mean (about 30 calls) |
| Time per request, shared gateway under load | 67 s p50 | 104 s p50 | 138 s p50 | 117 s p50, 141 s p95 |
gw3 confusion (label → tool): met 89 met, 4 needs an endorsement, 4 can't tell; needs an endorsement 13 right, 1 called met; can't-be-done-by-a-certificate 4 right, 1 called needs an endorsement; papers don't say 3 right, 1 called needs an endorsement.
- The one wrong "met" (gw3): a pollution limit where the umbrella the contract lets count excludes pollution; the policy wording is quoted, so a reader sees why.
- The common miss: blanket waivers of subrogation (GL and WC) came back "needs an endorsement" or "papers don't say" on 6 requirements: safe in direction (a person looks) but noisy.
- The extra rows (precision 0.83) are mostly asks split in two (a limit and its aggregate named together) and requirements the author folded into one.
- Dev rounds, for the record: 0.84 → 0.85 → 0.92 → 0.97 → 0.96 status accuracy (silence in a forms schedule means "not met"; quotes joined with "..." accepted when every fragment is found; "a policy of that line wasn't sent" is "papers don't say" by rule; the email-only-party rule in code).
Checkable properties of the sample runs (rehearsal)
tx-framing-gaps: completed-operations AI and the WC waiver come back "needs an endorsement"; the lender-as-AI ask is "a certificate can't do this"; the carrier request quotes contract sections D.2 and D.5.- Every "met" row carries a policy quote found in the papers.
fl-clean: no requirement needs an endorsement.- The record verifies at
POST /record/verify; no output contains an ACORD form or layout.
Caveats
- Synthetic papers: real policies are 100+ pages of ISO or carrier forms; agencies often hold only declarations and forms schedules. One author wrote all 27 triples; "met" is 78% of labels, so a tool leaning to "met" scores well on accuracy: read the per-status rows, not only the total.
- The label policy for notice-of-cancellation asks (certificate can't vs needs an endorsement) is itself debatable; the tool says "needs an endorsement" and, in the holder note, that a certificate can't promise notice.
- Not a coverage opinion: it compares wording, it doesn't say whether a claim would be covered.