156 · Finance and insurance · preview
Certificate request check
Eval results
Not held outRun 29 Sep 2026Eval write-up (decosa-api, access required)
- Status right on found requirements109 of 120 (0.91)test splitn = 120third test run after cold-user fixes (runs: 0.91, 0.84, 0.91); CI 0.88-0.95
- Requirements found120 of 124 (0.97)test splitn = 124precision 120 of 144
- 'Met' calls that were right102 of 103test splitn = 103a wrong 'met' is the E&O risk
- 'Met' calls without a policy quote0test splitn = 103
Dataset
27 synthetic contract, request and policy triples (354 labelled requirements, 17 states), written blind by a separate author; stratified split, dev 18, test 9.
Caveats
- Synthetic papers from one author; 78% of labels are 'met', so read the per-status numbers.
- The test set was run three times; the third run followed a fix prompted by the second, so it is not fully held out.
- Not a coverage opinion; it compares wording.
- Slow under load: about two minutes a request on the shared gateway.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 30 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 33 s
- Receipts
- 31
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.015
Self-host verification
Not yet verified on a fresh self-host setup.
Rehearsal bundle: certificate-check.zip (4 KB, 5 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted numbers are the production smoke check of the sample, run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
- Measured on 27 synthetic contract, request and policy triples written by one author; real policies run to 100+ pages of forms.
- One wrong 'met' on the test set (an umbrella that excludes pollution); blanket waivers of subrogation often come back as needing a person.
- Typed or pasted text only; scanned policies need the document reader first.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Lists each requirement from the contract and the email with a verbatim quote, then per requirement says whether the policy papers meet it, quoting the policy; the grounding judge checks each reasonQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 48 GB card (1)
- accuracy on this task: not measured yet
Standard · the hosted demo, one 96 GB card (2)
- status right on found requirements, test set (9 triples, 120 requirements, gateway; third run, not fully held out): 109 of 120 (0.91)decosa-api docs/evals/certificate-check.md, measured on our server 2026-09-29, gateway route
- requirements found / 'met' calls that were right (test, third run): 120 of 124 / 102 of 103decosa-api docs/evals/certificate-check.md, measured on our server 2026-09-29, gateway route
Best · DeepSeek-V4-Flash on two more cards (1)
- accuracy on this task: not measured yet
Wanted · two large judges from different families (1)
- accuracy on this task: not measured yet