Skip to content
decosa

156 · Finance and insurance · preview

Certificate request check

Open the toolJSON

Eval results

Not held outRun 29 Sep 2026Eval write-up (decosa-api, access required)

  • Status right on found requirements109 of 120 (0.91)test splitn = 120third test run after cold-user fixes (runs: 0.91, 0.84, 0.91); CI 0.88-0.95
  • Requirements found120 of 124 (0.97)test splitn = 124precision 120 of 144
  • 'Met' calls that were right102 of 103test splitn = 103a wrong 'met' is the E&O risk
  • 'Met' calls without a policy quote0test splitn = 103

Dataset

27 synthetic contract, request and policy triples (354 labelled requirements, 17 states), written blind by a separate author; stratified split, dev 18, test 9.

Caveats

  • Synthetic papers from one author; 78% of labels are 'met', so read the per-status numbers.
  • The test set was run three times; the third run followed a fix prompted by the second, so it is not fully held out.
  • Not a coverage opinion; it compares wording.
  • Slow under load: about two minutes a request on the shared gateway.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
30 Sep 2026
Latency, this run
n/a
p50 over passed runs
33 s
Receipts
31
Model calls
n/a
Tokens
n/a
Cost per run
$0.015

Self-host verification

Not yet verified on a fresh self-host setup.

Rehearsal bundle: certificate-check.zip (4 KB, 5 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted numbers are the production smoke check of the sample, run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
  • Measured on 27 synthetic contract, request and policy triples written by one author; real policies run to 100+ pages of forms.
  • One wrong 'met' on the test set (an umbrella that excludes pollution); blanket waivers of subrogation often come back as needing a person.
  • Typed or pasted text only; scanned policies need the document reader first.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Lists each requirement from the contract and the email with a verbatim quote, then per requirement says whether the policy papers meet it, quoting the policy; the grounding judge checks each reasonQwen3.8-27B (NVIDIA NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 48 GB card (1)
  • accuracy on this task: not measured yet
Standard · the hosted demo, one 96 GB card (2)
  • status right on found requirements, test set (9 triples, 120 requirements, gateway; third run, not fully held out): 109 of 120 (0.91)decosa-api docs/evals/certificate-check.md, measured on our server 2026-09-29, gateway route
  • requirements found / 'met' calls that were right (test, third run): 120 of 124 / 102 of 103decosa-api docs/evals/certificate-check.md, measured on our server 2026-09-29, gateway route
Best · DeepSeek-V4-Flash on two more cards (1)
  • accuracy on this task: not measured yet
Wanted · two large judges from different families (1)
  • accuracy on this task: not measured yet

How we measure · All tools