100 · Legal · live
Check their discovery responses
Eval results
Scored on a held-out or test splitRun 28 Sep 2026Eval write-up (decosa-api, access required)
- Deficient responses caught on unseen real cases (RECAP test, dockets never seen in dev)174 / 235 (74%)test splitn = 235Responses a motion to compel called deficient; 85% (174 / 205) of those the splitter found. All 43 test sets: 292 / 373 (78%).
- Responses found by the splitter on unseen real filings75.3%test splitn = 92827 docket-disjoint RECAP test sets; 78.8% on all 43; 99.3% on the 25 dev filings it was tuned on. Missed responses are listed as warnings.
- False flags on clean responses, blind synthetic test13 / 197 (6.6%)held outn = 197Clean responses flagged as deficient. On real filings a blind review of 80 flags the motions did not raise found 45 correct, 11 debatable, 24 wrong (15 were pre-2015 responses, since fixed).
- Deficient-or-not precision, blind synthetic test0.897held outn = 325113 correct flags, 13 false alarms, 15 missed, 184 clean left alone; sets written by a separate blind author.
- Deficient-or-not recall, blind synthetic test0.883held outn = 325
- Per-category F1, blind synthetic test0.792held outn = 325precision 0.715, recall 0.886 before post-test fixes (0.832 after, no longer held out).
- Recall vs blind Claude Opus 5.5, 90 responses0.952 vs 0.935held outn = 90precision 0.756 vs 0.879
- Blind partner review: tool letter vs hand-drafted, 9 comparisons0 / 9 preferred (scores 2-5 vs 8-9)test splitn = 9Three rounds, two of them after fixes; reviewers penalised template points instead of request-specific argument.
Dataset
68 sets of real written discovery responses from CourtListener RECAP (46 federal dockets; 25 dev, 43 test by hash, 27 of them on dockets with no dev set) labelled from the motions to compel; 24 synthetic sets (325 responses) written by a separate blind author; 3 synthetic demo sets (dev).
Caveats
- Motion labels undercount what is wrong, so precision against them is a lower bound; a blind adjudicator judged a sample of the other flags.
- The adjudicator, frontier judge, letter reviewer and cold users are Claude Opus 5.5 sub-agents, not practising lawyers.
- All real sets are federal; California and Texas are measured on synthetic sets only.
- About a third of the real sets were rebuilt from quotes in the motion (disputed items only), and some dockets were split into correlated sets; the docket-disjoint figures are the stricter ones.
- Fixes made after the held-out runs are reported separately and not counted as held out.
- The drafted letter lost every blind comparison with a hand-drafted letter.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 28 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 4.8 s
- Receipts
- 10
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.006
Self-host verification
Verified on 28 Sep 2026: fresh clone into a clean directory, api image from docker/api/Dockerfile, compose with a named volume, direct route to the local Qwen3.8-27B, local signing; torn down after
The rehearsal bundle passed 11/11 in 2.8 s; the three samples ran in 2.4-3.6 s with attested receipts.
Rehearsal bundle: discovery-deficiency.zip (4 KB, 11 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- The splitter finds about 3 in 4 responses in real court-filed PDFs it has not seen (75-79%); missing ones are listed as warnings.
- The letter draft is a starting point, not ready to send: a blind partner review preferred hand-drafted letters on every set (9 of 9), citing missing request-specific argument.
- Rule text only: no case law, local rules or standing orders. California uses the CCP 135 court calendar; Texas deadlines use the federal holiday list.
- Hosted numbers are from the pre-release server before merge; production numbers follow the nightly check.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Splitter, set checks (verification, signature, deadlines), flag rules, quote location, rules pack, letter and fix list, signed record (no model; CPU)decosa-api discovery check (decosa_api/verticals/discovery), with the dates block, the drafting editor's Word writer and the signed record (07)AGPL-3.0-or-later
- Reads each response once and describes it as JSON: objection grounds and whether each gives specifics, withholding statement, production date, answer shape, admission shapeQwen3.8-27B (NVFP4)Apache-2.0
- Scanned PDFs only: page images to textDocument reader (Docling layout heron + PaddleOCR-VL-1.6)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 32 GB card (1)
- Same model and prompts as standard: not measured separatelyestimate: identical pipeline without the document reader
Standard · Qwen3.8-27B and the document reader (hosted demo) (6)
- Responses correctly called deficient or not, blind synthetic test (24 sets by another author, 325 responses, federal, California, Texas): precision 0.897, recall 0.883; 184 clean responses left alonedecosa-api docs/evals/discovery-deficiency.md, held out, run once
- Real responses a motion to compel called deficient, caught (27 docket-disjoint RECAP test sets): 174 / 235 (74%); 174 / 205 (85%) of the responses the splitter founddecosa-api docs/evals/discovery-deficiency.md
- Clean responses wrongly flagged, blind synthetic test: 13 / 197 (6.6%)decosa-api docs/evals/discovery-deficiency.md, held out
- Splitter coverage on real filings (held out): 75.3% of responses found on docket-disjoint sets (78.8% on all 43; 99.3% on the dev filings it was tuned on)decosa-api docs/evals/discovery-deficiency.md
- Against a blind frontier judge (Claude Opus 5.5), 90 held-out responses: recall 0.952 vs 0.935; precision 0.756 vs 0.879decosa-api docs/evals/discovery-deficiency.md
- Draft letter vs a hand-drafted letter, blind partner review: a starting point, not ready to send: it lost 9 of 9 blind comparisons with a hand-drafted letter (2-5 vs 8-9 of 10)decosa-api docs/evals/discovery-deficiency.md