Skip to content
decosa

100 · Legal · live

Check their discovery responses

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 28 Sep 2026Eval write-up (decosa-api, access required)

  • Deficient responses caught on unseen real cases (RECAP test, dockets never seen in dev)174 / 235 (74%)test splitn = 235Responses a motion to compel called deficient; 85% (174 / 205) of those the splitter found. All 43 test sets: 292 / 373 (78%).
  • Responses found by the splitter on unseen real filings75.3%test splitn = 92827 docket-disjoint RECAP test sets; 78.8% on all 43; 99.3% on the 25 dev filings it was tuned on. Missed responses are listed as warnings.
  • False flags on clean responses, blind synthetic test13 / 197 (6.6%)held outn = 197Clean responses flagged as deficient. On real filings a blind review of 80 flags the motions did not raise found 45 correct, 11 debatable, 24 wrong (15 were pre-2015 responses, since fixed).
  • Deficient-or-not precision, blind synthetic test0.897held outn = 325113 correct flags, 13 false alarms, 15 missed, 184 clean left alone; sets written by a separate blind author.
  • Deficient-or-not recall, blind synthetic test0.883held outn = 325
  • Per-category F1, blind synthetic test0.792held outn = 325precision 0.715, recall 0.886 before post-test fixes (0.832 after, no longer held out).
  • Recall vs blind Claude Opus 5.5, 90 responses0.952 vs 0.935held outn = 90precision 0.756 vs 0.879
  • Blind partner review: tool letter vs hand-drafted, 9 comparisons0 / 9 preferred (scores 2-5 vs 8-9)test splitn = 9Three rounds, two of them after fixes; reviewers penalised template points instead of request-specific argument.

Dataset

68 sets of real written discovery responses from CourtListener RECAP (46 federal dockets; 25 dev, 43 test by hash, 27 of them on dockets with no dev set) labelled from the motions to compel; 24 synthetic sets (325 responses) written by a separate blind author; 3 synthetic demo sets (dev).

Caveats

  • Motion labels undercount what is wrong, so precision against them is a lower bound; a blind adjudicator judged a sample of the other flags.
  • The adjudicator, frontier judge, letter reviewer and cold users are Claude Opus 5.5 sub-agents, not practising lawyers.
  • All real sets are federal; California and Texas are measured on synthetic sets only.
  • About a third of the real sets were rebuilt from quotes in the motion (disputed items only), and some dockets were split into correlated sets; the docket-disjoint figures are the stricter ones.
  • Fixes made after the held-out runs are reported separately and not counted as held out.
  • The drafted letter lost every blind comparison with a hand-drafted letter.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
28 Sep 2026
Latency, this run
n/a
p50 over passed runs
4.8 s
Receipts
10
Model calls
n/a
Tokens
n/a
Cost per run
$0.006

Self-host verification

Verified on 28 Sep 2026: fresh clone into a clean directory, api image from docker/api/Dockerfile, compose with a named volume, direct route to the local Qwen3.8-27B, local signing; torn down after

The rehearsal bundle passed 11/11 in 2.8 s; the three samples ran in 2.4-3.6 s with attested receipts.

Rehearsal bundle: discovery-deficiency.zip (4 KB, 11 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • The splitter finds about 3 in 4 responses in real court-filed PDFs it has not seen (75-79%); missing ones are listed as warnings.
  • The letter draft is a starting point, not ready to send: a blind partner review preferred hand-drafted letters on every set (9 of 9), citing missing request-specific argument.
  • Rule text only: no case law, local rules or standing orders. California uses the CCP 135 court calendar; Texas deadlines use the federal holiday list.
  • Hosted numbers are from the pre-release server before merge; production numbers follow the nightly check.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Splitter, set checks (verification, signature, deadlines), flag rules, quote location, rules pack, letter and fix list, signed record (no model; CPU)decosa-api discovery check (decosa_api/verticals/discovery), with the dates block, the drafting editor's Word writer and the signed record (07)AGPL-3.0-or-later
  • Reads each response once and describes it as JSON: objection grounds and whether each gives specifics, withholding statement, production date, answer shape, admission shapeQwen3.8-27B (NVFP4)Apache-2.0
  • Scanned PDFs only: page images to textDocument reader (Docling layout heron + PaddleOCR-VL-1.6)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 32 GB card (1)
  • Same model and prompts as standard: not measured separatelyestimate: identical pipeline without the document reader
Standard · Qwen3.8-27B and the document reader (hosted demo) (6)
  • Responses correctly called deficient or not, blind synthetic test (24 sets by another author, 325 responses, federal, California, Texas): precision 0.897, recall 0.883; 184 clean responses left alonedecosa-api docs/evals/discovery-deficiency.md, held out, run once
  • Real responses a motion to compel called deficient, caught (27 docket-disjoint RECAP test sets): 174 / 235 (74%); 174 / 205 (85%) of the responses the splitter founddecosa-api docs/evals/discovery-deficiency.md
  • Clean responses wrongly flagged, blind synthetic test: 13 / 197 (6.6%)decosa-api docs/evals/discovery-deficiency.md, held out
  • Splitter coverage on real filings (held out): 75.3% of responses found on docket-disjoint sets (78.8% on all 43; 99.3% on the dev filings it was tuned on)decosa-api docs/evals/discovery-deficiency.md
  • Against a blind frontier judge (Claude Opus 5.5), 90 held-out responses: recall 0.952 vs 0.935; precision 0.756 vs 0.879decosa-api docs/evals/discovery-deficiency.md
  • Draft letter vs a hand-drafted letter, blind partner review: a starting point, not ready to send: it lost 9 of 9 blind comparisons with a hand-drafted letter (2-5 vs 8-9 of 10)decosa-api docs/evals/discovery-deficiency.md

How we measure · All tools