Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: discovery deficiency check (use case 100)

28 Sep 2026. Model: Qwen3.8-27B through the model gateway, temperature 0, thinking off, one JSON call per response (prompt discovery.v3). Code assigns categories and rule cites from the rules pack (100+ verbatim quotes of FRCP, CCP and TRCP text, read 28 Sep 2026). Scripts: scripts/discovery_eval.py (sets samples, synth, recap); outputs in docs/evals/discovery-deficiency/ (the model cache is kept out of git).

Data

Set What Split Size Labels Licence
Demo samples 3 synthetic sets written by the builder (federal RFP, California special interrogatories, our draft RFAs) dev (tuned on) 26 responses the builder's CC0
RECAP Real written responses attached as exhibits to motions to compel, from CourtListener RECAP (public federal court records; 46 dockets, 21 districts, filed 2003-2026), collected by a separate agent 25 sets dev / 43 sets test (fixed by sha256 of the set id; 16 test sets share a docket with a dev set, see below) 615 / 928 responses what the motion (or order) says is deficient, mapped to the taxonomy by the collecting agent, not by the builder public records
Synthetic test 24 sets written by a blind sub-agent playing a litigation associate (12 federal, 6 California, 6 Texas), with planted defects, clean controls and traps (e.g. "reasonably calculated" in California, weekend deadlines) test only; the builder never read the sets or keys before scoring 325 responses (197 clean) the author's key CC0

Raw RECAP data: our server (motions, orders, responses text, labels, the collector's README with its method and biases; 30 sets have the court's ruling order).

Motions raise only what the movant chose to raise, so on RECAP a flag the motion doesn't mention is not necessarily wrong: precision against motion labels is a lower bound. A blind adjudicator checked a sample of those flags (below).

Held-out results (frozen code 02d3417, before any change prompted by test data)

Blind synthetic test (24 sets, 325 responses, run once):

  • Parse coverage: 311 of 325 responses found (95.7%); one California set (ca_rfp_01) did not split.
  • Per response, deficient or not: precision 0.897, recall 0.883 (113 flagged correctly, 13 false alarms, 15 missed, 184 clean responses left alone).
  • Per (response, category): precision 0.715, recall 0.886, F1 0.792.
  • By category (caught / labelled, false flags): objections without specifics 50/53, 6; incorporated general objections 27/27, 0; unstated withholding 23/27, 6; no production date 13/19, 0; evasive or incomplete 19/22, 22; improper RFA answers 20/21, 3; privilege without a log 12/17, 31; pre-2015 standard 7/7, 0 (no false flags on the California and Texas traps).
  • Set level: general objections 6/6, missing verification 4/4, missing signature 1/1, late service 4/6 (no false flags).
  • Most false flags come from two rules adopted from the real movants' practice on the RECAP dev split: responses that incorporate a general-objections block claiming privilege are flagged "privilege without a log" (19 of the 31), and objections-only responses with no specifics are flagged "no answer" (most of the 22). The synthetic author labels those only as incorporated/boilerplate objections. It is a disagreement between label sources as much as an error.

RECAP test (43 sets of real responses, 928 responses listed by the labeller, run once):

  • Parse coverage 78.8% (736 responses found); two sets did not split. On dev the splitter reached 99.3% after tuning on dev layouts, so the splitter does not generalise as well as the model's judgment.
  • Per response, deficient or not, against the motions: recall 0.783 (292 of 373). Among responses that were split, recall is 0.896 (292 of 326); 47 of the 81 misses are responses that were not split. Response-level precision cannot be measured here: only responses a motion called deficient are scored, so a false alarm is impossible by construction. False flags are measured on the synthetic set (13 of 197 clean responses flagged, 6.6%) and by the blind adjudication below.
  • Per (response, category), against motion labels: precision 0.557 (lower bound), recall 0.498.
  • Docket overlap (found after the run): 16 of the 43 test sets share a docket with a dev set (the collector split some dockets into one set per party), so they share lawyers and layouts with what the splitter was tuned on. On the 27 docket-disjoint test sets (the stricter held-out figure): splitter coverage 75.3%; deficient responses caught 0.740 (174 of 235), 0.849 on the responses that were split (174 of 205); per (response, category) precision 0.496 (lower bound), recall 0.475.
  • Rebuilt texts: in about 20 of the 68 sets (8 test sets) the served responses were never filed, so the collector rebuilt responses.txt from the verbatim request/response quotes in the motion or joint stipulation. Those texts hold only the disputed items, usually without general objections, signature or verification, which is why set-level recall on RECAP is low and why a copy that ends at the last response now skips the signature and verification checks.
  • Blind adjudication of 80 flags the motions did not raise (a sample of 1,282; Claude Code Opus 5.5 sub-agent, told the filing date, shown the response, the flag and its reason): 45 correct, 11 debatable, 24 wrong. 15 of the 24 wrong ones were the 2015-amendment duties (withholding statement, production date, "reasonably calculated") applied to responses served in 2006-2014; others: a request-specific objection called boilerplate, text from a neighbouring interrogatory, a cross-reference read as "see documents", privacy or relevance read as privilege.

Frontier comparison (blind): Claude Code Opus 5.5 as a sub-agent, given only the category definitions and 90 responses sampled from the synthetic test set (seed 98), judged each one; scored by us against the keys.

On the same 90 responses Precision Recall F1 Deficient-or-not P / R
Blind Opus 5.5 0.879 0.935 0.906 1.00 / 0.95
This tool (Qwen3.8-27B + code) 0.756 0.952 0.843 0.97 / 0.92

Recall is on par; the gap is precision, mostly the incorporated-privilege rule (10 of 19 false flags), which the judge could not apply because it saw each response without the general-objections block.

Fixes after the held-out runs (disclosed: these numbers are no longer held out)

The blind letter reviews, the cold-user tests and the adjudication above led to fixes (commits 0a835ec, f11b023): the 2015 duties are not checked for federal responses served before 1 Dec 2015 (or with a pre-2015 case number and no service date); the incorporated-privilege rule is limited to requests for production; California court holidays (CCP 135, 12a); PDF header noise is never quoted; Bates ranges with mixed-case prefixes are recognised; plus letter fixes. Re-scored on the same model outputs where the prompt did not change:

  • Synthetic test: precision 0.784, recall 0.886, F1 0.832 (deficient-or-not unchanged, 0.897 / 0.883).
  • RECAP test: per (response, category) precision 0.590, recall 0.496; deficient-or-not unchanged.
  • Demo samples (dev): 34 / 34 category flags, 0 false.

Dev results (tuned on)

  • Demo samples: 34 / 34, 0 false flags, 8 clean responses clean.
  • RECAP dev: parse 99.3% (611 / 615); deficient-or-not recall 0.987; per category recall 0.858, precision 0.625 (lower bound).

The letter (blind reviews)

The letter is written by code from the flags. A blind senior-partner persona (Claude Code Opus 5.5 sub-agent, not told which letter came from where) compared it with letters hand-drafted by a blind associate sub-agent on the same sets:

Round Sets Tool letter scores Hand-drafted scores Partner preferred
1 (frozen code) 2 RECAP test sets + the California sample 3, 3, 3 (75-90 min to edit) 8, 8, 9 hand-drafted, 3 of 3
2 (after fixes) same sets 4, 3, 5 (60-100 min) 8, 8, 9 hand-drafted, 3 of 3
3 (fresh sets, code as of f11b023) 2 other RECAP test sets (N.D. Ga. 2014, N.D. Ill.) + a Texas synthetic set 2, 3, 3 (90-120 min) 8, 8, 8 (30 min) hand-drafted, 3 of 3

What the reviewers credited: correct rule text, correct California timing, waiver and verification analysis (round 2), quoting the response. What they penalised: template repetition instead of request-specific argument, no relevance or merits argument, missed narrowing of production and protective-order holdbacks, and (round 1) quoting PDF layout noise, asking for support of objections it had found waived, and offering to extend a motion deadline that had passed. The letter is a structured first draft of the deficiency list; it does not replace the associate's argument. Round 3 also found plain errors, fixed afterwards but not re-reviewed (commit 53c0b42): Rule 36(a)(5) paraphrased as requiring "specificity", a served Texas withholding statement not recognised (a log demanded contrary to TRCP 193.3(b)), TRCP 196.2(b) applied as a general production-date rule, and "answer the remainder" where nothing was answered. On messy real sets it also misattributed objection grounds between neighbouring responses (splitter bleed).

Time and cost

  • Model cost at gateway list price ($0.30 / $1.50 per million tokens): $0.00087 per response on real responses (747 calls, 1,099,967 prompt and 214,795 completion tokens for 736 responses); $0.00059 per response on the synthetic samples. The ten-response federal sample costs about $0.006.
  • Time: the federal sample took p50 4.8 s (5 runs, max 13.6 s) on the pre-release server through the shared gateway; the three samples took 7.5, 8.5 and 11.4 s uncached in the eval harness. Real sets of 10-100 responses took a median 52 s per set during the RECAP test run, which ran 3 sets at once (above the 6-concurrent-call guideline; later runs used one set at a time).
  • Manual baseline: a blind associate sub-agent estimated what a mid-level associate would bill to review and letter each set by hand: 165, 180 and 130 minutes (round 1 sets) and 170, 140 and 210 minutes (round 3 sets, a second blind associate). These are estimates, not measured human time.
  • Cold users (blind sub-agents, not real lawyers): a litigation paralegal estimated 1.5-2 hours saved per set; a solo litigator about 1 hour (review and calendaring saved; the letter still needs rewriting).

Checkable properties of the sample runs (for the rehearsal kit)

  1. harbor-rfp: set flags general_objections and untimely (due 2026-05-01, served 2026-05-04); withholding_unstated on RFP 2, 3, 5, 8 and 10; no flag on RFP 1, 4 and 7; every rule cite in the rules pack.
  2. marlowe-rogs: missing_verification (the proof of service's oath is ignored) and untimely (due 2026-04-17 by CCP 1013(a)); no outdated_standard; motion date 2026-06-09.
  3. delgado-rfa-draft (ours): rfa_improper on RFA 2, 6, 7; no flag on RFA 1, 3, 5; no letter; fix rows and a reminder to sign.
  4. Every model call has a receipt; the signed record verifies at /record/verify.
  5. Letter: no rule id outside the pack; no quote containing PDF header text.

Own model (page 72, M8)

Not needed now. The model's judgment on split responses is on par with a blind frontier judge for recall (0.95 vs 0.94 on 90 held-out responses) and costs under a tenth of a cent per response. The failure that matters is splitting real PDFs (75-79% held-out coverage), which a response-segmentation model (headings, request/response boundaries, OCR margin noise) trained on RECAP exhibits could fix; flagged as a capability opportunity.

Limits

  • Labels: motion labels undercount (lower-bound precision); the synthetic key is one author's convention; the adjudicator and reviewers are the same model family (Claude Opus 5.5), not practising lawyers.
  • Federal sets dominate the real data (all RECAP sets are federal courts); California and Texas are tested on synthetic sets only.
  • RECAP: dockets split into several sets are correlated, and 16 test sets share a docket with dev; the docket-disjoint figures are the ones to quote. About a third of the sets are rebuilt from motion quotes (disputed items only).
  • Rule text, not case law, local rules or standing orders. State deadlines: California uses CCP 135 holidays; Texas uses the federal holiday list.

Verdict

Would a buyer pay? For the triage, maybe; for the letter, no.

  • What works: on text that splits cleanly, the per-response triage is good and cheap. Held out: deficient-or-not precision 0.897 / recall 0.883 on a blind synthetic set, recall 0.85-0.90 on real split responses (0.849 docket-disjoint), recall on par with a blind frontier judge, $0.0009 per response, seconds per set. Deadlines (including California court holidays and the 45-day clock), verification and waiver logic, and verbatim rule text are things a chat assistant does not give a lawyer.
  • What does not: (1) the splitter found only 75-79% of responses in held-out real PDFs (75.3% on docket-disjoint sets); (2) the drafted letter lost all 9 blind comparisons with a hand-drafted letter (scores 2-5 vs 8-9), because it lists template points instead of arguing each request; (3) per-category precision on real sets is middling (about 0.6 against motions, 56-70% of other flags judged correct by a blind adjudicator before the pre-2015 fix).
  • Cold users (sub-agent personas): both "maybe". The paralegal saw 1.5-2 hours saved per set and a firm price of $750-1,500 a month, but only for a hosted tier cleared for client material; the solo saw about an hour saved and would pay $49-79 a month or $10-15 a set for a zero-retention hosted version. Neither would self-host a GPU box.

Recommendation: keep it on the site as a triage tool, and rework before promoting it. In order: (a) sell it as "triage, deadlines and the deficiency list", with the letter as an export of that list, not a drafted letter; (b) make the splitter robust on real PDFs (a model fallback when headings are missed, or our own segmentation model trained on RECAP exhibits); (c) run it on the Confidential tier (single Qwen3.8 call type, like the tools proven there on 28 Sep), which is the purchase trigger both personas named; (d) only then invest in request-specific letter argument (a model-written paragraph per request, checked against the response).