Eval: discovery deficiency check (use case 100)
28 Sep 2026. Model: Qwen3.8-27B through the model gateway, temperature 0, thinking off, one JSON call per
response (prompt discovery.v3). Code assigns categories and rule cites from the rules pack (100+ verbatim quotes of FRCP,
CCP and TRCP text, read 28 Sep 2026). Scripts: scripts/discovery_eval.py (sets samples, synth, recap); outputs in
docs/evals/discovery-deficiency/ (the model cache is kept out of git).
Data
| Set | What | Split | Size | Labels | Licence |
|---|---|---|---|---|---|
| Demo samples | 3 synthetic sets written by the builder (federal RFP, California special interrogatories, our draft RFAs) | dev (tuned on) | 26 responses | the builder's | CC0 |
| RECAP | Real written responses attached as exhibits to motions to compel, from CourtListener RECAP (public federal court records; 46 dockets, 21 districts, filed 2003-2026), collected by a separate agent | 25 sets dev / 43 sets test (fixed by sha256 of the set id; 16 test sets share a docket with a dev set, see below) | 615 / 928 responses | what the motion (or order) says is deficient, mapped to the taxonomy by the collecting agent, not by the builder | public records |
| Synthetic test | 24 sets written by a blind sub-agent playing a litigation associate (12 federal, 6 California, 6 Texas), with planted defects, clean controls and traps (e.g. "reasonably calculated" in California, weekend deadlines) | test only; the builder never read the sets or keys before scoring | 325 responses (197 clean) | the author's key | CC0 |
Raw RECAP data: our server (motions, orders, responses text, labels, the collector's README with its method and biases; 30 sets have the court's ruling order).
Motions raise only what the movant chose to raise, so on RECAP a flag the motion doesn't mention is not necessarily wrong: precision against motion labels is a lower bound. A blind adjudicator checked a sample of those flags (below).
Held-out results (frozen code 02d3417, before any change prompted by test data)
Blind synthetic test (24 sets, 325 responses, run once):
- Parse coverage: 311 of 325 responses found (95.7%); one California set (ca_rfp_01) did not split.
- Per response, deficient or not: precision 0.897, recall 0.883 (113 flagged correctly, 13 false alarms, 15 missed, 184 clean responses left alone).
- Per (response, category): precision 0.715, recall 0.886, F1 0.792.
- By category (caught / labelled, false flags): objections without specifics 50/53, 6; incorporated general objections 27/27, 0; unstated withholding 23/27, 6; no production date 13/19, 0; evasive or incomplete 19/22, 22; improper RFA answers 20/21, 3; privilege without a log 12/17, 31; pre-2015 standard 7/7, 0 (no false flags on the California and Texas traps).
- Set level: general objections 6/6, missing verification 4/4, missing signature 1/1, late service 4/6 (no false flags).
- Most false flags come from two rules adopted from the real movants' practice on the RECAP dev split: responses that incorporate a general-objections block claiming privilege are flagged "privilege without a log" (19 of the 31), and objections-only responses with no specifics are flagged "no answer" (most of the 22). The synthetic author labels those only as incorporated/boilerplate objections. It is a disagreement between label sources as much as an error.
RECAP test (43 sets of real responses, 928 responses listed by the labeller, run once):
- Parse coverage 78.8% (736 responses found); two sets did not split. On dev the splitter reached 99.3% after tuning on dev layouts, so the splitter does not generalise as well as the model's judgment.
- Per response, deficient or not, against the motions: recall 0.783 (292 of 373). Among responses that were split, recall is 0.896 (292 of 326); 47 of the 81 misses are responses that were not split. Response-level precision cannot be measured here: only responses a motion called deficient are scored, so a false alarm is impossible by construction. False flags are measured on the synthetic set (13 of 197 clean responses flagged, 6.6%) and by the blind adjudication below.
- Per (response, category), against motion labels: precision 0.557 (lower bound), recall 0.498.
- Docket overlap (found after the run): 16 of the 43 test sets share a docket with a dev set (the collector split some dockets into one set per party), so they share lawyers and layouts with what the splitter was tuned on. On the 27 docket-disjoint test sets (the stricter held-out figure): splitter coverage 75.3%; deficient responses caught 0.740 (174 of 235), 0.849 on the responses that were split (174 of 205); per (response, category) precision 0.496 (lower bound), recall 0.475.
- Rebuilt texts: in about 20 of the 68 sets (8 test sets) the served responses were never filed, so the collector
rebuilt
responses.txtfrom the verbatim request/response quotes in the motion or joint stipulation. Those texts hold only the disputed items, usually without general objections, signature or verification, which is why set-level recall on RECAP is low and why a copy that ends at the last response now skips the signature and verification checks. - Blind adjudication of 80 flags the motions did not raise (a sample of 1,282; Claude Code Opus 5.5 sub-agent, told the filing date, shown the response, the flag and its reason): 45 correct, 11 debatable, 24 wrong. 15 of the 24 wrong ones were the 2015-amendment duties (withholding statement, production date, "reasonably calculated") applied to responses served in 2006-2014; others: a request-specific objection called boilerplate, text from a neighbouring interrogatory, a cross-reference read as "see documents", privacy or relevance read as privilege.
Frontier comparison (blind): Claude Code Opus 5.5 as a sub-agent, given only the category definitions and 90 responses sampled from the synthetic test set (seed 98), judged each one; scored by us against the keys.
| On the same 90 responses | Precision | Recall | F1 | Deficient-or-not P / R |
|---|---|---|---|---|
| Blind Opus 5.5 | 0.879 | 0.935 | 0.906 | 1.00 / 0.95 |
| This tool (Qwen3.8-27B + code) | 0.756 | 0.952 | 0.843 | 0.97 / 0.92 |
Recall is on par; the gap is precision, mostly the incorporated-privilege rule (10 of 19 false flags), which the judge could not apply because it saw each response without the general-objections block.
Fixes after the held-out runs (disclosed: these numbers are no longer held out)
The blind letter reviews, the cold-user tests and the adjudication above led to fixes (commits 0a835ec, f11b023): the 2015 duties are not checked for federal responses served before 1 Dec 2015 (or with a pre-2015 case number and no service date); the incorporated-privilege rule is limited to requests for production; California court holidays (CCP 135, 12a); PDF header noise is never quoted; Bates ranges with mixed-case prefixes are recognised; plus letter fixes. Re-scored on the same model outputs where the prompt did not change:
- Synthetic test: precision 0.784, recall 0.886, F1 0.832 (deficient-or-not unchanged, 0.897 / 0.883).
- RECAP test: per (response, category) precision 0.590, recall 0.496; deficient-or-not unchanged.
- Demo samples (dev): 34 / 34 category flags, 0 false.
Dev results (tuned on)
- Demo samples: 34 / 34, 0 false flags, 8 clean responses clean.
- RECAP dev: parse 99.3% (611 / 615); deficient-or-not recall 0.987; per category recall 0.858, precision 0.625 (lower bound).
The letter (blind reviews)
The letter is written by code from the flags. A blind senior-partner persona (Claude Code Opus 5.5 sub-agent, not told which letter came from where) compared it with letters hand-drafted by a blind associate sub-agent on the same sets:
| Round | Sets | Tool letter scores | Hand-drafted scores | Partner preferred |
|---|---|---|---|---|
| 1 (frozen code) | 2 RECAP test sets + the California sample | 3, 3, 3 (75-90 min to edit) | 8, 8, 9 | hand-drafted, 3 of 3 |
| 2 (after fixes) | same sets | 4, 3, 5 (60-100 min) | 8, 8, 9 | hand-drafted, 3 of 3 |
| 3 (fresh sets, code as of f11b023) | 2 other RECAP test sets (N.D. Ga. 2014, N.D. Ill.) + a Texas synthetic set | 2, 3, 3 (90-120 min) | 8, 8, 8 (30 min) | hand-drafted, 3 of 3 |
What the reviewers credited: correct rule text, correct California timing, waiver and verification analysis (round 2), quoting the response. What they penalised: template repetition instead of request-specific argument, no relevance or merits argument, missed narrowing of production and protective-order holdbacks, and (round 1) quoting PDF layout noise, asking for support of objections it had found waived, and offering to extend a motion deadline that had passed. The letter is a structured first draft of the deficiency list; it does not replace the associate's argument. Round 3 also found plain errors, fixed afterwards but not re-reviewed (commit 53c0b42): Rule 36(a)(5) paraphrased as requiring "specificity", a served Texas withholding statement not recognised (a log demanded contrary to TRCP 193.3(b)), TRCP 196.2(b) applied as a general production-date rule, and "answer the remainder" where nothing was answered. On messy real sets it also misattributed objection grounds between neighbouring responses (splitter bleed).
Time and cost
- Model cost at gateway list price ($0.30 / $1.50 per million tokens): $0.00087 per response on real responses (747 calls, 1,099,967 prompt and 214,795 completion tokens for 736 responses); $0.00059 per response on the synthetic samples. The ten-response federal sample costs about $0.006.
- Time: the federal sample took p50 4.8 s (5 runs, max 13.6 s) on the pre-release server through the shared gateway; the three samples took 7.5, 8.5 and 11.4 s uncached in the eval harness. Real sets of 10-100 responses took a median 52 s per set during the RECAP test run, which ran 3 sets at once (above the 6-concurrent-call guideline; later runs used one set at a time).
- Manual baseline: a blind associate sub-agent estimated what a mid-level associate would bill to review and letter each set by hand: 165, 180 and 130 minutes (round 1 sets) and 170, 140 and 210 minutes (round 3 sets, a second blind associate). These are estimates, not measured human time.
- Cold users (blind sub-agents, not real lawyers): a litigation paralegal estimated 1.5-2 hours saved per set; a solo litigator about 1 hour (review and calendaring saved; the letter still needs rewriting).
Checkable properties of the sample runs (for the rehearsal kit)
- harbor-rfp: set flags
general_objectionsanduntimely(due 2026-05-01, served 2026-05-04);withholding_unstatedon RFP 2, 3, 5, 8 and 10; no flag on RFP 1, 4 and 7; every rule cite in the rules pack. - marlowe-rogs:
missing_verification(the proof of service's oath is ignored) anduntimely(due 2026-04-17 by CCP 1013(a)); nooutdated_standard; motion date 2026-06-09. - delgado-rfa-draft (ours):
rfa_improperon RFA 2, 6, 7; no flag on RFA 1, 3, 5; no letter; fix rows and a reminder to sign. - Every model call has a receipt; the signed record verifies at /record/verify.
- Letter: no rule id outside the pack; no quote containing PDF header text.
Own model (page 72, M8)
Not needed now. The model's judgment on split responses is on par with a blind frontier judge for recall (0.95 vs 0.94 on 90 held-out responses) and costs under a tenth of a cent per response. The failure that matters is splitting real PDFs (75-79% held-out coverage), which a response-segmentation model (headings, request/response boundaries, OCR margin noise) trained on RECAP exhibits could fix; flagged as a capability opportunity.
Limits
- Labels: motion labels undercount (lower-bound precision); the synthetic key is one author's convention; the adjudicator and reviewers are the same model family (Claude Opus 5.5), not practising lawyers.
- Federal sets dominate the real data (all RECAP sets are federal courts); California and Texas are tested on synthetic sets only.
- RECAP: dockets split into several sets are correlated, and 16 test sets share a docket with dev; the docket-disjoint figures are the ones to quote. About a third of the sets are rebuilt from motion quotes (disputed items only).
- Rule text, not case law, local rules or standing orders. State deadlines: California uses CCP 135 holidays; Texas uses the federal holiday list.
Verdict
Would a buyer pay? For the triage, maybe; for the letter, no.
- What works: on text that splits cleanly, the per-response triage is good and cheap. Held out: deficient-or-not precision 0.897 / recall 0.883 on a blind synthetic set, recall 0.85-0.90 on real split responses (0.849 docket-disjoint), recall on par with a blind frontier judge, $0.0009 per response, seconds per set. Deadlines (including California court holidays and the 45-day clock), verification and waiver logic, and verbatim rule text are things a chat assistant does not give a lawyer.
- What does not: (1) the splitter found only 75-79% of responses in held-out real PDFs (75.3% on docket-disjoint sets); (2) the drafted letter lost all 9 blind comparisons with a hand-drafted letter (scores 2-5 vs 8-9), because it lists template points instead of arguing each request; (3) per-category precision on real sets is middling (about 0.6 against motions, 56-70% of other flags judged correct by a blind adjudicator before the pre-2015 fix).
- Cold users (sub-agent personas): both "maybe". The paralegal saw 1.5-2 hours saved per set and a firm price of $750-1,500 a month, but only for a hosted tier cleared for client material; the solo saw about an hour saved and would pay $49-79 a month or $10-15 a set for a zero-retention hosted version. Neither would self-host a GPU box.
Recommendation: keep it on the site as a triage tool, and rework before promoting it. In order: (a) sell it as "triage, deadlines and the deficiency list", with the letter as an export of that list, not a drafted letter; (b) make the splitter robust on real PDFs (a model fallback when headings are missed, or our own segmentation model trained on RECAP exhibits); (c) run it on the Confidential tier (single Qwen3.8 call type, like the tools proven there on 28 Sep), which is the purchase trigger both personas named; (d) only then invest in request-specific letter argument (a model-written paragraph per request, checked against the response).