{"schema_version":"1","site":"https://decosa.ai","id":"discovery-deficiency","num":"100","name":"Check their discovery responses","tool_name":"Check their discovery responses","short":"Discovery responses","blurb":"For litigation associates and paralegals on either side of written discovery. Paste or upload the other side's interrogatory answers or responses to requests for production or admission. It flags objections with no specifics, incorporated general objections, responses that don't say whether anything is being withheld, promises to produce with no date, evasive or incomplete answers, admission responses that neither admit nor deny, privilege claims with no log, a missing verification or signature, and late service, with the deadline worked out. Each flag quotes the response and cites the rule text (federal courts, California, Texas). Then it turns the flags into a first-draft deficiency letter with the rule for each point: a starting point to rewrite, not a letter ready to send. Run it on your own draft responses before you serve them: it suggests fixes and never removes an objection.","status":"live","labels":{"industry":["legal"],"job":["review","draft"],"input":["text","files"],"deploy":["hosted","selfhost"],"status":"live","output":["text","record"],"data":["privileged","pii"],"hardware":"gpu-96","licence":"permissive"},"industries":["legal"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["document-reader","signed-record"],"models":"Qwen3.8-27B (one structured read of each response) · code (splitting, rule cites, deadlines, the letter) · document reader for scanned PDFs","where":"Hosted for invented or public responses; self-host for client material","hardware":"1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; splitting, rule cites, deadlines, the letter and the record run on CPU","final_artifact":"Flags per request (quote, what is missing, rule text), the response and motion deadlines, a meet-and-confer letter as Word (or a pre-service fix list), and a signed decosa.record.v1 of the flags and receipt ids.","self_host_first":true,"verification":{"receipt_coverage":"full","summary":"Receipt per model call; every quote found in the response at a character offset; every rule cite from a rules pack of verbatim rule text (read 28 Sep 2026); signed hash-chained record","manual_qa":{"hosted":{"date":"2026-09-28","result":"pass","p50_ms":4756,"p95_ms":13634,"runs":5,"receipts_per_run":10,"cost_per_run_usd":0.0059},"selfhost":{"date":"2026-09-28","result":"pass","method":"fresh clone into a clean directory, api image from docker/api/Dockerfile, compose with a named volume, direct route to the local Qwen3.8-27B, local signing; torn down after","notes":"The rehearsal bundle passed 11/11 in 2.8 s; the three samples ran in 2.4-3.6 s with attested receipts."},"known_limits":["The splitter finds about 3 in 4 responses in real court-filed PDFs it has not seen (75-79%); missing ones are listed as warnings.","The letter draft is a starting point, not ready to send: a blind partner review preferred hand-drafted letters on every set (9 of 9), citing missing request-specific argument.","Rule text only: no case law, local rules or standing orders. California uses the CCP 135 court calendar; Texas deadlines use the federal holiday list.","Hosted numbers are from the pre-release server before merge; production numbers follow the nightly check."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Deficient responses caught on unseen real cases (RECAP test, dockets never seen in dev)","value":"174 / 235 (74%)","unit":null,"n":235,"split":"test","note":"Responses a motion to compel called deficient; 85% (174 / 205) of those the splitter found. All 43 test sets: 292 / 373 (78%)."},{"name":"Responses found by the splitter on unseen real filings","value":"75.3%","unit":null,"n":928,"split":"test","note":"27 docket-disjoint RECAP test sets; 78.8% on all 43; 99.3% on the 25 dev filings it was tuned on. Missed responses are listed as warnings."},{"name":"False flags on clean responses, blind synthetic test","value":"13 / 197 (6.6%)","unit":null,"n":197,"split":"heldout","note":"Clean responses flagged as deficient. On real filings a blind review of 80 flags the motions did not raise found 45 correct, 11 debatable, 24 wrong (15 were pre-2015 responses, since fixed)."},{"name":"Deficient-or-not precision, blind synthetic test","value":"0.897","unit":null,"n":325,"split":"heldout","note":"113 correct flags, 13 false alarms, 15 missed, 184 clean left alone; sets written by a separate blind author."},{"name":"Deficient-or-not recall, blind synthetic test","value":"0.883","unit":null,"n":325,"split":"heldout","note":null},{"name":"Per-category F1, blind synthetic test","value":"0.792","unit":null,"n":325,"split":"heldout","note":"precision 0.715, recall 0.886 before post-test fixes (0.832 after, no longer held out)."},{"name":"Recall vs blind Claude Opus 5.5, 90 responses","value":"0.952 vs 0.935","unit":null,"n":90,"split":"heldout","note":"precision 0.756 vs 0.879"},{"name":"Blind partner review: tool letter vs hand-drafted, 9 comparisons","value":"0 / 9 preferred (scores 2-5 vs 8-9)","unit":null,"n":9,"split":"test","note":"Three rounds, two of them after fixes; reviewers penalised template points instead of request-specific argument."}],"dataset":"68 sets of real written discovery responses from CourtListener RECAP (46 federal dockets; 25 dev, 43 test by hash, 27 of them on dockets with no dev set) labelled from the motions to compel; 24 synthetic sets (325 responses) written by a separate blind author; 3 synthetic demo sets (dev).","held_out":true,"caveats":["Motion labels undercount what is wrong, so precision against them is a lower bound; a blind adjudicator judged a sample of the other flags.","The adjudicator, frontier judge, letter reviewer and cold users are Claude Opus 5.5 sub-agents, not practising lawyers.","All real sets are federal; California and Texas are measured on synthetic sets only.","About a third of the real sets were rebuilt from quotes in the motion (disputed items only), and some dockets were split into correlated sets; the docket-disjoint figures are the stricter ones.","Fixes made after the held-out runs are reported separately and not counted as held out.","The drafted letter lost every blind comparison with a hand-drafted letter."],"date":"2026-09-28","doc_url":"https://decosa.ai/metrics/evals/discovery-deficiency"},"stack":{"summary":"For litigation associates and paralegals on either side of written discovery. It splits a set of interrogatory answers or responses to requests for production or admission into its numbered responses, reads each one with an open model, and flags what the rules require and the response lacks: specific objections, a statement of whether anything is withheld, a production date, a complete answer, a proper admission or denial, a privilege log, a verification, a signature, timely service. Each flag quotes the response and cites the rule text for federal court, California or Texas; the rule cites come from a pack of verbatim rule text, never from the model. It then turns the flags into a first-draft deficiency letter: a starting point with the rule for each point, not a letter ready to send. On your own draft responses it lists fixes before service and never removes an objection.","tagline":"The other side's written discovery responses in; every objection without specifics, unstated withholding, evasive answer, missing verification and late service out, quoted and cited to the rule, with the deadlines worked out and a first-draft letter to rewrite.","deployment":"hosted-or-self-host","regulatory_note":"Checked 28 Sep 2026. A review aid for lawyers and litigation staff, not legal advice and not for people without a lawyer: an attorney decides which deficiencies to raise and signs the letter (Fed. R. Civ. P. 26(g)(1) makes the signer certify discovery papers). Rules cited: Fed. R. Civ. P. 26(b)(1), 26(b)(5)(A), 26(g), 33(b), 33(d), 34(b)(2)(A)-(C), 36(a)(3)-(6), 37(a)(1), (a)(4), 6(a), 6(d) and the 2015 Advisory Committee Notes to Rules 26 and 34 (https://www.law.cornell.edu/rules/frcp); Cal. Code Civ. Proc. 2016.040, 2017.010, 2030.210-2030.300, 2031.210-2031.310, 2033.210-2033.280, 1010.6, 1013 (https://leginfo.legislature.ca.gov); Tex. R. Civ. P. 191-198 and 21a (https://www.txcourts.gov/rules-forms/rules-standards/). Every quote in the rules pack was matched word for word against those pages on 28 Sep 2026; the Texas rules came from the official PDF, which carries no single effective-date stamp. It does not apply case law (courts differ on whether general or boilerplate objections are waived), local rules or a judge's standing order; state deadlines are counted on the US federal holiday list, which can differ from a state court's. Confidentiality: discovery responses carry client facts and are often under a protective order; ABA Formal Opinion 512 (29 Jul 2024) asks lawyers to understand how a tool uses what they put in and to protect client information, so run real matters self-hosted, where nothing leaves the machine. The hosted demo takes invented or public responses only. Model licence: Apache-2.0 (Qwen3.8-27B).","components":[{"id":"discovery","role":"Splitter, set checks (verification, signature, deadlines), flag rules, quote location, rules pack, letter and fix list, signed record (no model; CPU)","name":"decosa-api discovery check (decosa_api/verticals/discovery), with the dates block, the drafting editor's Word writer and the signed record (07)","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12; rules pack of 100 verbatim quotes (FRCP, CCP, TRCP) read 28 Sep 2026; poppler pdftotext for PDFs","receipt_coverage":"partial","in_hosted_demo":null,"tiers":[],"alternative_to":null},{"id":"model","role":"Reads each response once and describes it as JSON: objection grounds and whether each gives specifics, withholding statement, production date, answer shape, admission shape","name":"Qwen3.8-27B (NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP NVFP4, GDN/attention FP8) + FP8 KV cache; MTP head, 3 draft tokens","vram_gb":20,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, temperature 0, thinking off","receipt_coverage":"strong","in_hosted_demo":null,"tiers":[],"alternative_to":null},{"id":"reader","role":"Scanned PDFs only: page images to text","name":"Document reader (Docling layout heron + PaddleOCR-VL-1.6)","hf_repo":"PaddlePaddle/PaddleOCR-VL-1.6","license":"Apache-2.0","params":"0.9B","quant":"BF16","vram_gb":6,"memory_gb_estimate":null,"engine":"vLLM 0.29 (parser) + Docling 2.130 service","receipt_coverage":"partial","in_hosted_demo":true,"tiers":[],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · one 32 GB card","summary":"Text, Word and text-layer PDFs only (no scans): the model on one consumer card, everything else on CPU.","components":["discovery","model"],"hardware":"1x RTX 5090 32 GB (estimate)","quality_evidence":[{"metric":"Same model and prompts as standard","value":"not measured separately","source":"estimate: identical pipeline without the document reader"}],"latency_note":"not measured","in_hosted_demo":false,"receipt_coverage":"strong","receipt_note":"Every model call is receipted by the instance key.","hosting":null},{"id":"standard","label":"Standard · Qwen3.8-27B and the document reader (hosted demo)","summary":"What the hosted demo runs: Qwen3.8-27B reads each response, code assigns flags, cites rules, computes deadlines and writes the letter; the document reader handles scans.","components":["discovery","model","reader"],"hardware":"1x RTX PRO 6000 96 GB (measured on shared cards)","quality_evidence":[{"metric":"Responses correctly called deficient or not, blind synthetic test (24 sets by another author, 325 responses, federal, California, Texas)","value":"precision 0.897, recall 0.883; 184 clean responses left alone","source":"decosa-api docs/evals/discovery-deficiency.md, held out, run once"},{"metric":"Real responses a motion to compel called deficient, caught (27 docket-disjoint RECAP test sets)","value":"174 / 235 (74%); 174 / 205 (85%) of the responses the splitter found","source":"decosa-api docs/evals/discovery-deficiency.md"},{"metric":"Clean responses wrongly flagged, blind synthetic test","value":"13 / 197 (6.6%)","source":"decosa-api docs/evals/discovery-deficiency.md, held out"},{"metric":"Splitter coverage on real filings (held out)","value":"75.3% of responses found on docket-disjoint sets (78.8% on all 43; 99.3% on the dev filings it was tuned on)","source":"decosa-api docs/evals/discovery-deficiency.md"},{"metric":"Against a blind frontier judge (Claude Opus 5.5), 90 held-out responses","value":"recall 0.952 vs 0.935; precision 0.756 vs 0.879","source":"decosa-api docs/evals/discovery-deficiency.md"},{"metric":"Draft letter vs a hand-drafted letter, blind partner review","value":"a starting point, not ready to send: it lost 9 of 9 blind comparisons with a hand-drafted letter (2-5 vs 8-9 of 10)","source":"decosa-api docs/evals/discovery-deficiency.md"}],"latency_note":"measured on our pre-release server through the shared gateway: seconds for the sample","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Every model call is receipted.","hosting":null}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"GET /discovery/info, /discovery/samples; POST /discovery/check (SSE or JSON), /discovery/letter.docx; POST /record/verify. Keeps no response text."},{"name":"decosa-llm","port":8000,"image":"${DECOSA_REGISTRY}/decosa-llm:0.1.0","purpose":"vLLM OpenAI endpoint for Qwen3.8-27B."},{"name":"decosa-docreader","port":8497,"image":"built from services/docreader (no published image yet)","purpose":"Scanned PDFs only."}],"tools":[{"name":"Rules pack (FRCP, California CCP, Texas TRCP)","url":"https://www.law.cornell.edu/rules/frcp","license":"US federal rules: public domain; California statutes: public; Texas rules: published by the Supreme Court of Texas","purpose":"The rule text each flag cites, verbatim and linked; read 28 Sep 2026."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured on our server: Qwen3.8-27B NVFP4 (about 20 GB of weights) with room for the document reader (about 6 GB)."},{"tier":"1x RTX 5090 32 GB","fits":true,"notes":"Estimate: Qwen3.8-27B NVFP4 with a modest KV cache; responses are short, so a 16k context is enough. Not measured."},{"tier":"CPU only","fits":false,"notes":"The model needs a GPU. Splitting, deadlines, rule cites, the letter and the record run on CPU."}],"latency":[{"lane":"the ten-response federal RFP sample, hosted gateway route","typical_ms":4756,"source":"measured on our server 2026-09-28 (pre-release server, shared gateway), p50 of 5 runs"},{"lane":"real response sets of 10-100 responses (RECAP), hosted gateway route","typical_ms":51900,"source":"measured on our server 2026-09-28, median over 43 sets, three sets running at once"},{"lane":"the California rehearsal set, self-hosted direct route to the local Qwen3.8","typical_ms":2800,"source":"measured on our server 2026-09-28 (fresh clone, compose, direct route)"}],"benchmark":null,"notes":["Measured on real public filings (CourtListener RECAP exhibits to motions to compel, labelled from the motion) and on synthetic sets written by a separate blind author; see the eval for splits and caveats.","The splitter is the weak point on real PDFs: on held-out filings from dockets it never saw it found 75.3% of responses. Missing responses are listed as warnings so you can see them.","The drafted letter is a structured first draft of the deficiency list; a blind partner review preferred hand-drafted letters on every set tested."]},"buyer_facts":[{"label":"Data retention","value":"Nothing stored: the responses live in memory for the request. The signed record holds the document hash, each flag's category, place and rule ids and the receipt ids, never the response text; logs carry counts and timings only."},{"label":"What leaves the box","value":"Self-hosted on the direct route: nothing. The model and the checks run on the same machine. Hosted: model calls go through the Decosa API, and only invented or public responses are accepted."},{"label":"What it checks","value":"Per response: objections without specifics, incorporated general objections, no statement of whether anything is withheld, no production date, evasive or incomplete answers (including \"see documents produced\"), admission responses that neither admit nor deny, privilege claims with no log, the pre-2015 \"reasonably calculated\" test (federal only). Per set: general objections, verification, signature, late service."},{"label":"What it is not","value":"Legal advice, a motion or a filing. It cites rule text, not case law, local rules or standing orders, and does not judge whether the requests themselves were proper. An attorney decides and signs."},{"label":"Time and cost per response set","value":"Model cost at list price: a fraction of a cent per response (one Qwen3.8 call per response; the letter and deadlines are code). The sample runs in seconds; real sets took about a minute on a shared gateway. A blind associate estimated hours to review and letter a set by hand; the letter draft is a starting point, not ready to send (see the eval)."},{"label":"Confidential tier","value":"Not yet run in the attested enclave. Every model call here is the same kind of single Qwen3.8 chat call that other tools ran sealed on 28 Sep 2026; the proven setup is your own self-hosted API with its model calls sealed to the enclave. Scanned PDFs also need the document reader, which runs in the enclave's capability build. Until a sealed run is recorded for this tool, use self-hosting for client material."}],"data_handling":{"page":"/data#discovery-deficiency","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Nothing stored: the responses live in memory for the request. The signed record holds the document hash, each flag's category, place and rule ids and the receipt ids, never the response text; logs carry counts and timings only.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/legal/discovery-deficiency","input":"discovery","lanes":[{"id":"grid","title":"Requests by defect","kind":"list"},{"id":"flags","title":"Flags with quotes and rules","kind":"list"},{"id":"deadlines","title":"Deadlines","kind":"list"},{"id":"letter","title":"Meet-and-confer letter","kind":"markdown"}],"samples":[{"n":1,"id":"harbor-rfp","title":"Their RFP responses (federal, N.D. Ill.)","deep_link":"/legal/discovery-deficiency?sample=1&autorun=0"},{"n":2,"id":"marlowe-rogs","title":"Their special interrogatory answers (California)","deep_link":"/legal/discovery-deficiency?sample=2&autorun=0"},{"n":3,"id":"delgado-rfa-draft","title":"Our draft RFA responses (pre-service, W.D. Tex.)","deep_link":"/legal/discovery-deficiency?sample=3&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/discovery-deficiency-hosted.md","selfhost":"/prompts/discovery-deficiency-selfhost.md","assemble":"/prompts/discovery-deficiency-assemble.md","mac":null},"rehearsal":{"bundle":"/samples/discovery-deficiency.zip","bundle_url":"https://decosa.ai/samples/discovery-deficiency.zip","folder":"/samples/discovery-deficiency/","expected":"/samples/discovery-deficiency/expected.json","files":["/samples/discovery-deficiency/expected.json","/samples/discovery-deficiency/inputs/marlowe-rogs.txt"],"bytes":3605,"checks":["the set is deficient","late service is flagged (due 17 Apr, served 20 Apr)","the missing verification is flagged (the proof of service oath does not count)","the \"will supplement\" answer to interrogatory 5 is flagged as evasive","\"See documents produced\" in answer 8 is flagged","answers 3 and 7 are not flagged","no pre-2015 standard flag in California","the 45-day motion date is worked out","a letter is drafted","every model call has a signed receipt","the signed record verifies"],"licence":"Written for this bundle (CC0); every party, lawyer, case number and fact is invented. Part of decosa-api, AGPL-3.0-or-later.","about":"Synthetic responses on California pleading paper (every party, lawyer and fact invented). The requests were served by mail on 13 Mar 2026, so the answers were due 17 Apr (30 days + 5 for mail within California, CCP 1013(a)); they were served 20 Apr, 3 days late, which waives the objections (CCP 2030.290(a)). There is no verification: the oath in the proof of service is the server's, not the party's. Answers 2 and 5 are evasive, answer 8 says \"See documents produced\", answers 3 and 7 are fine. California still uses the \"reasonably calculated\" test (CCP 2017.010), so it must not be flagged as outdated. The check must flag all of this with quotes and rule cites, work out the 45-day motion date, draft a letter, attach a receipt to every model call and sign a record that verifies.","run":{"containers":"docker compose exec api python scripts/rehearse.py discovery-deficiency","checkout":"python scripts/rehearse.py discovery-deficiency --bundle discovery-deficiency.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py discovery-deficiency"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=discovery-deficiency","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":63.6,"basis":"stack","unknown":[]}],"mac":null},"links":{"page":"/legal/discovery-deficiency","json":"/use-cases/discovery-deficiency.json","metrics":"/metrics/discovery-deficiency","console":"/legal/discovery-deficiency","console_sample":"/legal/discovery-deficiency?sample=1&autorun=0","stack":"/legal/discovery-deficiency#stack","try_live":"/legal/discovery-deficiency","watch":"/legal/discovery-deficiency","build":"/legal/discovery-deficiency#build","self_host":"/legal/discovery-deficiency#self-host","prompts":{"hosted":"/prompts/discovery-deficiency-hosted.md","selfhost":"/prompts/discovery-deficiency-selfhost.md","assemble":"/prompts/discovery-deficiency-assemble.md","mac":null}}}