Check their discovery responses
Every objection without specifics, unstated withholding, evasive answer and late or unverified set, quoted and cited to the rule, with the deadlines worked out and a first-draft letter to rewrite.
Built on: Document reader, Signed record
Deficiencies
LivePick a response set and check it. Each numbered response is read for objections without specifics, unstated withholding, missing production dates, evasive answers, improper admission responses and privilege claims without a log; the set for general objections, a missing verification or signature, and late service.
Every flag quotes the response and cites the rule text. Their responses get a meet-and-confer letter; your own draft gets a fix list. An attorney decides.
Watch a recorded run first
Watch: their responses checked and the letter drafted
Replay · not liveLoading the recorded runs…
Pick a response set and check it. Each numbered response is read for objections without specifics, unstated withholding, missing production dates, evasive answers, improper admission responses and privilege claims without a log; the set for general objections, a missing verification or signature, and late service.
Every flag quotes the response and cites the rule text. Their responses get a meet-and-confer letter; your own draft gets a fix list. An attorney decides.
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; splitting, rule cites, deadlines, the letter and the record run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the check their discovery responses API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real client material belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- discovery-deficiency
Use the hosted API
# Decosa discovery-response check: use the hosted API
You are wiring Decosa's discovery-response check into this project (a litigation support tool, a document workflow or a
script). It takes a set of written discovery responses (interrogatory answers, responses to requests for production or
admission; text, PDF or Word), splits it into its numbered responses and returns flags: objections without specifics,
incorporated general objections, no statement of whether anything is withheld, no production date, evasive or incomplete
answers, admission responses that neither admit nor deny, privilege claims with no log, and for the set general
objections, a missing verification or signature and late service. Each flag quotes the response and cites the rule text
(federal, California or Texas). For the other side's responses it writes a meet-and-confer letter; for your own draft, a
fix list. Every model call has a signed receipt and the run ends in a signed record. Use only what is listed below; if
you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes invented or public responses only** (`"synthetic": true` is required), for example exhibits
already filed on a public docket. Client material belongs on the self-hosted version.
- It is a review aid for lawyers, not legal advice: an attorney decides what to raise and signs the letter.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool's page. Keep it in `DECOSA_API_KEY`, never in code.
Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "discovery-deficiency"}` returns `{"token", "expires_at", "budget"}`.
Demo sessions are limited per IP per hour and carry a generated-token budget: 429 with `Retry-After` over a limit, 402
when the budget left is too small, 409 when a demo token already has a run going.
## Check a response set
- `POST /discovery/check` (token). Body, one of:
- `{"text": "...", "synthetic": true, "jurisdiction": "federal", "mode": "theirs", "requests_served_on": "2026-05-01",
"requests_service": "email", "privilege_log": "not_served", "parties": {"responding_party": "...", "our_client": "..."}}`
(or `"file_b64"` with a PDF, .docx or .txt and `"filename"` instead of `"text"`; `jurisdiction`: federal, california,
texas; `mode`: theirs or ours; `requests_service`: email, electronic, mail, hand, overnight; `privilege_log`: served,
not_served, unknown; `responses_served_on` overrides the date on the certificate of service);
- `{"sample_id": "harbor-rfp"}` (bundled invented sets: also `marlowe-rogs`, `delgado-rfa-draft`).
- Keep the request headings in the text ("REQUEST FOR PRODUCTION NO. 3:" / "RESPONSE:", "INTERROGATORY NO. 3" / "ANSWER:").
- Limits: 400,000 characters, 250 responses, 12 MB per file.
- The JSON response has `status` (deficient, check or no_flags), `grid` (each request with the flag ids per category),
`flags` (each with `category`, `label`, `severity`, `item`, `quote`, `span`, `why`, `found_by` and `rules`: id, URL and
verbatim rule text), `deadline` and `motion_deadline` (with the counting steps), `letter` (`paragraphs`, `text`,
`reply_by`) or `fixes`, `counts`, `warnings`, `record` and `receipt_events`.
- With `Accept: text/event-stream` (or `"stream": true`): `reading`/`read` for a file, `parsed`, `set`, a `receipt` per
model call, an `item` event per response with its flags, `result`, `budget` and `done`.
- `POST /discovery/letter.docx` (token) `{"paragraphs": [...]}` returns the edited letter as a Word file (no model call).
- `POST /record/verify` (no token) `{"record": {...}}` re-checks the signed record.
- `GET /discovery/info` (categories, rules pack sources and read date, limits), `GET /discovery/samples` and
`GET /attest/signing-key` need no token.
## Example (Python, `pip install httpx`)
```python
import base64, httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
f = pathlib.Path("responses.pdf")
body = {"file_b64": base64.b64encode(f.read_bytes()).decode(), "filename": f.name, "synthetic": True,
"jurisdiction": "federal", "mode": "theirs", "requests_served_on": "2026-05-01", "requests_service": "email"}
r = httpx.post(f"{API}/discovery/check", json=body, headers=H, timeout=600)
r.raise_for_status()
js = r.json()
for fl in js["flags"]:
print(fl["item"] or "set", fl["label"], "|", fl["quote"], "|", fl["rules"][0]["id"] if fl["rules"] else "")
pathlib.Path("letter.txt").write_text(js["letter"]["text"] if js["letter"] else "")
```
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa discovery-response check: run it yourself (containers)
You are setting up Decosa's discovery-response check on this machine, so client material never leaves it. It splits a set
of written discovery responses into its numbered responses, has an open model read each one, flags what the rules require
and the response lacks (each flag quoted and cited to the rule text for federal court, California or Texas), computes the
deadlines, writes the meet-and-confer letter (or a fix list for your own draft) and signs a record. Nothing is sent to
Decosa's hosted API. It is a review aid for lawyers: an attorney decides and signs.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/discovery-deficiency.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py discovery-deficiency` (the api image carries the same bundle under /app/rehearsal/discovery-deficiency/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py discovery-deficiency --bundle discovery-deficiency.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the set is deficient", "late service is flagged (due 17 Apr, served 20 Apr)", "the missing verification is flagged (the proof of service oath does not count)"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and
`DECOSA_DISCOVERY_SYNTHETIC_ONLY=0` (so this box accepts client material), and bind every port to 127.0.0.1. Never set
the gateway route on a box that holds client material.
3. Scanned PDFs (optional): add the document reader (`services/docreader` plus PaddleOCR-VL-1.6 on vLLM, about 6 GB)
and set `DECOSA_DOCREADER_URL`. Without it, send text, Word files or PDFs with a text layer.
4. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check.
5. Check: `curl -fsS http://127.0.0.1:<PORT>/discovery/info` shows `synthetic_only: false`; `GET /attest/signing-key`
shows this box's public key.
6. Smoke test: get a token with `POST /demo/session {"vertical":"discovery-deficiency"}` and send
`{"sample_id": "harbor-rfp"}` to `POST /discovery/check`. Expect `status: deficient`, `untimely` and
`general_objections` for the set, `withholding_unstated` on requests 2, 3 and 10, and no flags on requests 1, 4 and 7.
Send the record to `POST /record/verify`: `ok` must be true.
7. Report back: the public key, the flag counts, and how long the run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds client
material. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVFP4) needs a GPU.
- GeForce RTX 4090lite tierRuns with a smaller tier
The standard tier does not fit: Needs about 26 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes.
- GeForce RTX 5090lite tierRuns with a smaller tier
The standard tier does not fit: Needs about 34 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (63.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (63.6 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBlite tierRuns with a smaller tier
The standard tier can't be checked: Document reader (Docling layout heron + PaddleOCR-VL-1.6) has no mapped Apple Silicon build The lite tier fits with changes.
- Apple M5 Max, 64 GBlite tierRuns with a smaller tier
The standard tier can't be checked: Document reader (Docling layout heron + PaddleOCR-VL-1.6) has no mapped Apple Silicon build The lite tier fits with changes.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"discovery-deficiency"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py discovery-deficiency
Download the mock-data bundle (4 KB, 11 checks)expected.json
Synthetic responses on California pleading paper (every party, lawyer and fact invented). The requests were served by mail on 13 Mar 2026, so the answers were due 17 Apr (30 days + 5 for mail within California, CCP 1013(a)); they were served 20 Apr, 3 days late, which waives the objections (CCP 2030.290(a)). There is no verification: the oath in the proof of service is the server's, not the party's. Answers 2 and 5 are evasive, answer 8 says "See documents produced", answers 3 and 7 are fine. California still uses the "reasonably calculated" test (CCP 2017.010), so it must not be flagged as outdated. The check must flag all of this with quotes and rule cites, work out the 45-day motion date, draft a letter, attach a receipt to every model call and sign a record that verifies.
What the rehearsal checks
- the set is deficient
- late service is flagged (due 17 Apr, served 20 Apr)
- the missing verification is flagged (the proof of service oath does not count)
- the "will supplement" answer to interrogatory 5 is flagged as evasive
- "See documents produced" in answer 8 is flagged
- answers 3 and 7 are not flagged
- no pre-2015 standard flag in California
- the 45-day motion date is worked out
- a letter is drafted
- every model call has a signed receipt
- the signed record verifies
Licence: Written for this bundle (CC0); every party, lawyer, case number and fact is invented. Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa discovery-response check: run it yourself (containers)
You are setting up Decosa's discovery-response check on this machine, so client material never leaves it. It splits a set
of written discovery responses into its numbered responses, has an open model read each one, flags what the rules require
and the response lacks (each flag quoted and cited to the rule text for federal court, California or Texas), computes the
deadlines, writes the meet-and-confer letter (or a fix list for your own draft) and signs a record. Nothing is sent to
Decosa's hosted API. It is a review aid for lawyers: an attorney decides and signs.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/discovery-deficiency.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py discovery-deficiency` (the api image carries the same bundle under /app/rehearsal/discovery-deficiency/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py discovery-deficiency --bundle discovery-deficiency.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the set is deficient", "late service is flagged (due 17 Apr, served 20 Apr)", "the missing verification is flagged (the proof of service oath does not count)"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and
`DECOSA_DISCOVERY_SYNTHETIC_ONLY=0` (so this box accepts client material), and bind every port to 127.0.0.1. Never set
the gateway route on a box that holds client material.
3. Scanned PDFs (optional): add the document reader (`services/docreader` plus PaddleOCR-VL-1.6 on vLLM, about 6 GB)
and set `DECOSA_DOCREADER_URL`. Without it, send text, Word files or PDFs with a text layer.
4. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check.
5. Check: `curl -fsS http://127.0.0.1:<PORT>/discovery/info` shows `synthetic_only: false`; `GET /attest/signing-key`
shows this box's public key.
6. Smoke test: get a token with `POST /demo/session {"vertical":"discovery-deficiency"}` and send
`{"sample_id": "harbor-rfp"}` to `POST /discovery/check`. Expect `status: deficient`, `untimely` and
`general_objections` for the set, `withholding_unstated` on requests 2, 3 and 10, and no flags on requests 1, 4 and 7.
Send the record to `POST /record/verify`: `ok` must be true.
7. Report back: the public key, the flag counts, and how long the run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds client
material. If I ask for it later, follow the Provide page instead of improvising.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Runs with a smaller tierCheck their discovery responses on GeForce RTX 5090: use the Lite · one 32 GB card tier
The standard tier does not fit: Needs about 34 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.
What this tool's stack says about this hardware:
- 1x RTX 5090 32 GB (fits): Estimate: Qwen3.8-27B NVFP4 with a modest KV cache; responses are short, so a 16k context is enough. Not measured.
Lite · one 32 GB card: what changesuses estimates
- Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Splitter, set checks: decosa-api discovery check (decosa_api/verticals/discovery), with the dates block, the drafting editor's Word writer and the signed record (07). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Reads each response once and describes it as...: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Check their discovery responses, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Check their discovery responses on my hardware Fetch https://decosa.ai/prompts/discovery-deficiency-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=discovery-deficiency) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Lite · one 32 GB card (lite). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Splitter, set checks: decosa-api discovery check (decosa_api/verticals/discovery), with the dates block, the drafting editor's Word writer and the signed record (07), CPU - Reads each response once and describes it as...: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/discovery-deficiency-assemble.md
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 28 Sep 2026 · measured 28 Sep 2026: · p50 4.8 s · p95 14 s (5 runs) · ~$0.006 per run · 10 receipts
Loading the nightly status…
Self-host: verified 28 Sep 2026 · fresh clone into a clean directory, api image from docker/api/Dockerfile, compose with a named volume, direct route to the local Qwen3.8-27B, local signing; torn down after
Measured cost to run: about $0.087 per 100 responses (hosted, 28 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The rehearsal bundle passed 11/11 in 2.8 s; the three samples ran in 2.4-3.6 s with attested receipts.
Known limits (4)
- The splitter finds about 3 in 4 responses in real court-filed PDFs it has not seen (75-79%); missing ones are listed as warnings.
- The letter draft is a starting point, not ready to send: a blind partner review preferred hand-drafted letters on every set (9 of 9), citing missing request-specific argument.
- Rule text only: no case law, local rules or standing orders. California uses the CCP 135 court calendar; Texas deadlines use the federal holiday list.
- Hosted numbers are from the pre-release server before merge; production numbers follow the nightly check.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; splitting, rule cites, deadlines, the letter and the record run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the check their discovery responses API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real client material belongs on your own hardware.
The other side's written discovery responses in; every objection without specifics, unstated withholding, evasive answer, missing verification and late service out, quoted and cited to the rule, with the deadlines worked out and a first-draft letter to rewrite.
For litigation associates and paralegals on either side of written discovery. It splits a set of interrogatory answers or responses to requests for production or admission into its numbered responses, reads each one with an open model, and flags what the rules require and the response lacks: specific objections, a statement of whether anything is withheld, a production date, a complete answer, a proper admission or denial, a privilege log, a verification, a signature, timely service. Each flag quotes the response and cites the rule text for federal court, California or Texas; the rule cites come from a pack of verbatim rule text, never from the model. It then turns the flags into a first-draft deficiency letter: a starting point with the rule for each point, not a letter ready to send. On your own draft responses it lists fixes before service and never removes an objection.
- Deployment
- Hosted or self-host
- Regulatory
- Checked 28 Sep 2026. A review aid for lawyers and litigation staff, not legal advice and not for people without a lawyer: an attorney decides which deficiencies to raise and signs the letter (Fed. R. Civ. P. 26(g)(1) makes the signer certify discovery papers). Rules cited: Fed. R. Civ. P. 26(b)(1), 26(b)(5)(A), 26(g), 33(b), 33(d), 34(b)(2)(A)-(C), 36(a)(3)-(6), 37(a)(1), (a)(4), 6(a), 6(d) and the 2015 Advisory Committee Notes to Rules 26 and 34 (https://www.law.cornell.edu/rules/frcp); Cal. Code Civ. Proc. 2016.040, 2017.010, 2030.210-2030.300, 2031.210-2031.310, 2033.210-2033.280, 1010.6, 1013 (https://leginfo.legislature.ca.gov); Tex. R. Civ. P. 191-198 and 21a (https://www.txcourts.gov/rules-forms/rules-standards/). Every quote in the rules pack was matched word for word against those pages on 28 Sep 2026; the Texas rules came from the official PDF, which carries no single effective-date stamp. It does not apply case law (courts differ on whether general or boilerplate objections are waived), local rules or a judge's standing order; state deadlines are counted on the US federal holiday list, which can differ from a state court's. Confidentiality: discovery responses carry client facts and are often under a protective order; ABA Formal Opinion 512 (29 Jul 2024) asks lawyers to understand how a tool uses what they put in and to protect client information, so run real matters self-hosted, where nothing leaves the machine. The hosted demo takes invented or public responses only. Model licence: Apache-2.0 (Qwen3.8-27B).
Text description
A set of written discovery responses goes to decosa-api, which splits it into numbered requests and responses in code, checks the set (general objections, verification, signature, timeliness), has Qwen3.8-27B describe each response as JSON, assigns flags and rule cites in code from a verified rules pack, writes the meet-and-confer letter or fix list in code, and signs a record. Every model call gets a receipt.
At a glance
- Data retention
- Nothing stored: the responses live in memory for the request. The signed record holds the document hash, each flag's category, place and rule ids and the receipt ids, never the response text; logs carry counts and timings only.
- What leaves the box
- Self-hosted on the direct route: nothing. The model and the checks run on the same machine. Hosted: model calls go through the Decosa API, and only invented or public responses are accepted.
- What it checks
- Per response: objections without specifics, incorporated general objections, no statement of whether anything is withheld, no production date, evasive or incomplete answers (including "see documents produced"), admission responses that neither admit nor deny, privilege claims with no log, the pre-2015 "reasonably calculated" test (federal only). Per set: general objections, verification, signature, late service.
- What it is not
- Legal advice, a motion or a filing. It cites rule text, not case law, local rules or standing orders, and does not judge whether the requests themselves were proper. An attorney decides and signs.
- Time and cost per response set
- Model cost at list price: a fraction of a cent per response (one Qwen3.8 call per response; the letter and deadlines are code). The sample runs in seconds; real sets took about a minute on a shared gateway. A blind associate estimated hours to review and letter a set by hand; the letter draft is a starting point, not ready to send (see the eval).
- Confidential tier
- Not yet run in the attested enclave. Every model call here is the same kind of single Qwen3.8 chat call that other tools ran sealed on 28 Sep 2026; the proven setup is your own self-hosted API with its model calls sealed to the enclave. Scanned PDFs also need the document reader, which runs in the enclave's capability build. Until a sealed run is recorded for this tool, use self-hosting for client material.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
one 32 GB card
Text, Word and text-layer PDFs only (no scans): the model on one consumer card, everything else on CPU.
- Models
- decosa-api discovery check (decosa_api/verticals/discovery), with the dates block, the drafting editor's Word writer and the signed record (07)
- Qwen3.8-27B (NVFP4)
- Hardware
- 1x RTX 5090 32 GB (estimate)
- Quality evidence
- Same model and prompts as standardnot measured yetestimate: identical pipeline without the document reader
- Latency
- not measured
- Verification
- Proof: strongSelf-host onlyEvery model call is receipted by the instance key.
- In the hosted demo
Standard
Qwen3.8-27B and the document reader (hosted demo)
What the hosted demo runs: Qwen3.8-27B reads each response, code assigns flags, cites rules, computes deadlines and writes the letter; the document reader handles scans.
- Models
- decosa-api discovery check (decosa_api/verticals/discovery), with the dates block, the drafting editor's Word writer and the signed record (07)
- Qwen3.8-27B (NVFP4)
- Document reader (Docling layout heron + PaddleOCR-VL-1.6)
- Hardware
- 1x RTX PRO 6000 96 GB (measured on shared cards)
- Quality evidence
- Responses correctly called deficient or not, blind synthetic test (24 sets by another author, 325 responses, federal, California, Texas)precision 0.897, recall 0.883; 184 clean responses left alonedecosa-api docs/evals/discovery-deficiency.md, held out, run once
- Real responses a motion to compel called deficient, caught (27 docket-disjoint RECAP test sets)174 / 235 (74%); 174 / 205 (85%) of the responses the splitter founddecosa-api docs/evals/discovery-deficiency.md
- Clean responses wrongly flagged, blind synthetic test13 / 197 (6.6%)decosa-api docs/evals/discovery-deficiency.md, held out
- Splitter coverage on real filings (held out)75.3% of responses found on docket-disjoint sets (78.8% on all 43; 99.3% on the dev filings it was tuned on)decosa-api docs/evals/discovery-deficiency.md
- Against a blind frontier judge (Claude Opus 5.5), 90 held-out responsesrecall 0.952 vs 0.935; precision 0.756 vs 0.879decosa-api docs/evals/discovery-deficiency.md
- Draft letter vs a hand-drafted letter, blind partner reviewa starting point, not ready to send: it lost 9 of 9 blind comparisons with a hand-drafted letter (2-5 vs 8-9 of 10)decosa-api docs/evals/discovery-deficiency.md
- Latency
- measured on our pre-release server through the shared gateway: seconds for the sample
- Verification
- Proof: strongEvery model call is receipted.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Splitter, set checks (verification, signature, deadlines), flag rules, quote location, rules pack, letter and fix list, signed record (no model; CPU)decosa-api discovery check (decosa_api/verticals/discovery), with the dates block, the drafting editor's Word writer and the signed record (07) 0 GBProof: partial | LiteStandard | 0 GB | Proof: partial | |
| ||||
Reads each response once and describes it as JSON: objection grounds and whether each gives specifics, withholding statement, production date, answer shape, admission shapeQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strong | LiteStandard | 27.8B · 20 GB | Proof: strong | |
| ||||
Scanned PDFs only: page images to textDocument reader (Docling layout heron + PaddleOCR-VL-1.6)PaddlePaddle/PaddleOCR-VL-1.6 on Hugging Face (opens in a new tab) 0.9B · 6 GBProof: partialIn the hosted demo | Standard | 0.9B · 6 GB | Proof: partialIn the hosted demo | |
| ||||
Tools, services and hardware
Tools
- Rules pack (FRCP, California CCP, Texas TRCP) (opens in a new tab)US federal rules: public domain; California statutes: public; Texas rules: published by the Supreme Court of Texas
The rule text each flag cites, verbatim and linked; read 28 Sep 2026.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /discovery/info, /discovery/samples; POST /discovery/check (SSE or JSON), /discovery/letter.docx; POST /record/verify. Keeps no response text.
- decosa-llm:8000
${DECOSA_REGISTRY}/decosa-llm:0.1.0vLLM OpenAI endpoint for Qwen3.8-27B.
- decosa-docreader:8497
built from services/docreader (no published image yet)Scanned PDFs only.
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server: Qwen3.8-27B NVFP4 (about 20 GB of weights) with room for the document reader (about 6 GB).
- 1x RTX 5090 32 GB Fits
Estimate: Qwen3.8-27B NVFP4 with a modest KV cache; responses are short, so a 16k context is enough. Not measured.
- CPU only Does not fit
The model needs a GPU. Splitting, deadlines, rule cites, the letter and the record run on CPU.
Latency per lane
- the ten-response federal RFP sample, hosted gateway route4.8 s
Measuredmeasured on our server 2026-09-28 (pre-release server, shared gateway), p50 of 5 runs
- real response sets of 10-100 responses (RECAP), hosted gateway route51.9 s
Measuredmeasured on our server 2026-09-28, median over 43 sets, three sets running at once
- the California rehearsal set, self-hosted direct route to the local Qwen3.82.8 s
Measuredmeasured on our server 2026-09-28 (fresh clone, compose, direct route)
Notes
- Measured on real public filings (CourtListener RECAP exhibits to motions to compel, labelled from the motion) and on synthetic sets written by a separate blind author; see the eval for splits and caveats.
- The splitter is the weak point on real PDFs: on held-out filings from dockets it never saw it found 75.3% of responses. Missing responses are listed as warnings so you can see them.
- The drafted letter is a structured first draft of the deficiency list; a blind partner review preferred hand-drafted letters on every set tested.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa discovery-response check on this machine
You are setting up a self-hosted review aid on this Linux machine, for a litigation team. It takes a set of written
discovery responses (interrogatory answers, responses to requests for production or admission; text, PDF or Word) and
returns:
- each numbered request and response, split in code;
- flags per response, each with a quote from the response at a character offset, what is missing, and the rule text it
falls short of (Federal Rules of Civil Procedure, California Code of Civil Procedure or Texas Rules of Civil Procedure,
from a rules pack of verbatim quotes read on 28 Sep 2026): objections with no specifics, incorporated general objections,
no statement of whether anything is withheld, no production date, evasive or incomplete answers, admission responses
that neither admit nor deny, privilege claims with no log;
- set-level flags: general objections, no verification under oath, no signature, late service (the deadline is computed
from the date the requests were served and the service method); in California, the 45-day motion-to-compel date;
- for the other side's responses, a meet-and-confer letter written in code from the flags (downloadable as Word); for our
own draft, a fix list that never removes an objection;
- a signed, hash-chained record of the document hash, the flags, their rule ids and every model receipt id.
Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.
**Before anything else, remind me:**
- Discovery responses carry client facts and are often under a protective order. On this box nothing leaves the machine:
the language model and the checks run here. Never point it at a hosted gateway while it holds client material.
- It is a review aid, not legal advice. It cites rule text, not case law, local rules or a judge's standing order; an
attorney decides what to raise and edits and signs the letter.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/discovery-deficiency.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py discovery-deficiency` (the api image carries the same bundle under /app/rehearsal/discovery-deficiency/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py discovery-deficiency --bundle discovery-deficiency.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the set is deficient", "late service is flagged (due 17 Apr, served 20 Apr)", "the missing verification is flagged (the proof of service oath does not count)"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## What you are building
| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
Scanned PDFs (no text layer) need the document reader too: add the `parser` and `docreader` services exactly as in the
label-consistency-check or medical-chronology assemble prompts and set `DECOSA_DOCREADER_URL`. PDFs with a text layer,
Word files and pasted text need only the two services above.
## 1. Check the GPU, driver and Docker
1. `nvidia-smi`: one NVIDIA GPU with at least 32 GB (Qwen3.8-27B NVFP4 is about 20 GB of weights plus KV cache), driver
580 or newer. NVFP4 needs Blackwell; on Hopper set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main` (not
measured). Smaller cards: tell me.
2. `docker --version`, `docker compose version`, `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the NVIDIA
Container Toolkit is missing, install them from the official Docker and NVIDIA repositories after asking me.
3. About 40 GB of free disk.
## 2. Images and source
The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`. If a pull
fails, clone `decosa-api` (access required) into `~/decosa-discovery/decosa-api`, check out the newest tag that contains
`decosa_api/verticals/discovery/` (`main` until one does), and build
`docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`.
## 3. The compose file
`~/decosa-discovery/.env`:
```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_GPU_UTIL=0.55
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example LLP litigation support>"
```
`~/decosa-discovery/docker-compose.yml`:
```yaml
name: decosa-discovery
x-health: &health { interval: 15s, timeout: 5s, retries: 5 }
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
"--gpu-memory-utilization", "${LLM_GPU_UTIL}", "--max-num-seqs", "16", "--seed", "0", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
api:
image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
restart: unless-stopped
depends_on: { llm: { condition: service_healthy } }
environment:
DECOSA_LLM_ROUTE: direct # local model only; receipts are signed by this box's key ("attested")
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LOCAL_SIGNING: "on" # Ed25519 key created at /data/attest/ on first start
DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
DECOSA_DISCOVERY_SYNTHETIC_ONLY: "0" # this box accepts client material
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_LLM_TOKENS: "400000" # per session; a 25-request set uses about 5,000 generated tokens
DECOSA_SESSION_TTL_S: "28800"
ports: ["127.0.0.1:8445:8445"]
volumes: [decosa-data:/data]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/discovery/info', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```
Use the named volumes exactly as written: the api image runs as an unprivileged user, and a root-owned host bind mount
makes it fail on `/data/keys.sqlite`. Run `docker compose up -d` and poll `docker compose ps` until both services are
healthy (the LLM takes 5-10 minutes the first time). Back up the signing key with
`docker compose cp api:/data/attest ./attest-backup`; keep it private and never print it.
## 4. Smoke test on the bundled invented cases
```bash
API=localhost:8445
curl -s $API/discovery/info | jq '{synthetic_only, rules_read_on, jurisdictions}'
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"discovery-deficiency"}' | jq -r .token)
curl -s $API/discovery/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample_id":"harbor-rfp"}' > /tmp/h.json
jq -r '.status, .counts.flags, (.flags[] | "\(.item // "set")\t\(.category)\t\(.rules[0].id)")' /tmp/h.json
jq '{record}' /tmp/h.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
curl -s $API/discovery/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample_id":"marlowe-rogs"}' | jq '.status, .motion_deadline.due, [.flags[] | select(.level=="set") | .category]'
```
Pass if: `synthetic_only` is `false`; the federal sample is `deficient`, with `untimely` and `general_objections` for the
set, `withholding_unstated` on requests 2, 3 and 10, and no flag on requests 1, 4 and 7; the record verifies; and the
California sample shows `missing_verification` and `untimely` with a motion deadline of `2026-06-09` and no
`outdated_standard` flag. Every receipt should be `attested`. Each sample takes a few seconds on an idle GPU; tell me what
you measure.
## 5. Check a real set
```bash
python3 - <<'PY' > req.json
import base64, json, pathlib
f = pathlib.Path("responses.pdf") # or .docx / .txt
print(json.dumps({"file_b64": base64.b64encode(f.read_bytes()).decode(), "filename": f.name,
"jurisdiction": "federal", "mode": "theirs", # "ours" for our own draft before service
"requests_served_on": "2026-05-01", "requests_service": "email",
"privilege_log": "not_served",
"parties": {"responding_party": "Defendant Co.", "our_client": "Plaintiff LLC"}}))
PY
curl -s $API/discovery/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @req.json > out.json
jq -r .letter.text out.json > letter.txt
jq '{paragraphs: .letter.paragraphs}' out.json | curl -s $API/discovery/letter.docx -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @- -o letter.docx
```
Keep the request headings in the document ("REQUEST FOR PRODUCTION NO. 3:", "RESPONSE:", "INTERROGATORY NO. 3", "ANSWER:"):
responses are found by them. Limits: 400,000 characters, 250 responses, 12 MB per file.
## 6. Point tools at the local API
- Base URL `http://localhost:8445`; for a web app set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`.
- `POST /discovery/check` returns JSON, or Server-Sent Events with `Accept: text/event-stream`. `GET /discovery/info` lists
the categories, the rules pack sources and read date, and what it does not check.
- Keep the record with the matter file; anyone can re-check it with `POST /record/verify` against the key at
`GET /attest/signing-key`. Keep the API on 127.0.0.1; for other users, put a TLS reverse proxy with authentication in front.
## 7. Keep the direct route
`DECOSA_LLM_ROUTE=gateway` would send response text to the hosted Decosa API. Leave it off for client material.Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- A discovery deficiency letter starts with finding what is wrong: this tool reads each response, flags boilerplate objections, unstated withholding, evasive answers, improper admission responses and late service with the rule text, and turns them into a first-draft letter you rewrite.
- Who it's for
- Litigation associates, paralegals and solo litigators who review the other side's written discovery responses or check their own before serving.
- Where it runs
- Hosted for invented or public responses; self-host for client material
- Key numbers
On unseen real cases it caught 74% of the responses a motion to compel called deficient, and its splitter found 75% of responses; 6.6% of clean synthetic responses were wrongly flagged.
- 13 / 197 (6.6%) False flags on clean responses, blind synthetic test (held out, n = 197)
- 0.897 Deficient-or-not precision, blind synthetic test (held out, n = 325)
- 0.883 Deficient-or-not recall, blind synthetic test (held out, n = 325)
- 4.8 s Median end-to-end run, hosted (QA sweep 2026-09-28)
- Models
- Qwen3.8-27B (one structured read of each response) · code (splitting, rule cites, deadlines, the letter) · document reader for scanned PDFs
- Where
- Hosted for invented or public responses; self-host for client material
- Checks
- Receipt per model call; every quote found in the response at a character offset; every rule cite from a rules pack of verbatim rule text (read 28 Sep 2026); signed hash-chained record
- Industry
- Legal
- Output
- Notes, reports and drafts · Signed record or verdict
- Data
- Privileged or legal · Personal data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Decosa hosted · Self-host
- Built from
- Document reader · Signed record
Questions people ask
Does it write the discovery deficiency letter for me?
It drafts a starting point from its flags, quoting each response and citing the rule text, with the dates worked out. It is not ready to send: in blind reviews a partner preferred hand-drafted letters every time, because the draft lists points instead of arguing each request. Use it as the checklist for your own letter.
Which rules does it cite?
Federal Rules of Civil Procedure 26, 33, 34, 36 and 37 (with the 2015 amendments), the California Code of Civil Procedure discovery sections and the Texas Rules of Civil Procedure 191-198, quoted word for word from the official texts read on 28 Sep 2026. The model never picks a rule; code does. It does not apply case law or local rules.
Can I use it on client material?
Not on the hosted demo, which takes invented or public responses only and keeps nothing. Self-host it on your own GPU box, where nothing leaves the machine.
Does it get California deadlines right?
It counts 30 days plus the CCP 1013 or 1010.6 extension, rolls past California court holidays (CCP 135 and 12a), and shows the 45-day date to move to compel further responses. Check local court closures yourself.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Check their discovery responses
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…