Check a brief before filing
A checklist pointing at the exact words in the brief: fake or misread authorities, misquotations, bad record cites and personal identifiers left in.
Built on: Grounding, Signed record, Citation index
Got the other side's brief instead? Check the other side's brief
1. Pick a sample
Your own documents: run the tool on your own hardware, or request confidential access. The demo takes samples or made-up data only.
2. Run it
On production the sample took 12 s (median of 5 runs, 2026-09-29; slowest 19 s). Slower when the service is busy.
Result
The answer appears here first, then what it found, the draft, and how long it took. Sample: Fictional brief with planted errors (Doe v. Harbor Point).
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the judge; parsing, lookups, quotes and the privacy check run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the filing pre-flight API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real client material belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- filing-preflight
Use the hosted API
# Decosa Filing pre-flight: use the hosted API
You are wiring Decosa's filing pre-flight into this project. It takes a brief and the record excerpts it cites and runs
five checks: every cited case, statute and rule exists; the cited opinion supports what the brief says; record cites
resolve to the excerpts; quotations are verbatim; and no full SSN, birth date, financial-account number or minor's name
is left in (Fed. R. Civ. P. 5.2). Each proposition check is one model call with its own signed receipt, and the run ends
with a signed record. Use only what is listed below. If you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- The hosted API is for public filings. A draft brief is confidential and may be privileged: for drafts, use the
self-host prompt instead. The hosted service sends only citations to public case-law sources, never the text.
- It is a check, not legal advice: "ok" means found and matching, not that the argument is right. Say so wherever you
show results.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
`DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "filing-preflight"}` returns `{"token", "expires_at", "budget"}`.
The session's token allowance is its `budget`; how many sessions one network may start an hour is `demo_sessions` in
`GET https://api.decosa.ai/healthz`. Over a limit: HTTP 429 with `Retry-After`.
3. One run at a time per demo token (409 otherwise).
## Endpoints
- `POST /preflight/parse?name=brief.pdf` (token): the raw file as the body (`application/pdf`, a .docx, or `text/plain`;
up to 15 MB) → `{upload_id, kind, chars, text, redaction: {ran, pages: [{page, boxes, hidden_chars}]}}`. `hidden_chars`
counts text still selectable under black boxes. The text is kept for one hour under `upload_id`.
- `POST /preflight/check` (token). Body:
`{"text" | "upload_id" | "sample_id": "...", "record"?: [{"label": "App.", "text" | "upload_id": "...", "first_page"?: 1}], "authorities"?: [{"cite": "469 U.S. 325", "name"?: "...", "text": "..."}], "title"?: "...", "stream"?: true}`
- Limits: brief up to 80,000 characters; up to 8 record excerpts (600,000 characters in all). Record labels are what
the brief cites (`App.`, `J.A.`, `ECF No. 45`, `Ex. 7`, `Tr.`). Mark pages in pasted excerpts with lines like `[[212]]`;
PDFs sent by `upload_id` keep their pages.
- With `"stream": true` (or `Accept: text/event-stream`) it streams `ready` (every citation, quotation, proposition and
record cite with its offsets in the brief), `stage`, `item` events (replace by `id`), a `receipt` after each model call,
then `report`, `done` and `budget`. With `"stream": false`: one JSON object `{run_id, decision, totals, items, report, budget}`.
- Each item: `{id, check: authority|proposition|record|quote|privacy, status: problem|review|unverified|ok|skipped, title, detail, start, end, ...}`.
`problem` means positive evidence of an error; `review` needs a lawyer's eye; `unverified` means the free sources do
not cover it (never treat it as ok).
- `GET /preflight/runs/{run_id}/export?format=md` → a Markdown review for the partner; `format=record` → the signed record.
- `POST /record/verify` (no token) `{"record": {...}}` → `{ok, summary, bad}`.
- `GET /preflight/info`, `GET /preflight/samples`, `GET /preflight/samples/{id}`, `GET /attest/signing-key` (no token).
Offsets are Python string indices (Unicode code points) into the brief as sent.
## Example: stop a filing with problems (Python, `pip install httpx`)
```python
import httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
up = httpx.post(f"{API}/preflight/parse?name=brief.pdf", content=pathlib.Path("brief.pdf").read_bytes(),
headers={**H, "Content-Type": "application/pdf"}, timeout=120).json()
app = httpx.post(f"{API}/preflight/parse?name=appendix.pdf", content=pathlib.Path("appendix.pdf").read_bytes(),
headers={**H, "Content-Type": "application/pdf"}, timeout=120).json()
r = httpx.post(f"{API}/preflight/check", headers=H, timeout=900,
json={"upload_id": up["upload_id"], "record": [{"label": "App.", "upload_id": app["upload_id"]}], "stream": False})
r.raise_for_status()
run = r.json()
for it in run["items"]:
if it["status"] in ("problem", "review"):
print(it["status"].upper(), it["check"], it["title"], "--", it.get("detail"))
pathlib.Path("preflight-review.md").write_text(httpx.get(f"{API}{run['export']['md']}", headers=H).text)
if run["decision"] == "problems":
raise SystemExit("do not file: fix the problems first")
```
## Honest limits
- It does not check whether a case is still good law (no citator), and free sources miss Westlaw/Lexis-only decisions,
many unpublished orders, agency decisions and state codes: those come back `unverified`.
- Proposition verdicts come from an open model and are a triage list; see the measured rates on the Stack tab.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa Filing pre-flight: run it yourself (containers)
You are setting up the Decosa filing pre-flight on this machine, so draft briefs never leave it. It checks a brief's
citations, quotations, record cites and Rule 5.2 identifiers and signs a record. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/filing-preflight.zip (6 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py filing-preflight` (the api image carries the same bundle under /app/rehearsal/filing-preflight/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py filing-preflight --bundle filing-preflight.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the Word file is read as a brief", "the appendix PDF is read and its redaction check runs", "the pre-flight decision is: problems"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). For the GPU judge, also install the NVIDIA
container toolkit and check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and bind every port to
127.0.0.1. Ask me whether citation lookups may go out (only the citations are sent, to static.case.law,
courtlistener.com, uscode.house.gov, ecfr.gov and law.cornell.edu). If not, set `DECOSA_PREFLIGHT_OFFLINE=1`: cases are
then checked only against opinion texts I send in `authorities`, or a local mirror of the Caselaw Access Project
(`DECOSA_PREFLIGHT_CAP_BASE`). Without a GPU, set `DECOSA_PREFLIGHT_JUDGE=off` (no proposition check).
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/preflight/info` lists the five checks, `offline` and `judge`;
`GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"filing-preflight"}` and run
`POST /preflight/check {"sample_id": "demo-planted", "stream": false}`. With lookups on, expect `decision: "problems"`
with the invented Varghese cite, the one-word T.L.O. misquote, the Redding holding, App. 4 and App. 9, and the SSN,
birth date, account number and minor's name. Then `POST /record/verify` with `report.record`: `ok` must be true.
6. Report back: the public key and key id, the smoke-test decision and totals, and how long the run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
client drafts. If I ask for it later, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/filing-preflight-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMlite tierRuns with a smaller tier
The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits (48 of 96 GB).
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits (48 of 64 GB).
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"filing-preflight"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py filing-preflight
Download the mock-data bundle (6 KB, 12 checks)expected.json
A short fictional appellate brief (Doe v. Harbor Point Unified School District) as a .docx and its six-page appendix as a PDF, both sent as raw files the way a firm would upload them. The brief has planted problems: a record cite the appendix contradicts (App. 4), a cite to a page not in the appendix (App. 9), a synthetic SSN, birth date, account number and a minor's full name, an invented case and a misquote. The check must find the planted record-cite and privacy problems and end in a signed record that verifies. Case lookups go to the free public sources (Caselaw Access Project, CourtListener, LII, uscode.house.gov); the brief's text is never sent to them.
What the rehearsal checks
- the Word file is read as a brief
- the appendix PDF is read and its redaction check runs
- the pre-flight decision is: problems
- the invented case (Varghese v. China Southern Airlines, from Mata v. Avianca) is a problem
- the one-word misquote of T.L.O. is a problem
- the record cite the appendix contradicts (App. 4) is a problem
- the cite to a page not in the appendix (App. 9) is a problem
- at least four privacy items (SSN, birth date, account number, minor's name) are problems
- the partner-review Markdown lists the invented case as a problem
- the signed record verifies
- a record with its decision changed to ok no longer verifies
- every model call has a signed receipt
Licence: Fictional: the brief, the appendix, the parties and every identifier were written for Decosa (no real case or people); the Supreme Court cases it cites are public domain. Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa Filing pre-flight: run it yourself (containers)
You are setting up the Decosa filing pre-flight on this machine, so draft briefs never leave it. It checks a brief's
citations, quotations, record cites and Rule 5.2 identifiers and signs a record. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/filing-preflight.zip (6 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py filing-preflight` (the api image carries the same bundle under /app/rehearsal/filing-preflight/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py filing-preflight --bundle filing-preflight.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the Word file is read as a brief", "the appendix PDF is read and its redaction check runs", "the pre-flight decision is: problems"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). For the GPU judge, also install the NVIDIA
container toolkit and check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and bind every port to
127.0.0.1. Ask me whether citation lookups may go out (only the citations are sent, to static.case.law,
courtlistener.com, uscode.house.gov, ecfr.gov and law.cornell.edu). If not, set `DECOSA_PREFLIGHT_OFFLINE=1`: cases are
then checked only against opinion texts I send in `authorities`, or a local mirror of the Caselaw Access Project
(`DECOSA_PREFLIGHT_CAP_BASE`). Without a GPU, set `DECOSA_PREFLIGHT_JUDGE=off` (no proposition check).
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/preflight/info` lists the five checks, `offline` and `judge`;
`GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"filing-preflight"}` and run
`POST /preflight/check {"sample_id": "demo-planted", "stream": false}`. With lookups on, expect `decision: "problems"`
with the invented Varghese cite, the one-word T.L.O. misquote, the Redding holding, App. 4 and App. 9, and the SSN,
birth date, account number and minor's name. Then `POST /record/verify` with `report.record`: `ok` must be true.
6. Report back: the public key and key id, the smoke-test decision and totals, and how long the run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
client drafts. If I ask for it later, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/filing-preflight-mac.md instead.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsFiling pre-flight on GeForce RTX 5090: use the Standard · one GPU for the judge (hosted demo) tier
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
What this tool's stack says about this hardware:
- 1x RTX 5090 32 GB (fits): Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for prompts of up to about 16k tokens. Estimate: same judge as the grounding check, not run here for this vertical.
Standard · one GPU for the judge (hosted demo): what changesuses estimates
- Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Checker: decosa-api filing pre-flight (decosa_api/verticals/preflight). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Judge: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Filing pre-flight, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Filing pre-flight on my hardware Fetch https://decosa.ai/prompts/filing-preflight-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=filing-preflight) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · one GPU for the judge (hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Checker: decosa-api filing pre-flight (decosa_api/verticals/preflight), CPU - Judge: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/filing-preflight-assemble.md
Or on a Mac Studio
No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 48 GB of unified memory or more. Measured speeds and what runs where
From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh
Mac prompt for your coding agent
# Decosa Filing pre-flight: run it on this Mac (Apple Silicon, no NVIDIA GPU) You are setting up the Decosa Filing pre-flight on this Mac, natively on Apple Silicon. The models run on the Mac's GPU through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API. Every model this tool needs runs on the Mac. It needs 48 GB of unified memory or more. Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop or kill a process this setup did not start; if a port is taken, pick another one. ## Step 0: set up with a coding agent, rehearse on mock data, then go private This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works. Work in this order: 1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to "test with something realistic". 2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool, https://decosa.ai/samples/filing-preflight.zip (6 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json), show me what is in it, and run the rehearsal against the local API: `.venv/bin/python scripts/rehearse.py filing-preflight` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key). It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the Word file is read as a brief", "the appendix PDF is read and its redaction check runs", "the pre-flight decision is: problems"). Show me the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json` to make a check pass. 3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this machine. For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent can read. Switch to your own data only after the rehearsal has passed and the agent's work is done. ## What runs where | Part | On an NVIDIA GPU | On this Mac | Status | |---|---|---|---| | Checker: citation parsing, lookups, quotation matching, record cites, privacy scan, signed record (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured | | Judge: one call per proposition and per record cite, plus one call for minors' names | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured | ## Steps 1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and `sysctl -n hw.memsize` should be at least 48 GB for this tool. Check about 30 GB of free disk with `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory. 2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`. 3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`. Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me. 4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api) and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them. If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`. 5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key: show it to me, because it is what others pin to check the receipts and records this Mac signs. 6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py filing-preflight`. It runs the tool's own sample end to end against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts. `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found. 7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`, the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`. 8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of `scripts/mac/setup.sh status`. ## Good to know - Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a self-hosted Mac. - The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published evals use. Expect small differences in wording and scores. - Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel calls; typed judgments then use sampling because oMLX returns no log-probabilities). - Long briefs make long prompts: the model server's footprint peaked near 43 GB in the measured run before the setup script capped the prompt cache. - Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details: `docs/self-host-mac.md` in the checkout.
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 29 Sep 2026 · measured 29 Sep 2026: · p50 12 s · p95 19 s (5 runs) · ~$0.015 per run · 15 receipts
Loading the nightly status…
Self-host: verified 25 Sep 2026 · fresh clone, compose up, sample against local model servers
Measured cost to run: about $0.015 per brief (hosted, 29 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. Offline (the default in the prompt) the planted record-cite and privacy problems are found and the cases come back unverified, as documented; with lookups on, the same 9 problems as hosted. The signed record verifies and a tampered entry is named.
Known limits (2)
- No citator: it does not say whether a case is still good law. Westlaw- or Lexis-only decisions, many unpublished orders, agency decisions and state codes come back unverified, never OK.
- Hosted speed depends on load on the shared service: the time shown is the median of our latest production runs of the fictional sample.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the judge; parsing, lookups, quotes and the privacy check run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the filing pre-flight API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real client material belongs on your own hardware.
One check before a brief is filed: authorities exist and say what the brief claims, record cites resolve, quotations are verbatim, and Rule 5.2 identifiers are gone. Ends in a signed record.
Send a brief (pasted, PDF or DOCX) and the record excerpts it cites. The checker reads every citation with eyecite, looks each case, statute and rule up in free public sources (only the citation leaves the machine), compares every quotation with the source text by string match, resolves App., J.A., ECF and Tr. cites to the pages you gave, and scans for full SSNs, birth dates, account numbers, minors' names and text left under black boxes. An open judge model then checks each proposition against the opinion and each record cite against its page, one receipted call each. The result is a checklist that points at the exact words in the brief, a Markdown review for the partner, and a signed hash-chained record. It is a triage list for the lawyer who signs, not a verdict.
- Deployment
- Hosted or self-host
- Regulatory
- A pre-flight check, not legal advice and not a substitute for reading the authorities: under Fed. R. Civ. P. 11(b)(2) the lawyer who signs still certifies that the legal contentions are warranted by existing law, and courts have sanctioned lawyers for filing citations an AI tool invented (Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023); Park v. Kim, 91 F.4th 610 (2d Cir. 2024); Wadsworth v. Walmart Inc., No. 2:23-cv-00118 (D. Wyo. Feb. 24, 2025)). Many judges' standing orders now require disclosure of AI use or a certification that a person checked every citation; this tool does not write that certification, and its signed record is evidence that a check ran, not the certification. It does not check whether a case is still good law. Privacy: Fed. R. Civ. P. 5.2(a) (and Fed. R. Crim. P. 49.1, Fed. R. Bankr. P. 9037) allow only the last four digits of Social Security, taxpayer and financial-account numbers, the year of birth and a minor's initials; the check finds likely identifiers but cannot know which Rule 5.2(b) exemptions apply. Confidentiality: a draft brief is client information and may be work product; ABA Formal Opinion 512 (29 Jul 2024) asks lawyers to understand how a tool uses what they put in and to protect it, so self-host drafts (offline mode sends nothing out). The hosted demo is for public filings and the fictional sample, and sends only citations to the public sources. Model licence: Apache-2.0 (Qwen3.8-27B). Checked 25 Sep 2026.
Text description
A brief, its record excerpts and optionally the authorities' text go to the checker, which parses citations, quotations, propositions and record cites, looks each citation up in public sources (only the citation leaves the machine; offline mode sends nothing), matches quotations by string, resolves record cites to pages and scans for Rule 5.2 identifiers and text under black boxes. The Qwen3.8-27B judge checks each proposition against the opinion and each record cite against its page. The checker seals everything into a signed record. Outputs: a five-part checklist, a Markdown review and the signed record. Each hosted judge call gets a receipt that our gateway countersigns.
At a glance
- Data retention
- Hosted: the brief and excerpts live in memory; an uploaded file's text and a finished run are kept for one hour (for the export links), then dropped. Logs carry counts only.
- What leaves the box
- Hosted: every model call goes through our gateway to the GPU serving Qwen3.8-27B, and its receipt (hashes, token counts, no text) is kept by the gateway and this API. Self-hosted on the direct route: the brief never leaves the box; case lookups send only the citation to CAP, CourtListener, LII and uscode.house.gov; offline mode sends nothing and leaves case items unverified.
- Input formats
- Brief as text, PDF or DOCX (up to 15 MB; 80,000 characters checked); up to 8 record excerpts as text with [[page]] markers or PDFs that keep their pages.
- Model calls per brief
- One receipted judge call per proposition and per record cite; citation lookups use free public case-law data. The measured cost per brief is the cost per task at the top of the page.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
CPU only, no GPU
Existence, quotations, record pages and the privacy patterns. No proposition check, and minors' names are only flagged for review.
- Models
- decosa-api filing pre-flight (decosa_api/verticals/preflight)
- Hardware
- Any Linux or macOS machine with Python 3.11+
- Quality evidence
- Fake citations caught (existence and name checks need no model)17/19 problem, same as standarddocs/evals/filing-preflight.md, 9 held-out Solicitor General briefs; the existence check is identical without the judge
- One-word misquotations (string match, no model)14/22 problem, 21/22 problem or review, same as standarddocs/evals/filing-preflight.md, 9 held-out Solicitor General briefs
- Citations for a holding the case does not containnot checked in this tierno judge
- Latency
- estimate: seconds plus the lookups; not timed separately.
- Verification
- No proof yetSelf-host onlyNo model call, so no model receipts; the record is still signed by the box.
- In the hosted demo
Standard
one GPU for the judge (hosted demo)
Adds the proposition and record-support checks with Qwen3.8-27B, one receipted call each. This is what the hosted API runs.
- Models
- decosa-api filing pre-flight (decosa_api/verticals/preflight)
- Qwen3.8-27B (NVFP4)
- Hardware
- 1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)
- Quality evidence
- Fake citations caught (the six Mata v. Avianca fakes plus real cites with the first page moved): problem / problem or review17/19 (89%) / 18/19 (95%)docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs
- Citations for a holding the case does not contain, caught as problem12/13 (92%)docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs
- One-word misquotations: problem / problem or review14/22 (64%) / 21/22 (95%); the 7 reviews are singular/plural changesdocs/evals/filing-preflight.md, 9 held-out Solicitor General briefs
- False problems on the same briefs unaltered (511 checked items)19 (3.7%): quotations tied to the wrong source 11, authorities the free indexes lack 6, name parse 1, proposition 1docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs
- Propositions read by the judge on real briefs: supported / partly / not supported by passages read / contradicted29 / 24 / 25 / 4 of 82 (all but 1 shown as review, not problem)docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs
- Privacy patterns on the full text of 12 real briefs (570,911 characters)0 false findingsdocs/evals/filing-preflight.md
- Latency
- Measured on our server, shared card, gateway route: most of the time is judge calls and CourtListener lookups. The sample's measured time per task is at the top of the page; a real argument section's is in the latency table.
- Verification
- Proof: strongEvery judge call is a separate gateway call with a gateway-signed receipt; the signed record lists them.
- Needs more compute
Wanted: the best setup
a GLM-5.3-Flash judge on your own hardware
A stronger open judge from another family for propositions, where the 27B is least sure. Briefs can be privileged, so it runs on your hardware, never on community providers. Not served yet.
- Models
- decosa-api filing pre-flight (decosa_api/verticals/preflight)
- Qwen3.8-27B (NVFP4)
- GLM-5.3-Flash (NVFP4)
- Hardware
- Your own hardware: 2x 96 GB cards (NVFP4, about 170-186 GB, unconfirmed) or a Mac with 192 GB or more (MLX 4-bit, 165 GB). Estimate; GLM's SGLang SM120 build hung on our server.
- Quality evidence
- This eval, same protocolnot measured yet
- Latency
- not measured yet
- Verification
- No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Also runs on
- A second judge on one cardNemotron-3-Super-120B-A12B (NVFP4)not servedA cross-family second judge that fits beside the 27B, so we would host it ourselves. sm_120 support unconfirmed. Hardware: 1x RTX PRO 6000 96 GB (80 GB of weights).
We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Checker: citation parsing, lookups, quotation matching, record cites, privacy scan, signed record (no model; CPU)decosa-api filing pre-flight (decosa_api/verticals/preflight) 0 GBProof: partial | LiteStandardWanted | 0 GB | Proof: partial | |
| ||||
Judge: one call per proposition and per record cite, plus one call for minors' namesQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | StandardWanted | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Stronger judge for propositions (wanted)GLM-5.3-Flash (NVFP4) about 170 GB (estimate)No proof yetSelf-host only | Wanted | about 170 GB (estimate) | No proof yetSelf-host only | |
| ||||
One-card second judge from another model familyNemotron-3-Super-120B-A12B (NVFP4)nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 on Hugging Face (opens in a new tab) 120B (12B active) · about 80 GB (estimate)No proof yetSelf-host only | Alternate | 120B (12B active) · about 80 GB (estimate) | No proof yetSelf-host only | |
| ||||
Tools, services and hardware
Tools
- eyecite (Free Law Project) (opens in a new tab)BSD-2-Clause
Finds full, short, id. and supra citations to cases, statutes and journals; reporters-db and courts-db (BSD-2-Clause) come with it.
- Caselaw Access Project (static.case.law) (opens in a new tab)Public domain (CC0 data)
Case existence by volume and first page, each case's page range, and opinion text with the reporter's page numbers, for most reporters up to about 2018-2020. Also downloadable in bulk for an offline mirror.
- Local case-law citation index (CourtListener bulk data + Caselaw Access Project metadata) (opens in a new tab)Public domain (CourtListener bulk data: Public Domain Mark; CAP metadata: CC0 1.0)
Every case citation is looked up here first, with no network call: 18.1 M citations, 10.1 M opinions, 6.9 M CAP page ranges and 34.9 M docket names, quarterly snapshot (30 Jun 2026).
- CourtListener search API (Free Law Project) (opens in a new tab)Free API (anonymous use is rate-limited); opinions are public domain
Only when the local index misses or a volume is newer than its snapshot: the citation search, a name search, and the opinion PDF from CourtListener's public storage.
Whether a cited U.S.C. section exists, and its text for quotations.
- eCFR API (opens in a new tab)Public domain
Whether a cited C.F.R. section exists in the current edition, and its text.
- Legal Information Institute, federal rules (opens in a new tab)Rule text is public domain
Fed. R. Civ. P., Crim. P., App. P., Bankr. P. and Evid.: whether the rule and subdivision exist, and their text.
- scripts/preflight_eval.py and docs/evals/filing-preflight.mdApache-2.0
The eval: planted fake citations, wrong propositions and misquotes in real Solicitor General briefs, and the false-problem rate on the same briefs unaltered.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /preflight/info, /preflight/samples; POST /preflight/parse (PDF, DOCX or text), /preflight/check (SSE or JSON); GET /preflight/runs/{id}/export?format=md|record. Keeps briefs in memory for one hour, never on disk.
- vLLM (judge):8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).
Hardware
- Any CPU, no GPU Fits
CPU-only tier (DECOSA_PREFLIGHT_JUDGE=off): existence, quotations, record-page resolution and the privacy patterns; no proposition check. Estimate: not timed separately.
- 1x RTX 5090 32 GB Fits
Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for prompts of up to about 16k tokens. Estimate: same judge as the grounding check, not run here for this vertical.
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server: the hosted demo and the eval ran on this card, shared with other services.
Latency per lane
- 9,000-character argument section of a real brief, end to end35.0 s
Measuredmeasured on our server 2026-09-25 over the eval's briefs: the median of the test split (docs/evals/filing-preflight.md)
- citation lookup, cached vs uncached400 ms
Measuredmeasured on our server 2026-09-28: a case from the local index, or from CAP or CourtListener when the index misses; lookups are cached for a day
Notes
- Quotations and record pages are compared by code: both sides are reduced to letters and digits, so line breaks, hyphenation and curly quotes do not count, but a changed, added or dropped word does, and the item shows the word diff. Bluebook alterations ([W]arrants, ellipses, bracketed insertions) are honoured.
- Statuses are conservative. 'Problem' needs positive evidence; a source that is missing, unreachable or does not cover an authority gives 'unverified', never 'OK'. A proposition is a 'problem' only when the sentence attributes a holding to the court and the judge finds it contradicted or unsupported with high certainty; otherwise 'review'.
- Most false problems on real briefs are quotations tied to the wrong source (words quoted from the record or a statute next to a case cite) and Supreme Court orders or recent Westlaw-only decisions that the free indexes do not hold.
- It does not check whether a case is still good law (no citator), and it does not write the judge's AI-use certification.
- Case citations are looked up in a local index (CourtListener and CAP public data, quarterly; snapshot 30 Jun 2026) with no network call; CourtListener's API is asked only on a miss or for a volume newer than the snapshot, and a free account token (DECOSA_PREFLIGHT_COURTLISTENER_TOKEN) raises its limit. Offline mode (DECOSA_PREFLIGHT_OFFLINE=1) sends nothing out and still uses the local index if one is installed.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa filing pre-flight on this machine
You are setting up a pre-flight check for court filings. It takes a brief (text, PDF or DOCX) and the record excerpts it
cites, and checks five things: every cited case, statute and rule exists; the cited opinion says what the brief claims;
record cites resolve to the excerpts; quotations are verbatim; and no full Social Security number, birth date,
financial-account number or minor's name is left in (Fed. R. Civ. P. 5.2). It ends with a record signed by this box's
own key. Work step by step, show me each command before you run anything with `sudo`, and stop to ask if a check fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/filing-preflight.zip (6 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py filing-preflight` (the api image carries the same bundle under /app/rehearsal/filing-preflight/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py filing-preflight --bundle filing-preflight.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the Word file is read as a brief", "the appendix PDF is read and its redaction check runs", "the pre-flight decision is: problems"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- Model: Qwen3.8-27B (Apache-2.0) as the judge for propositions and record cites. Everything else is CPU code:
decosa-api (AGPL-3.0-or-later), eyecite (BSD-2-Clause) for citations, pdfplumber (MIT) for the redaction check, and poppler's
`pdftotext` (GPL-2.0, run as a separate program).
- A draft brief is confidential and may be privileged or work product. Bind every port to 127.0.0.1. The service keeps
briefs and results in memory for one hour and never writes them to disk or logs; keep it that way.
- Decide with me which mode to run:
- **Offline** (`DECOSA_PREFLIGHT_OFFLINE=1`): nothing leaves this machine. Authorities are checked only against texts
I supply with the request (or a local mirror of the Caselaw Access Project bulk data, step 2).
- **Online lookups** (the default): only citations (volume, reporter, page; statute sections; a case name when a cite
is not found) go to static.case.law, courtlistener.com, uscode.house.gov, ecfr.gov and law.cornell.edu. Never the
brief's text.
- Be honest about what it does: "OK" means found and matching, not that the argument is right. It is not legal advice,
and a lawyer still reads the authorities.
## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB for the judge (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV
cache; opinions up to 60,000 characters go to the judge whole, about 15k tokens). Driver 570 or newer; Blackwell cards
run NVFP4, older cards use the FP8 weights.
2. No GPU? Run the CPU-only tier: set `DECOSA_PREFLIGHT_JUDGE=off`. Existence, quotations, record-cite resolution and
the privacy patterns still run; propositions are reported as not checked, and minors' names as "review".
3. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
from the official Docker and NVIDIA repositories after asking me, then run
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
4. Disk: about 30 GB free for the model and images. A local CAP mirror adds tens of GB; only fetch the reporters you need.
## 2. Images, weights and (optional) a local case-law mirror
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
`git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out the newest release tag that contains
`decosa_api/verticals/preflight/` (`main` until one does: v0.1.0 predates it), and build `docker/api/Dockerfile` (it installs the `preflight` extra and poppler).
Without Docker: `pip install ".[preflight]"` and `apt install poppler-utils`.
- `vllm/vllm-openai:v0.29.0` for the judge; weights `nvidia/Qwen3.8-27B-NVFP4` (or `Qwen/Qwen3.8-27B-FP8`).
- Optional mirror for offline case lookups: copy the reporters you cite from `https://static.case.law/` (for each
`<reporter>/<volume>/`: `CasesMetadata.json` and `html/`, plus `ReportersMetadata.json` at the root) into
`~/decosa/cap/` and serve it read-only on 127.0.0.1:8480 (`python -m http.server 8480 --bind 127.0.0.1 -d ~/decosa/cap`).
Then set `DECOSA_PREFLIGHT_CAP_BASE=http://host.docker.internal:8480` (or the host's address on the compose network).
## 3. docker-compose.yml
Write this in `~/decosa/preflight/`:
```yaml
services:
llm:
image: vllm/vllm-openai:v0.29.0
command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
"--enable-prefix-caching"]
ports: ["127.0.0.1:8114:8000"]
volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
api:
image: ${DECOSA_REGISTRY}/decosa-api:<tag>
ports: ["127.0.0.1:8445:8445"]
environment:
DECOSA_HOST: 0.0.0.0
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_BUDGET_LLM_TOKENS: "60000" # per demo session; a long brief needs about 170 tokens per proposition
DECOSA_PREFLIGHT_OFFLINE: "1" # remove to allow citation lookups (see step 0)
DECOSA_PREFLIGHT_MAX_CONCURRENT: "2"
volumes: ["decosa-data:/data"]
depends_on: { llm: { condition: service_healthy } }
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/preflight/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
decosa-data:
```
The api keeps its state (keys, receipts, this box's signing key) in the named volume `decosa-data`, not in a
host folder: the image runs as an unprivileged user (uid 10001), and a host folder that Docker creates is owned by
root, which stops the api with `PermissionError: [Errno 13] Permission denied: '/data/keys.sqlite'`. Then start everything: `docker compose up -d`.
Prefix caching matters: every proposition cited to the same opinion sends that opinion first, so the server reuses it.
On first start the api service creates this box's Ed25519 key in the `decosa-data` volume (`/data/attest/` in the api container, mode 0600). Back it up and
never print it. Records and model calls are signed with it: an attestation by me, the operator, not a proof.
## 4. Smoke test
1. `curl -s localhost:8445/preflight/info | jq '{checks, offline, judge}'` lists the five checks and the mode.
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"filing-preflight"}' | jq -r .token)`.
3. Run the fictional sample:
`curl -s -XPOST localhost:8445/preflight/check -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample_id":"demo-planted","stream":false}' > run.json`.
With lookups on, expect `decision: "problems"` and problems for: the invented case (Varghese, 925 F.3d 1339), the
changed word in the T.L.O. quotation ("permit" for "require"), the Redding holding, the App. 4 record cite (the page
says the nurse ordered the search), App. 9 (not in the excerpt), and the SSN, birth date, account number and the
minor's full name. Offline, the case items are "unverified" unless you pass the opinions in `authorities`.
4. Stream it with `-H 'accept: text/event-stream' -N` and `"stream": true`: `ready`, `stage`, `item` events, a `receipt`
after each model call, then `report`, `done` and `budget`.
5. `jq '{record: .report.record}' run.json | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @-`
must say `ok: true`. Change one item's `status` in the record and verify again: it must fail and name that entry.
6. Upload a PDF: `curl -s -XPOST 'localhost:8445/preflight/parse?name=brief.pdf' -H "authorization: Bearer $T" -H 'content-type: application/pdf' --data-binary @brief.pdf | jq '{upload_id, chars, redaction}'`,
then check it with `{"upload_id": "..."}`. A PDF with a black box drawn over text shows `hidden_chars` > 0.
7. Time it and tell me how long the fictional sample and one real argument section take end to end on this machine
(longer with lookups on). Expect minutes, not seconds, on a busy card: most of the time is judge calls.
## 5. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`, or call `POST /preflight/check` from
your document system before filing and keep the signed record and the Markdown review
(`GET /preflight/runs/{id}/export?format=md|record`) in the matter file. Contract: `API_CONTRACT.md`, section
"Filing pre-flight".
Off by default. Joining serves other people's requests on this GPU; never do it on a box that holds client drafts.
If I ask for it, follow the provider guide at `/provide` on the site, and do not enable it without my explicit yes.Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
1 law, rule and guidance page cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- One check before a brief is filed: authorities exist and say what the brief claims, record cites resolve, quotations are verbatim, and Rule 5.2 identifiers are gone. Ends in a signed record.
- Who it's for
- Teams in legal.
- Where it runs
- Self-host for drafts; hosted for public filings
- Key numbers
- 89% Fake citations caught as problem (strict) (test split, n = 19)
- 92% Citations for a holding the case does not contain, caught (test split, n = 13)
- 64% One-word misquotations caught as problem (strict) (test split, n = 22)
- 12.1 s Median end-to-end run, hosted (QA sweep 2026-09-29)
- Models
- Qwen3.8-27B
- Where
- Self-host for drafts; hosted for public filings
- Checks
- Exact-match quotes; receipt per proposition; signed record
- Industry
- Legal
- Output
- Signed record or verdict · Structured data
- Data
- Privileged or legal · Personal data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Built from
- Grounding · Signed record · Citation index
Questions people ask
Does the filing pre-flight catch fake AI citations?
On 9 held-out real Solicitor General briefs with planted errors, Decosa's filing pre-flight caught 17 of 19 fake citations as a problem (89% strict, 95% lenient), including Mata v. Avianca fakes, and 12 of 13 citations for a holding the case does not contain. The test set is small, so the rates have wide intervals (Wilson 95% 0.69-0.97 for fakes), and the planted wrong holdings lean toward clear reversals.
What does it check besides whether a case exists?
Decosa's filing pre-flight compares every quotation with the source text by string match, resolves App., J.A., ECF and Tr. record cites to the excerpts you supply, has an open judge model check each proposition against the opinion, and scans for full SSNs, birth dates, account numbers, minors' names and text left under black boxes, which Fed. R. Civ. P. 5.2(a) limits. It cannot know which Rule 5.2(b) exemptions apply.
Does it tell me whether a case is still good law?
No. Decosa's filing pre-flight has no citator, so it does not say whether a case was overruled, reversed or distinguished. Westlaw- or Lexis-only decisions, many unpublished orders, agency decisions and state codes come back unverified, never OK. On propositions it is a triage list: 47% of real-brief propositions came back "review". The lawyer who signs still certifies the legal contentions under Fed. R. Civ. P. 11(b)(2).
How often does the citation checker raise false problems?
On 9 unaltered real Solicitor General briefs, Decosa's filing pre-flight marked 19 of 511 checked items (3.7%) as a problem when nothing was wrong, plus 67 items (13%) marked review; the Wilson 95% interval is 0.02-0.06. The real briefs come from one careful filer, so briefs citing more unpublished or Westlaw-only decisions are covered less well and will see more unverified and review items.
Does my draft brief leave the firm?
Self-hosted on the direct route, the brief never leaves the box: Decosa's filing pre-flight sends only each citation to CAP, CourtListener, LII and uscode.house.gov, and offline mode sends nothing and leaves case items unverified. The hosted route keeps an uploaded file's text and a finished run for one hour, then drops them, and is meant for public filings, because a draft brief can be privileged or work product.
Is the signed record the AI certification a judge asks for?
No. Decosa's filing pre-flight ends in a signed hash-chained record that is evidence a check ran and what it found; change one status and verification names the entry. Many judges' standing orders require disclosure of AI use or a certification that a person checked every citation. The tool does not write that certification, and its record is not the certification.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Filing pre-flight
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…