Skip to content
decosa
LiveHostedSelf-host

Check the other side's brief

Findings for your reply: cases the public sources don't have, changed quotations and instructions hidden for AI tools, each tied to a page and a source.

Held-out test24 / 31 (77%)Non-existent citations flagged (held-out test)
On production8.3 smedian on production (2026-09-30); slower when the service is busy
List price~$0.013 per briefmeasured, at list price

Built on: Citation index, Grounding

Checking your own brief before you file? Check a brief before filing

Their filing

A two-page fictional opposition to a summary-judgment motion (M.R. v. Lakeview Unified School District) that cites real Supreme Court cases. Built as a PDF, the way a filing arrives. Planted: a case no public source has at a page inside another case, a Westlaw cite no public source has, a one-word change inside a T.L.O. quotation, a holding Redding does not contain, white text telling AI reviewers not to flag citations, 1-point text telling a summariser what to say, an instruction in the file's Keywords, and a student's full birth date. Anderson, Celotex and Harlow are quoted correctly. Open the PDF

Findings to verify, not legal advice. A case the public sources lack is “not found in the sources searched”, never “fake”: pull the reporter before you say so to a court. Only the citations go to the public case-law sources, never the text. A filed brief is public; for anything that is not, self-host.

Findings

Live

Results appear here: findings for your reply, each tied to a page of their filing and a public source, then a Word memo and a signed record.

Watch a recorded run first

Watch: the other side's brief checked

Replay · not live

Recorded from real runs of check-their-brief on production (api.decosa.ai) on 30 Sep 2026: Qwen3.8-27B through the gateway, lookups in the Caselaw Access Project and CourtListener. Replayed here without a backend; timings are compressed.

Planted: a case no public source has at a page inside another case, a Westlaw cite no public source has, a one-word change inside a T.L.O. quotation, a holding Redding does not contain, white text telling AI reviewers not to flag citations, 1-point text telling a summariser what to say, an instruction in the file's Keywords, and a student's full birth date. Anderson, Celotex and Harlow are quoted correctly.

Results appear here: findings for your reply, each tied to a page of their filing and a public source, then a Word memo and a signed record.

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the check their brief API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) or 1x RTX 5090 (32 GB) for the judge; the hidden-text scan, lookups and quotation match run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
check-their-brief

Use the hosted API

# Decosa "Check the other side's brief": use the hosted API

You are wiring Decosa's check of an opposing filing into this project. It takes the other side's filing (a PDF, a Word
file or text) and returns findings for a reply: citations the public sources do not have (with the URLs searched),
quotations that differ from the opinion, holdings the opinion does not contain, identifiers left in (Fed. R. Civ. P.
5.2), and text the page hides (white, tiny or off-page text, invisible characters, metadata, comments), with a flag on
any of it that is aimed at AI tools. Each model call has its own signed receipt, and the run ends with a signed record
and a Word findings memo. Use only what is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- A filed brief is a public record, so the hosted API is fine for it. Your notes, and papers served but not filed, belong
  on your own box: use the self-host prompt for those. Only citations and case names go to the public case-law sources.
- These are findings for a lawyer to verify, not legal advice. Never describe a citation as "fake": show the finding's
  own wording ("not found in the public sources searched") and its search trail.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page, kept in `DECOSA_API_KEY`, never in code.
   Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "check-their-brief"}` returns `{"token", "expires_at", "budget"}`.
   The session's token allowance is its `budget`; how many sessions one network may start an hour is `demo_sessions` in
   `GET https://api.decosa.ai/healthz`. Over a limit: HTTP 429 with `Retry-After`. One run at a time per demo token (409 otherwise).

## Endpoints
- `POST /theirbrief/parse?name=opp.pdf` (token): the raw file as the body (`application/pdf`, a .docx or `text/plain`;
  up to 20 MB) → `{upload_id, kind, chars, text, hidden: {ran, pages, findings: [{kind, how, page, text, chars, detail}]}, ocr}`.
  A PDF is rendered page by page and compared with its text layer; a scan is read by the document reader (`ocr` set).
  Text under a black box is counted, never returned. The text is kept for one hour under `upload_id`.
- `POST /theirbrief/check` (token). Body: `{"upload_id" | "text" | "sample_id": "...", "record"?: [{"label": "App.", "text": "..."}], "authorities"?: [{"cite": "469 U.S. 325", "name"?: "...", "text": "..."}], "title"?: "...", "stream"?: true}`.
  Up to 120,000 characters. Send the file by `upload_id` to get the hidden-text checks; pasted text can only be checked
  for invisible characters.
  - Streaming (`"stream": true` or `Accept: text/event-stream`): `ready`, `stage`, `item` events (replace by `id`), a
    `receipt` after each model call, then `report`, `done` and `budget`. `"stream": false`: one JSON object
    `{run_id, decision, totals, items, report, budget}`; `decision` is `findings`, `review` or `clear`.
  - Each item: `{id, check: hidden|authority|quote|proposition|record|privacy, status: problem|review|unverified|ok|skipped, headline, title, detail, page, their_sentence?, hidden_text?, trail?: [{source, what, url, status}], evidence?, diff?, receipt_ids?}`.
    `page` is the page of their filing. `unverified` means the free sources do not cover it: never show it as a finding
    or as fine.
- `GET /theirbrief/runs/{run_id}/export?format=docx` → the Word findings memo; `format=md` → Markdown; `format=record` → the signed record.
- `POST /record/verify` (no token) `{"record": {...}}` → `{ok, summary, bad}`.
- `GET /theirbrief/info`, `GET /theirbrief/samples`, `GET /theirbrief/samples/{id}`, `GET /theirbrief/samples/{id}/file` (no token).

## Example: check an incoming filing and save the memo (Python, `pip install httpx`)
```python
import httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
up = httpx.post(f"{API}/theirbrief/parse?name=opposition.pdf", content=pathlib.Path("opposition.pdf").read_bytes(),
                headers={**H, "Content-Type": "application/pdf"}, timeout=180).json()
run = httpx.post(f"{API}/theirbrief/check", headers=H, timeout=900, json={"upload_id": up["upload_id"], "stream": False}).json()
for it in run["items"]:
    if it["status"] == "problem":
        print(f"p. {it.get('page') or '-'}  {it['headline']}: {it['title']}")
pathlib.Path("findings.docx").write_bytes(httpx.get(f"{API}{run['export']['docx']}", headers=H, timeout=60).content)
```

## Honest limits
- No citator: it does not say whether a case is still good law. Westlaw- and Lexis-only decisions, many unpublished
  orders and state codes are not in the free sources: they come back `review` or `unverified`, never `ok`.
- Holdings are checked by an open model and are a triage list; the measured rates are on the Stack tab.
- Instructions hidden in the file are caught well; instructions in plain view are caught when a pattern or an AI word
  flags the sentence first.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa "Check the other side's brief": run it yourself (containers)

You are setting up Decosa's check of opposing filings on this machine, so your notes and any unfiled papers never leave
it. It checks their citations, quotations, holdings and Rule 5.2 identifiers, scans the file for hidden text and
instructions aimed at AI tools, and signs a record. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/check-their-brief.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py check-their-brief` (the api image carries the same bundle under /app/rehearsal/check-their-brief/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py check-their-brief --bundle check-their-brief.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the PDF is read and its hidden-text scan runs", "the scan finds the hidden texts before any model runs", "the decision is: findings"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions (docs.docker.com/engine/install). For the GPU judge, also install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Read it.
   Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For `api` set `DECOSA_LLM_ROUTE=direct`,
   `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_THEIRBRIEF_OCR=off` (unless the document
   reader service runs next to it) and bind every port to 127.0.0.1. Ask me whether lookups may go out (only citations
   and case names, to static.case.law, courtlistener.com, uscode.house.gov, ecfr.gov and law.cornell.edu). If not, set
   `DECOSA_PREFLIGHT_OFFLINE=1`. If I have a free CourtListener token, set `DECOSA_PREFLIGHT_COURTLISTENER_TOKEN`.
   Without a GPU, set `DECOSA_THEIRBRIEF_JUDGE=off` (no holdings check; the hidden-text scan and patterns still run).
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/theirbrief/info` lists the six checks, `offline`, `judge` and `ocr`;
   `GET /attest/signing-key` shows this box's public key. Show me the key.
5. Smoke test: get a token with `POST /demo/session {"vertical":"check-their-brief"}` and run
   `POST /theirbrief/check {"sample_id": "demo-opposition", "stream": false}`. Expect `decision: "findings"` with three
   hidden instructions aimed at AI tools, the student's birth date, and (with lookups on) Hollis v. Brentwood School
   District not found, the T.L.O. misquote and the Redding holding. Then `POST /record/verify` with `report.record`:
   `ok` must be true. Download `GET /theirbrief/runs/{run_id}/export?format=docx`.
6. Report back: the public key and key id, the smoke-test decision and totals, and how long the run took.

## Network (optional)
Off by default. Joining as a provider serves other people's requests on this GPU; never on a box that holds client
work. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMlite tierRuns with a smaller tier

    The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits.

  • GeForce RTX 4090lite tierRuns with a smaller tier

    The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits.

  • GeForce RTX 5090lite tierRuns with a smaller tier

    The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits.

  • 2x GeForce RTX 5090lite tierRuns with a smaller tier

    The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits.

  • L40Slite tierRuns with a smaller tier

    The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits.

  • H100 80 GB (SXM)lite tierRuns with a smaller tier

    The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits.

  • RTX PRO 6000 Blackwell 96 GBlite tierRuns with a smaller tier

    The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits.

  • 2x RTX PRO 6000 Blackwell 96 GBlite tierRuns with a smaller tier

    The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits.

  • Apple M3 Ultra (Mac Studio), 96 GBlite tierRuns with a smaller tier

    The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) has no mapped Apple Silicon build The lite tier fits.

  • Apple M5 Max, 64 GBlite tierRuns with a smaller tier

    The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) has no mapped Apple Silicon build The lite tier fits.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"check-their-brief"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py check-their-brief

Download the mock-data bundle (4 KB, 11 checks)expected.json

A two-page fictional opposition to a summary-judgment motion (M.R. v. Lakeview Unified School District) as a PDF, sent as a raw file. It cites real Supreme Court cases and carries planted problems: white text telling AI reviewers not to flag citations, 1-point text telling a summariser what to say, an instruction in the file's Keywords property, a case at a page that sits inside another case (925 F.3d 1339), a Westlaw cite no public source has, a one-word change inside a T.L.O. quotation, a holding Redding does not contain, and a student's full birth date. The check must find the hidden instructions and the birth date, report the planted citation problems when the public sources answer, and end in a signed record that verifies. Lookups send only citations to the Caselaw Access Project and CourtListener, never the text.

What the rehearsal checks
  • the PDF is read and its hidden-text scan runs
  • the scan finds the hidden texts before any model runs
  • the decision is: findings
  • three hidden instructions aimed at AI tools are findings
  • the harmless hidden 'DRAFT' label is not called an instruction
  • the student's full birth date is a finding
  • nothing is called fake
  • the memo has the table for the reply
  • the signed record verifies
  • a record with its decision changed no longer verifies
  • every model call has a signed receipt

Licence: Fictional: the filing, parties and identifiers were written for Decosa (no real case or people); the Supreme Court cases it cites are public domain. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa "Check the other side's brief": run it yourself (containers)

You are setting up Decosa's check of opposing filings on this machine, so your notes and any unfiled papers never leave
it. It checks their citations, quotations, holdings and Rule 5.2 identifiers, scans the file for hidden text and
instructions aimed at AI tools, and signs a record. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/check-their-brief.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py check-their-brief` (the api image carries the same bundle under /app/rehearsal/check-their-brief/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py check-their-brief --bundle check-their-brief.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the PDF is read and its hidden-text scan runs", "the scan finds the hidden texts before any model runs", "the decision is: findings"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions (docs.docker.com/engine/install). For the GPU judge, also install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Read it.
   Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For `api` set `DECOSA_LLM_ROUTE=direct`,
   `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_THEIRBRIEF_OCR=off` (unless the document
   reader service runs next to it) and bind every port to 127.0.0.1. Ask me whether lookups may go out (only citations
   and case names, to static.case.law, courtlistener.com, uscode.house.gov, ecfr.gov and law.cornell.edu). If not, set
   `DECOSA_PREFLIGHT_OFFLINE=1`. If I have a free CourtListener token, set `DECOSA_PREFLIGHT_COURTLISTENER_TOKEN`.
   Without a GPU, set `DECOSA_THEIRBRIEF_JUDGE=off` (no holdings check; the hidden-text scan and patterns still run).
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/theirbrief/info` lists the six checks, `offline`, `judge` and `ocr`;
   `GET /attest/signing-key` shows this box's public key. Show me the key.
5. Smoke test: get a token with `POST /demo/session {"vertical":"check-their-brief"}` and run
   `POST /theirbrief/check {"sample_id": "demo-opposition", "stream": false}`. Expect `decision: "findings"` with three
   hidden instructions aimed at AI tools, the student's birth date, and (with lookups on) Hollis v. Brentwood School
   District not found, the T.L.O. misquote and the Redding holding. Then `POST /record/verify` with `report.record`:
   `ok` must be true. Download `GET /theirbrief/runs/{run_id}/export?format=docx`.
6. Report back: the public key and key id, the smoke-test decision and totals, and how long the run took.

## Network (optional)
Off by default. Joining as a provider serves other people's requests on this GPU; never on a box that holds client
work. If I ask for it later, follow the Provide page instead of improvising.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

Runs with a smaller tierCheck their brief on GeForce RTX 5090: use the Lite · CPU only, no GPU tier

The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB (fits): Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: not run here for this use case.

Lite · CPU only, no GPU: what changes

Nothing: it runs as listed in the stack.

Memory per component
  • Checker: decosa-api check-their-brief (decosa_api/verticals/theirbrief, on the filing pre-flight engine). CPU. Runs on CPU (vram_gb 0 in stack.json).

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Check their brief, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Check their brief on my hardware

Fetch https://decosa.ai/prompts/check-their-brief-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=check-their-brief)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Lite · CPU only, no GPU (lite). Fit check: runs, about 0 GB of 32 GB used.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Checker: decosa-api check-their-brief (decosa_api/verticals/theirbrief, on the filing pre-flight engine), CPU

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/check-their-brief-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 30 Sep 2026 · measured 30 Sep 2026: · p50 8.3 s · p95 8.4 s (5 runs) · ~$0.013 per run · 7 receipts

Loading the nightly status…

Self-host: verified 28 Sep 2026 · fresh clone, compose up, sample against local model servers

Measured cost to run: about $0.013 per brief (hosted, 30 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Fresh clone of a decosa-api pre-release build (6ee6b6b), the api image built from docker/api/Dockerfile (theirbrief extra), this prompt's compose with the direct route to the already-running local Qwen3.8-27B, OCR off, anonymous CourtListener. The prompt's smoke steps and the rehearsal bundle passed (11/11): 7 findings on the fictional PDF, 7 receipts, the record verifies and a tampered decision fails, the Word memo exports, a scan is refused with a clear message when OCR is off. 57.9 s, $0.0126. Torn down after.

Known limits (6)
  • Hosted numbers are the production smoke check of the sample, run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
  • No citator: it does not say whether a case is still good law.
  • Westlaw- and Lexis-only decisions, many unpublished orders and most state codes are not in the free sources: they come back 'look at' or 'not checkable', never 'fine'.
  • Case citations are looked up in a local index of CourtListener's and the Caselaw Access Project's public data (quarterly; snapshot 30 Jun 2026), with no network call. CourtListener's free API (250 searches a day) is asked only on a miss or for a volume newer than the snapshot; what nothing can answer is listed as 'not checked yet', never as fine.
  • Quotations and holdings of cases after about 2018 need the opinion PDF from CourtListener, one search per quoted case; if it cannot be asked, those checks stay 'not checked'.
  • Instructions hidden in the file are caught well; instructions written in plain view are caught only when a pattern or AI word flags the sentence first.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the check their brief API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) or 1x RTX 5090 (32 GB) for the judge; the hidden-text scan, lookups and quotation match run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

Upload the other side's filing: fake-looking citations, changed quotations, misstated holdings and hidden text aimed at AI tools, each tied to a page and a source.

For the associate or paralegal answering an opposition or reply. Upload their filing (PDF, Word or text). The checker renders every page and compares it with the text layer to find what a reader cannot see (white, tiny or off-page text, invisible characters, metadata, comments), and one receipted model call decides which of those texts speak to AI tools. It then reads every citation with eyecite, flags reporter series that do not exist, looks each case, statute and rule up in free public sources with a search trail, matches every quotation against the opinion by string, and has an open judge model check each holding against the opinion's text. Nothing is called fake: a miss is 'not found in the sources searched'. The result is findings for the reply, a Word memo and a signed record.

Deployment
Hosted or self-host
Regulatory
Findings for a lawyer to verify, not legal advice (as of 28 Sep 2026). Before telling a court that a cited case does not exist, pull the reporter: public sources miss some unpublished, very recent and Westlaw- or Lexis-only decisions. Fed. R. Civ. P. 11(b)(2) makes the signer of a filing certify that its legal contentions are warranted by existing law, and courts sanction invented citations (Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023)); a sanctions motion under Rule 11(c)(2) must be served 21 days before it is filed. If the scan finds text under a black box (a failed redaction), ABA Model Rule 4.4(b) requires a lawyer who receives inadvertently sent information to notify the sender; the memo never shows those words. A filed brief is a public record; your notes about it, and papers served but not filed, belong on your own box or the confidential tier.
Architecture
Text description

Their filing goes to the checker, which renders each PDF page against its text layer for hidden text and reads metadata and comments (scans go through the document reader first), blanks hidden text out, parses citations, flags reporters that do not exist, looks each citation up in public sources with a search trail, matches quotations by string and scans for Rule 5.2 identifiers. The Qwen3.8-27B judge checks each holding against the opinion and, once, which hidden texts speak to AI tools; each call gets a receipt. Outputs: findings for the reply tied to pages of their filing, a Word memo and a signed record.

Architecture

At a glance

Data retention
Hosted: the filing lives in memory; an uploaded file's text and a finished run are kept for one hour (for the export links), then dropped. Answers from the public sources (the citation queried, public opinion text; never the filing) are cached on disk for 7 days. Logs carry counts only.
What leaves the box
Hosted: each model call goes through the gateway to the GPU serving Qwen3.8-27B, with a signed receipt (hashes and token counts, no text). Self-hosted on the direct route: the filing never leaves the box; lookups send only citations and case names to CAP, CourtListener, LII, uscode.house.gov and eCFR; offline mode sends nothing.
Input formats
Their filing as a PDF (hidden text is checked), a Word file or pasted text; up to 20 MB and 120,000 characters. Scanned PDFs are read by the document reader (hosted, up to 40 pages).
Typical run
One model call reads the hidden texts and one checks each holding. The first check of a filing is slower, while the public sources are looked up; repeat lookups are cached. The page header shows the time and cost measured on production, and each run shows its own.
Wording
Nothing is called fake. A case the public sources lack is 'not found in the sources searched', with the URLs searched; a real case at the wrong cite is 'their cite is wrong'.
Confidential tier
For papers that are not public (served but unfiled, your notes): the same check can run inside an attested enclave, where neither Decosa nor the cloud host can read the text; enterprise pricing, ask us. A filed brief is public, so the hosted demo is fine for it.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    CPU only, no GPU

    The hidden-text scan with the injection patterns, existence, impossible reporters and quotations. No holdings check, and hidden texts that hit no pattern are listed without a model's reading.

    Models
    • decosa-api check-their-brief (decosa_api/verticals/theirbrief, on the filing pre-flight engine)
    Hardware
    Any Linux or macOS machine with Python 3.11+
    Quality evidence
    • Hidden-text scan and reporter/existence checkssame as standard (code, no model)docs/evals/check-their-brief.md
    • Hidden instructions without the model's readingnot measured yet
    Latency
    estimate: seconds plus the lookups; not timed separately.
    Verification
    No proof yetSelf-host onlyNo model call, so no model receipts; the record is still signed by the box.
  • In the hosted demo

    Standard

    one GPU for the judge (hosted demo)

    Adds the holdings check and the model's reading of hidden texts with Qwen3.8-27B, one receipted call each, and scanned filings through the document reader. This is what the hosted API runs.

    Models
    • decosa-api check-their-brief (decosa_api/verticals/theirbrief, on the filing pre-flight engine)
    • Qwen3.8-27B (NVFP4)
    • Decosa document reader (Docling layout + PaddleOCR-VL-1.6)
    Hardware
    1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate, without the document reader)
    Quality evidence
    • Citations to cases that do not exist, caught as a finding (strict) / as a finding or a look (lenient)24/31 (77%) / 25/31 (81%)docs/evals/check-their-brief.md, LePhantomCite held-out test (390 real brief excerpts, CC BY 4.0), run once
    • Case name and cite that belong to two different cases34/68 (50%) strict / 51/68 (75%) lenientsame
    • Quotations with a word swapped31/45 (69%) strict / 36/45 (80%) lenientsame
    • Wrong pin cites and misstated holdings (strict)2/55 and 4/131: not reliably caughtsame; the judge marks most misstated holdings 'look at', as it does 25% of holdings in error-free excerpts
    • False findings on error-free excerpts49 of 950 checked items (5.2%) as run; 33 (3.5%) with the post-test fixes simulated on the same outputssame
    • Hidden instructions aimed at AI tools (blind-written texts hidden in 12 real briefs, 13 techniques)71/74 flagged (the 3 misses were CJK text the planting tool could not encode); benign hidden texts flagged 3/75 as run, 1/75 in the regression re-run after the pattern fixdocs/evals/check-their-brief.md, section 3 (re-run after the real-PDF fixes: same numbers)
    • Instructions written in plain view2/6 flagged, 0/5 benign flagged: weaksame
    • Real filings courts criticised for invented citations (31 filings, 161 problems from the orders; CourtListener cache-only)Invented citations: 0 of 78 called fine; 30 findings and 3 looks among the 34 it could look up; 32 not checked because CourtListener was unavailable. All problems: 37/161 findings, 74/161 findings or looksdocs/evals/check-their-brief.md, section 2 (Charlotin database + RECAP; not held out from the fixes it exposed)
    • Same 54 filings with the local citation index, CourtListener off (gateway, 28 Sep 2026)Invented citations: 0 of 78 called fine; 32 findings and 63 findings or looks; 4 not checked (was 32). Case citations not checked: 97 of 2,030 (was 563). Uncriticised briefs: 47 findings on 2,006 checked items (2.3%). p50 17 s a filing.docs/evals/citation-index.md (not held out from the lookup fixes made during that run)
    • Findings on 23 uncriticised real briefs66 of 1,645 checked items (4.0%) as run; 49 of 1,666 (2.9%) after the fixes (direct-route re-run); many are real miscites in those briefs or quotes of the other side's invalid citesdocs/evals/check-their-brief.md, section 2
    Latency
    measured on our server, shared card, gateway route (see Latency).
    Verification
    Proof: strongEvery model call is a separate gateway call with a signed receipt; the signed record lists them.
  • Needs more compute

    Wanted: the best setup

    a GLM-5.3-Flash holdings judge on your own hardware

    A stronger open judge from another family for the holdings check, where the 27B is weakest here. Your notes can be privileged, so it runs on your hardware, never on community providers. Not served yet.

    Models
    • decosa-api check-their-brief (decosa_api/verticals/theirbrief, on the filing pre-flight engine)
    • Qwen3.8-27B (NVFP4)
    • GLM-5.3-Flash (NVFP4)
    Hardware
    Your own hardware: 2x 96 GB cards (NVFP4, about 170 GB, estimate) or a Mac with 192 GB or more (MLX 4-bit, estimate).
    Quality evidence
    • This eval, same protocolnot measured yet
    Latency
    not measured.
    Verification
    No proof yetSelf-host onlyNot a gateway model yet.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.

Also runs on

  • Our citation-support model as a second opinion on holdingsdecosa-citation-support-modernbert-large (own model M1, prototype; Apache-2.0)not servedTurns a holding the judge calls unsupported into a finding only when our own fine-tuned model agrees: on held-out pairs 80 of 129 misstated holdings at 7 false flags in 372, against 26 at 22 for the judge alone. Prototype weights, off by default (DECOSA_THEIRBRIEF_SUPPORT_URL). Hardware: CPU (8 threads, about 0.9 s per holding) or any small GPU.

We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.

Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Checker: hidden-text scan (rendered page against text layer, metadata, comments, invisible Unicode), citation parsing, impossible-reporter check, lookups with a search trail, quotation match, Rule 5.2 scan, findings memo, signed record (no model; CPU)decosa-api check-their-brief (decosa_api/verticals/theirbrief, on the filing pre-flight engine)
0 GBProof: partial
Judge: one call per holding checked against the opinion, and one call for which hidden or embedded texts speak to AI tools (texts quoted as data)Qwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Document reader for scanned filings (no text layer): layout plus OCR, then the same checksDecosa document reader (Docling layout + PaddleOCR-VL-1.6)PaddlePaddle/PaddleOCR-VL-1.6 on Hugging Face (opens in a new tab)
Proof: partialIn the hosted demo
Citation-support model (own model M1, prototype): does the opinion's passage support the brief's sentencedecosa-citation-support-modernbert-large (own model M1, prototype; Apache-2.0)decosaai/decosa-citation-support-modernbert-large on Hugging Face (opens in a new tab)
395M · 0 GBProof: partialSelf-host only
Stronger holdings judge (wanted)GLM-5.3-Flash (NVFP4)
about 170 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /theirbrief/info, /theirbrief/samples; POST /theirbrief/parse (PDF, DOCX or text; hidden-text scan; scans to the document reader), /theirbrief/check (SSE or JSON); GET /theirbrief/runs/{id}/export?format=md|docx|record. Filings stay in memory for one hour, never on disk.

  • vLLM (judge):8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind the gateway (hosted) or called directly (self-host).

Hardware

  • Any CPU, no GPU Fits

    Lite tier (DECOSA_THEIRBRIEF_JUDGE=off): hidden-text scan, the injection patterns, existence, impossible reporters and quotations; no holdings check. Estimate: not timed separately.

  • 1x RTX 5090 32 GB Fits

    Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: not run here for this tool.

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the hosted demo and the evals ran on this card, shared with other work.

Latency per lane

  • hidden-text scan of a PDF (render every page, compare with the text layer)150 ms

    Measuredmeasured on our server 2026-09-28: 0.08-0.15 s for 1-2 pages; 12 real briefs of 20-60 pages scanned in the planted eval

  • LePhantomCite excerpt (about 2,500 characters), end to end9.9 s

    Measuredmeasured on our server 2026-09-28, held-out test (390 excerpts, two at a time, cold lookups): p50 9.9 s, p95 61.2 s

Notes

  • Hidden text is measured, not guessed: a character counts as hidden when its box on the rendered page shows nothing, or when it is drawn in the page's colour over other text, in invisible render mode or at zero opacity. Tiny (under 3 pt) and off-page text, metadata, XMP, annotations, bookmarks, embedded files, JavaScript, Word hidden/white/tiny runs, tracked deletions and comments are read too.
  • An instruction to AI tools is flagged by strong patterns in code (never overruled) and by one receipted model call that reads the texts as quoted data. Hidden text is blanked out of the filing before the citation checks, so it never reaches the judge inside a sentence it is asked about.
  • Before a citation becomes 'not found', the case name is searched too; a real case at another cite is 'their cite is wrong', a Westlaw or Lexis cite no free source has is a 'look at', and a failed search is never reported as run.
  • Text under a black box (a failed redaction) is counted, never shown, with a note on Model Rule 4.4(b).
  • No citator: it does not say whether a case is still good law.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

check-their-brief/assemble-prompt.md139 lines
# Assemble "Check the other side's brief" on this machine

You are setting up a checker for the OTHER side's court filings. It takes their filing (a PDF, a Word file or text) and
returns findings for a reply: which cited cases, statutes and rules the public sources do not have (with the search
trail), which quotations differ from the opinion, which holdings the opinion does not contain, what the file hides
(text the page does not show, tiny or off-page text, invisible characters, metadata, comments) and whether any of that
hidden text is aimed at AI tools. It ends with a Word findings memo and a record signed by this box's own key. Work
step by step, show me each command before you run anything with `sudo`, and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/check-their-brief.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py check-their-brief` (the api image carries the same bundle under /app/rehearsal/check-their-brief/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py check-their-brief --bundle check-their-brief.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the PDF is read and its hidden-text scan runs", "the scan finds the hidden texts before any model runs", "the decision is: findings"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Model: Qwen3.8-27B (Apache-2.0) as the judge for holdings and for the injection check. Everything else is CPU code:
  decosa-api (AGPL-3.0-or-later), eyecite and reporters-db (BSD-2-Clause) for citations, pypdfium2 (Apache-2.0 or BSD-3-Clause,
  PDFium BSD-3-Clause) to render pages and read the text layer, pypdf (BSD-3-Clause) for metadata and annotations,
  pdfplumber (MIT), and poppler's `pdftotext` (GPL-2.0, run as a separate program).
- A filed brief is a public record; your notes about it, and any paper that is served but not filed, are not. Bind every
  port to 127.0.0.1. Filings and results stay in memory for one hour and are never written to disk or logs. The only
  thing written to disk is a cache of public-source answers (the citation asked, the public opinion text), never the
  filing.
- Decide with me which mode to run:
  - **Offline** (`DECOSA_PREFLIGHT_OFFLINE=1`): nothing leaves this machine. Authorities are checked only against
    opinion texts I send with the request, or a local mirror of the Caselaw Access Project (step 2).
  - **Online lookups** (default): only citations (volume, reporter, page; a case name when a cite is not found; statute
    sections) go to static.case.law, courtlistener.com, uscode.house.gov, ecfr.gov and law.cornell.edu. A free
    CourtListener account token (`DECOSA_PREFLIGHT_COURTLISTENER_TOKEN`) avoids its anonymous rate limit.
- Be honest about what it does: "not found" means the public sources searched do not have it, not that the case is
  invented. It is not legal advice; a lawyer verifies every finding before using it.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB for the judge (Qwen3.8-27B NVFP4 is about 20 GB of weights plus KV cache).
   Driver 570 or newer; Blackwell runs NVFP4, older cards use `Qwen/Qwen3.8-27B-FP8`.
2. No GPU? Set `DECOSA_THEIRBRIEF_JUDGE=off`: the hidden-text scan, existence, reporter and quotation checks and the
   pattern layer of the injection check still run; holdings are "not checked".
3. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
   from the official repositories after asking me, then run
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
4. Disk: about 30 GB free for the model and images.

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out the newest tag that contains
  `decosa_api/verticals/theirbrief/` (`main` until one does), and build `docker/api/Dockerfile` (it installs the
  `theirbrief` extra and poppler). Without Docker: `pip install ".[theirbrief]"` and `apt install poppler-utils`.
- `vllm/vllm-openai:v0.29.0` for the judge; weights `nvidia/Qwen3.8-27B-NVFP4`.
- Scanned filings (no text layer) need the document reader service; without it, set `DECOSA_THEIRBRIEF_OCR=off` and
  scans are refused with a clear message. Born-digital PDFs and Word files need nothing extra.
- Optional offline case-law mirror: copy the reporters you need from `https://static.case.law/` (`ReportersMetadata.json`
  at the root, and for each `<reporter>/<volume>/` its `CasesMetadata.json` and `html/`) into `~/decosa/cap/`, serve it
  read-only (`python -m http.server 8480 --bind 127.0.0.1 -d ~/decosa/cap`) and set `DECOSA_PREFLIGHT_CAP_BASE`.

## 3. docker-compose.yml
Write this in `~/decosa/theirbrief/`:

```yaml
services:
  llm:
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
              "--enable-prefix-caching"]
    ports: ["127.0.0.1:8114:8000"]
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>
    ports: ["127.0.0.1:8445:8445"]
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_BUDGET_LLM_TOKENS: "60000"         # per demo session; about 170 generated tokens per holding checked
      DECOSA_THEIRBRIEF_OCR: "off"              # on only if the document reader service runs next to it
      DECOSA_THEIRBRIEF_MAX_CONCURRENT: "2"
      # DECOSA_PREFLIGHT_OFFLINE: "1"           # uncomment to send nothing outside (see step 0)
      # DECOSA_PREFLIGHT_COURTLISTENER_TOKEN: "" # a free CourtListener token, if you have one
    volumes: ["decosa-data:/data"]
    depends_on: { llm: { condition: service_healthy } }
    healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/theirbrief/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
  decosa-data:
```

State (keys, receipts, this box's signing key, the public-source cache) lives in the named volume `decosa-data`, not a
host folder: the image runs as uid 10001, and a root-owned host folder stops the api with a PermissionError on
`/data/keys.sqlite`. Start everything: `docker compose up -d`. The signing key is created on first start under
`/data/attest/` (mode 0600): back it up, never print it. Records are signed with it: an attestation by you, the operator.

## 4. Smoke test
1. `curl -s localhost:8445/theirbrief/info | jq '{checks, offline, judge, ocr}'` lists the six checks and the mode.
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"check-their-brief"}' | jq -r .token)`.
3. Run the fictional sample (a PDF shipped with the api):
   `curl -s -XPOST localhost:8445/theirbrief/check -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample_id":"demo-opposition","stream":false}' > run.json`.
   Expect `decision: "findings"` and, among the problems: three hidden instructions aimed at AI tools (white text on
   page 1, 1-point text on page 2, the Keywords property), Hollis v. Brentwood School District at 925 F.3d 1339 not
   found (that page is inside J.D. v. Azar), the T.L.O. quotation with "any" for "strict", the Redding holding, and the
   student's full birth date. With lookups off, case items are "unverified" unless you pass opinions in `authorities`.
4. Your own file: `curl -s -XPOST 'localhost:8445/theirbrief/parse?name=opp.pdf' -H "authorization: Bearer $T" -H 'content-type: application/pdf' --data-binary @opp.pdf | jq '{upload_id, chars, hidden: .hidden.findings | length}'`,
   then check it with `{"upload_id": "..."}`.
5. Stream with `-H 'accept: text/event-stream' -N` and `"stream": true`: `ready`, `stage`, `item`, a `receipt` after
   each model call, then `report`, `done` and `budget`.
6. `jq '{record: .report.record}' run.json | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @-`
   must say `ok: true`. Change one item's status in the record and verify again: it must fail and name that entry.
7. The Word memo: `curl -s localhost:8445/theirbrief/runs/<run_id>/export?format=docx -H "authorization: Bearer $T" -o findings.docx`.
8. Time it and tell me. On our shared RTX PRO 6000 the fictional sample took about 30 s with lookups on (28 Sep 2026);
   most of the time is lookups and judge calls.

## 5. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`, or call `POST /theirbrief/check` from
your document system when an opposing filing arrives and keep the Word memo and the signed record in the matter file.
Contract: `API_CONTRACT.md`, section "Check their brief".

Off by default. It serves other people's requests on this GPU; never on a box that holds client work. Do not enable it
without my explicit yes; the guide is at `/provide` on the site.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Check opposing brief citations and hidden text before you reply: every citation looked up in public sources with a search trail, quotations matched word for word, and the PDF scanned for text the judge cannot see but an AI tool reads.
Who it's for
Litigators and paralegals answering an opposition or reply, especially against AI-drafted or pro se filings.
Where it runs
Hosted or self-host; the confidential tier for anything not yet public
Key numbers

On 390 held-out brief excerpts it caught 24 of 31 non-existent citations and 31 of 45 swapped-word misquotes as findings; misstated holdings are not reliably caught.

  • 24 / 31 (77%) Non-existent citations caught (strict) (held out, n = 31)
  • 34 / 68 (50%) Case name and cite from two different cases (strict) (held out, n = 68)
  • 31 / 45 (69%) Swapped-word misquotations (strict) (held out, n = 45)
  • 8.3 s Median end-to-end run, hosted (QA sweep 2026-09-30)
All results, datasets and caveats
Models
Qwen3.8-27B (holdings against the opinion; which hidden texts speak to AI tools) · document reader for scanned filings
Where
Hosted or self-host; the confidential tier for anything not yet public
Checks
Receipt per model call; search trail per citation; signed record
Industry
Legal
Output
Signed record or verdict · Notes, reports and drafts
Data
Personal data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

Can I check opposing brief citations for fake cases?

Yes. Each citation is looked up in a local index of CourtListener's and the Caselaw Access Project's public data, then CourtListener's search for anything newer, with the case name searched too before anything is called missing. In 31 real filings courts criticised for invented citations, none of the 78 invented citations was called fine. On 390 held-out excerpts of real briefs with injected errors, it caught 24 of 31 non-existent citations as findings. A miss is worded 'not found in the public sources searched', never 'fake'.

What is hidden text in a court filing?

Text a reader of the page cannot see but a text extractor and every AI tool reads: white text on white, 1-point text, text off the page, invisible characters, metadata and comments. The check renders each page and compares it with the text layer. In a blind test, 71 of 74 hidden instructions aimed at AI tools were flagged and 3 of 75 harmless hidden texts were.

Does it tell me whether a case is still good law?

No. There is no citator, so it does not report negative treatment or later history. Westlaw- and Lexis-only decisions, many unpublished orders and most state codes come back as 'look at' or 'not checkable', never as fine.

Does it catch misstated holdings?

Not reliably yet. The open judge marked only 4 of 131 misstated holdings as findings on the held-out test; most came back 'look at', like a quarter of accurate holdings. A prototype of our own citation-support model, combined with the judge, caught 80 of 129 at 7 false flags in 372 on held-out pairs, and is not switched on yet.

What does a check cost?

About $0.013 of model calls at list price for a short brief (7 calls for the fictional two-page opposition) and $0.012 for a 2,500-character excerpt. A paralegal's hand cite-check is an estimated 5-8 minutes per authority.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Check their brief

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.