Skip to content
decosa
LiveHostedSelf-hostMacSelf-host first for real data

Chart claim support in a patent spec

A claim chart naming the paragraphs that support each claim element, plus an issue list of missing antecedents and broken dependencies.

Held-out test93.4%Claim elements given the right supported-or-not call (held-out test)
On production7.8 smedian on production (2026-09-25); slower when the service is busy
List price~$0.19 per 100 claim elementsmeasured, at list price

Built on: Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; parsing and the code checks run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the patent claim-support checker API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real client material belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
patent-claim-support

Use the hosted API

# Decosa Patent claim-support checker: use the hosted API

You are wiring Decosa's patent claim-support checker into this project. It takes a patent specification and its
claims, and returns a claim chart that maps every claim element to the specification paragraphs that describe it
(cited by number, like [0012]) with a support strength, plus an issue list: elements with no support, "the X" with no
earlier "a X" (antecedent basis), dependent claims that point at a missing or later claim, and claim words the
specification never uses. It ends with a signed record. Each model call has its own signed receipt. Use only what is
listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- The hosted API is for published patents and applications and for testing your integration. An unfiled invention is
  confidential and may be export-controlled: for unpublished text use the self-host prompt instead. Never send an
  unfiled application here.
- It is evidence for a registered practitioner, not a legal opinion. Show every result as something to check.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "patent-claim-support"}` returns `{"token", "expires_at", "budget"}`.
   a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); each session has a token allowance (its `budget`) (about 120 claim elements). Over a limit: HTTP 429 with `Retry-After`.
3. One check at a time per demo token (409 otherwise).

## Endpoints
- `POST /patent/check` (token). Body:
  `{"spec": "...", "claims": "1. A device comprising ...\n2. The device of claim 1, ...", "only_claims"?: [1, 2, 3], "support"?: true, "title"?: "...", "stream"?: true}`,
  or `{"application": "..."}` (one text, cut at "What is claimed is:"), or `{"sample": "ladar-planted"}`.
  - Each claim starts on its own line with its number. Paragraphs numbered like `[0012]` are cited by that number.
  - Limits: 150,000 characters of specification, 40,000 of claims, 60 claims, 120 elements per request (use
    `only_claims` for more). `"support": false` runs the code checks only (no model call, no budget).
  - With `"stream": true` (or `Accept: text/event-stream`) it streams `ready` (claims split into elements with
    offsets, paragraphs, the code issues), a `receipt` after each model call, an `element` event per element
    (`{id: "2.1", strength, paragraphs: [{n, role}], reason}`), then `report`, `done` and `budget`. With
    `"stream": false`: one JSON object `{run_id, claims, elements, receipts, report, budget}`.
  - Strengths: `strong` (supported), `moderate` (supported, some doubt), `weak` (partly: a limitation is missing),
    `none` (no support found), `conflict` (the specification says otherwise), `error` (the call failed; nothing is guessed).
  - `report.issues`: `[{id, claim, kind: support|antecedent|dependency|multiple_dependency|numbering|term, severity: error|warning, message, offset?}]`.
- `POST /patent/parse` (token): the same body, code checks only, always JSON.
- `POST /patent/runs/{run_id}/decisions` `{"decisions": [{"target": "2.1" | "I3", "decision": "agree"|"disagree", "reviewer", "note"?}]}`: the practitioner's calls.
- `GET /patent/runs/{run_id}/export?format=csv` (claim chart), `md` (issues and chart), `record` (signed), `ledger` (signed, with the decisions).
- `POST /record/verify` (no token) `{"record": {...}}` → `{ok, summary, bad}`.
- `GET /patent/info`, `GET /patent/samples`, `GET /patent/samples/{id}`, `GET /attest/signing-key` (no token).

## Example: check a published patent and save the chart (Python, `pip install httpx`)
```python
import httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
s = httpx.get(f"{API}/patent/samples/filter").json()        # your own published patent has the same shape
r = httpx.post(f"{API}/patent/check", headers=H, timeout=600,
               json={"spec": s["spec"], "claims": s["claims"], "stream": False})
r.raise_for_status()
run = r.json()
for x in run["report"]["issues"]:
    print(x["severity"], x["kind"], x["message"])
pathlib.Path("claim-chart.csv").write_text(httpx.get(f"{API}{run['export']['csv']}", headers=H).text)
```

## Honest limits
- The support map finds the paragraphs that describe an element well, but misses many elements a practitioner would
  call only partly supported (a limitation described only in another embodiment, for example); see the measured rates
  on the Stack tab.
- The antecedent check is lenient on purpose; about half of what it flags on granted claims is a genuine slip, and a
  clean result is not proof of definiteness.
- Text only: paste or send the specification and claims; no DOCX, PDF or figures.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Patent claim-support checker: run it yourself (containers)

You are setting up the Decosa patent claim-support checker on this machine, so unfiled inventions never leave it. It
maps every claim element to the specification paragraphs that describe it, checks antecedent basis, claim references
and claim terms, and signs a record. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/patent-claim-support.zip (15 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py patent-claim-support` (the api image carries the same bundle under /app/rehearsal/patent-claim-support/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py patent-claim-support --bundle patent-claim-support.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "planted: claim 3 depends on a later claim (12)", "planted: claim 7 depends on a claim that does not exist (30)", "planted: "the clock signal" in claim 5 has no antecedent"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
   check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, with prefix caching) and the `api` service. For the `api`
   service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
   `DECOSA_PATENT_MAX_ELEMENTS=400` and bind every port to 127.0.0.1. Nothing in this tool needs the internet after
   the weights are downloaded.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/patent/info` shows `"route": "direct"`; `GET /attest/signing-key` shows
   this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"patent-claim-support"}` and run
   `POST /patent/check {"sample": "ladar-planted", "stream": false}`. Expect 13 elements; 2.1 (a cryocooler the
   specification never mentions) and 5.1 (its paragraphs were deleted) marked `none`; antecedent issues in claims 1, 5
   and 6; dependency errors for claims 3 and 7; every receipt `attested`. Then `POST /record/verify` with
   `report.record`: `ok` must be true.
6. Report back: the public key and key id, the totals, the issues and how long the run took.

Off, and keep it off on a box that holds unfiled inventions: joining serves other people's requests on this GPU.
If I ask for it later, on a separate machine, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/patent-claim-support-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMlite tierRuns with a smaller tier

    The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (32 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (32 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"patent-claim-support"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py patent-claim-support

Download the mock-data bundle (15 KB, 13 checks)expected.json

Two public-domain US patents as plain text. First, code-only checks on claims 1 to 8 of US 10,000,000 B2 with planted defects: claim 3 depends on a later claim, claim 7 on a claim that does not exist, "the clock signal" has no antecedent and "the data highway" is a term the specification never uses; all must be found. Then claim 1 of US 11,043,724 B2 is checked element by element against its specification: every element gets a receipted verdict, at least two are supported with a cited paragraph, "opening" is flagged as a claim term the specification never uses, and the signed record verifies.

What the rehearsal checks
  • planted: claim 3 depends on a later claim (12)
  • planted: claim 7 depends on a claim that does not exist (30)
  • planted: "the clock signal" in claim 5 has no antecedent
  • planted: "data highway" (claim 6) is a term the specification never uses
  • planted new matter: the Stirling cryocooler in claim 2 is not in the specification
  • claim 1 of the filter patent is split into four elements
  • no element's support check errors
  • element 1.3 is supported
  • element 1.4 is supported with a cited paragraph
  • "opening" is flagged as a claim term the specification never uses
  • the signed record verifies
  • a record with its warning count changed no longer verifies
  • every model call has a signed receipt

Licence: Public domain: US patent text from the USPTO (via the Google Patents public pages); the defects in the US 10,000,000 B2 copy were planted for Decosa. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa Patent claim-support checker: run it yourself (containers)

You are setting up the Decosa patent claim-support checker on this machine, so unfiled inventions never leave it. It
maps every claim element to the specification paragraphs that describe it, checks antecedent basis, claim references
and claim terms, and signs a record. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/patent-claim-support.zip (15 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py patent-claim-support` (the api image carries the same bundle under /app/rehearsal/patent-claim-support/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py patent-claim-support --bundle patent-claim-support.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "planted: claim 3 depends on a later claim (12)", "planted: claim 7 depends on a claim that does not exist (30)", "planted: "the clock signal" in claim 5 has no antecedent"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
   check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, with prefix caching) and the `api` service. For the `api`
   service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
   `DECOSA_PATENT_MAX_ELEMENTS=400` and bind every port to 127.0.0.1. Nothing in this tool needs the internet after
   the weights are downloaded.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/patent/info` shows `"route": "direct"`; `GET /attest/signing-key` shows
   this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"patent-claim-support"}` and run
   `POST /patent/check {"sample": "ladar-planted", "stream": false}`. Expect 13 elements; 2.1 (a cryocooler the
   specification never mentions) and 5.1 (its paragraphs were deleted) marked `none`; antecedent issues in claims 1, 5
   and 6; dependency errors for claims 3 and 7; every receipt `attested`. Then `POST /record/verify` with
   `report.record`: `ok` must be true.
6. Report back: the public key and key id, the totals, the issues and how long the run took.

Off, and keep it off on a box that holds unfiled inventions: joining serves other people's requests on this GPU.
If I ask for it later, on a separate machine, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/patent-claim-support-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsPatent claim-support checker on GeForce RTX 5090: use the Standard · one 96 GB card (measured; hosted demo) tier

The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB (fits): Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for a 48,000-character specification (about 12k tokens). Estimate: same model and prompts as the measured card, not run here on a 5090.

Standard · one 96 GB card (measured; hosted demo): what changesuses estimates

  • Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Checker: decosa-api patent module (decosa_api/verticals/patent). CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Model: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Patent claim-support checker, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Patent claim-support checker on my hardware

Fetch https://decosa.ai/prompts/patent-claim-support-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=patent-claim-support)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · one 96 GB card (measured; hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Checker: decosa-api patent module (decosa_api/verticals/patent), CPU
- Model: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/patent-claim-support-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh

Mac prompt for your coding agent

# Decosa Patent claim-support checker: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Patent claim-support checker on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/patent-claim-support.zip (15 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py patent-claim-support` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "planted: claim 3 depends on a later claim (12)", "planted: claim 7 depends on a claim that does not exist (30)", "planted: "the clock signal" in claim 5 has no antecedent"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Checker: claim parser, antecedent basis, claim references, claim terms, the claim chart, signed record and ledger (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured |
| Model: reads each claim element against the specification and names the supporting paragraphs | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py patent-claim-support`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 7.8 s · ~$0.020 per run · 13 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.

Measured cost to run: about $0.19 per 100 claim elements (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Verified on 2026-09-25: the image builds, the service starts healthy, the planted sample runs end to end on the direct route (13 attested calls, 3.4 s): claims 2 and 5 have no support, the antecedent, reference and term issues are found, the signed record verifies and fails when one strength is changed, the CSV export works, and no claim text reaches the logs. The model server's own startup was not re-verified (no new GPU load).

Known limits (5)
  • Labels are Claude's (an AI agent), not a registered practitioner's; antecedent flags were judged by separate Claude agents with a written rubric.
  • The support map finds supporting paragraphs well but flags only 5 of the 13 held-out elements the labels call partly supported; most misses are judgment calls such as a feature described only in another embodiment.
  • About half of the antecedent-basis flags on randomly drawn granted claims are genuine slips (36 of 66 held out); the rest are implicit or inherent references a practitioner would accept.
  • Run-to-run variation on the hosted gateway: the same element can come back partly supported in one run and unsupported in the next; both are flagged.
  • Text only (no DOCX, PDF or figures); specifications over 48,000 characters are read as the best-matching paragraphs, which can miss support far from the element's words.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; parsing and the code checks run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the patent claim-support checker API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real client material belongs on your own hardware.
The open stack

Every claim element mapped to the specification paragraphs that describe it, with antecedent-basis, claim-reference and claim-term checks, one receipt per element and a signed record.

Paste a specification and its claims. Code splits the claims into elements, checks that every "the" or "said" has an antecedent, that every dependent claim points at an earlier claim that exists, and that the claims use no word the specification never uses. An open model then reads each element against the specification, one receipted call per element, and names the paragraphs that describe it ([0012] style) with a support strength: supported, partly supported, or no support found. The result is a claim chart, an issue list and a signed record with hashes, verdicts, paragraph numbers and receipt ids but no claim text, plus a signed ledger of the practitioner's agree or disagree per element. It is evidence for a registered practitioner, who decides.

Deployment
Self-host first
Regulatory
A review aid for registered patent practitioners, not legal advice or a patentability opinion. The written-description and enablement requirements are in 35 U.S.C. 112(a) and MPEP 2163-2164; lack of antecedent basis is an indefiniteness question under 112(b) (MPEP 2173.05(e)); dependent and multiple dependent claims follow 112(d)-(e) and 37 CFR 1.75(c); claim terms must find clear support or antecedent basis in the description (37 CFR 1.75(d)(1)). A supported element here means the model found paragraphs that describe it, not that the claim is allowable. The USPTO's guidance on AI-based tools in practice (89 FR 25609, 11 Apr 2024; still the current practitioner guidance in the Federal Register as of 25 Sep 2026) applies the duty of candor and the signature rule (37 CFR 11.18) to AI-assisted papers: the signer must review them and relying on an AI tool alone is not a reasonable inquiry; it reminds practitioners to prevent disclosure of client information through AI tools (37 CFR 11.106) and warns that releasing controlled technical data to a foreign person can be a deemed export (15 CFR 734.13(b)); it says there is no general duty to tell the Office an AI tool was used unless asked, but information on AI use that is material to patentability or inventorship must be disclosed (the inventorship guidance it cites was revised on 28 Nov 2025, 90 FR 54636). Run unfiled applications on your own machine; the hosted demo takes published patents and applications only (US patent text is public domain). Model licence: Apache-2.0 (Qwen3.8-27B). Checked 25 Sep 2026.
Architecture
Text description

The specification (paragraphs numbered like [0012]) and the claims go to the checker, which runs on your own machine. Code splits the claims into elements and checks antecedent basis, claim references and claim words the specification never uses. Qwen3.8-27B (Apache-2.0) reads each element against the specification through the grounding checker's evidence and verdict parser, one call per element, and names the supporting paragraphs. Outputs: a claim chart with a support strength per element, an issue list, and a signed record and practitioner ledger with hashes, verdicts, paragraph numbers and receipt ids but no claim text. On the hosted route each model call gets a receipt that our gateway countersigns.

Architecture

At a glance

Data retention
The application is held in memory for the request; the run (chart, issues, no specification) for one hour, for the token or key that made it. Nothing is written to disk; logs carry counts only.
What leaves the box
Self-hosted: nothing. Hosted demo: the text goes to the model through our gateway, which is why the hosted demo is for published patents and applications only.
Model cost per application
A few cents per application on average at the gateway's list price (measured on granted patents); GPU time only when self-hosted.
Input
Specification and claims as text (or one pasted application cut at "What is claimed is:"). Paragraphs numbered like [0012] are cited by number. Hosted limits: 150,000 characters of specification, 60 claims, 120 elements per request; raise them on your own box.
Output
A claim chart (CSV or Markdown) with the supporting paragraphs and a strength per element, an issue list (support, antecedent basis, claim references, terms), a signed record and a signed ledger of the practitioner's decisions.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • In the hosted demo

    Lite

    code checks only, any CPU

    Antecedent basis, claim references and claim terms, with no model and no GPU (support: false, or POST /patent/parse). No support map.

    Models
    • decosa-api patent module (decosa_api/verticals/patent)
    Hardware
    Any CPU
    Quality evidence
    • Antecedent-basis flags that are genuine, on randomly drawn granted US claims (held out: rules frozen before the set was drawn)54.5% (36 of 66 flags; 55 patents, 891 claims)decosa-api docs/evals/patent-claim-support.md, 2026-09-25; each flag judged by a separate Claude agent with a written rubric, not a practitioner
    • Planted defects found by code in 6 test patents: antecedent errors / broken claim references / renamed terms11 of 11 / 12 of 12 / 6 of 6decosa-api docs/evals/patent-claim-support.md, 2026-09-25
    Latency
    measured: well under a second per application.
    Verification
    Proof: partialNo model calls, so no receipts; the record is signed by the instance.
  • In the hosted demo

    Standard

    one 96 GB card (measured; hosted demo)

    Qwen3.8-27B on an RTX PRO 6000 reads every element against the specification; the code checks run alongside. Also fits a 32 GB card (estimate).

    Models
    • decosa-api patent module (decosa_api/verticals/patent)
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX PRO 6000 96 GB
    Quality evidence
    • Elements where the model's supported-or-not call matches the labels, 6 held-out granted patents93.4% of 166 (100.0% of the 146 labelled as clear calls)decosa-api docs/evals/patent-claim-support.md, 2026-09-25; labels written by Claude (an AI agent) reading each specification, not by a practitioner
    • Supported elements where a cited paragraph is one the labels list / share of cited paragraphs that are100.0% / 93.8% (150 elements)decosa-api docs/evals/patent-claim-support.md, 2026-09-25
    • Elements the labels call only partly supported that the model also flags (recall) / flags that match a label (precision)5 of 13 / 5 of 8decosa-api docs/evals/patent-claim-support.md, 2026-09-25; most misses are limitations the labels mark as judgment calls, such as a feature described only in a different embodiment
    • Planted support removal: all paragraphs describing one element deleted, element flagged5 of 6 (the miss: a paragraph the removal left in still describes the element, checked by reading it; an earlier run missed a different patent the same way)decosa-api docs/evals/patent-claim-support.md, 2026-09-25
    • Dev patents (2, used to write the prompt): agreement / flag recall97.8% / 3 of 3decosa-api docs/evals/patent-claim-support.md, 2026-09-25
    Latency
    measured on our server, shared gateway: seconds to over a minute per whole granted patent, rising with its elements (the specification is served from the prefix cache).
    Verification
    Proof: strongHosted: every element's call has its own gateway-signed receipt, listed in the signed record.
  • Needs more compute

    Wanted: the best setup

    a GLM-5.3-Flash judge on your own hardware

    A stronger legal-reasoning model from another family for the partial-support calls the standard model misses (5 of 13 found). Specifications before filing stay on your hardware, never on community providers. Not served yet.

    Models
    • decosa-api patent module (decosa_api/verticals/patent)
    • Qwen3.8-27B (NVFP4)
    • GLM-5.3-Flash
    Hardware
    Your own hardware: 2x 96 GB cards (NVFP4, about 170-186 GB, unconfirmed) or a Mac with 192 GB or more (MLX 4-bit, 165 GB). Estimate; GLM's SGLang SM120 build hung on our server.
    Quality evidence
    • This eval, same protocolnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.

Also runs on

  • A second judge on one cardNemotron-3-Super-120B-A12B (NVFP4)not servedA cross-family second judge that fits beside the 27B, so we would host it ourselves. sm_120 support unconfirmed. Hardware: 1x RTX PRO 6000 96 GB (80 GB of weights).

We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.

Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Checker: claim parser, antecedent basis, claim references, claim terms, the claim chart, signed record and ledger (no model; CPU)decosa-api patent module (decosa_api/verticals/patent)
0 GBProof: partial
Model: reads each claim element against the specification and names the supporting paragraphsQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Stronger open model for the support map (wanted)GLM-5.3-Flash
about 170 GB (estimate)No proof yetSelf-host only
One-card second judge from another model familyNemotron-3-Super-120B-A12B (NVFP4)nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 on Hugging Face (opens in a new tab)
120B (12B active) · about 80 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

  • USPTO patent full text (via the Google Patents public page) (opens in a new tab)US patent text is a US government work in the public domain

    The demo samples and the eval: 8 granted patents with labelled claim elements, and 213 randomly drawn granted patents for the antecedent-basis check.

  • scripts/patent_eval.py and docs/evals/patent-claim-support.mdApache-2.0

    Runs the support-map eval against the labels, plants defects in six test patents, scores the antecedent check on four rounds of granted claims, and writes the metrics and cost.

  • POST /record/verifyApache-2.0

    Checks the signed record or ledger and names the first entry that was changed. The console also verifies it in your browser.

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /patent/info, /patent/samples; POST /patent/check (SSE or JSON), /patent/parse (code checks only); POST /patent/runs/{id}/decisions; GET /patent/runs/{id}/export?format=csv|md|record|ledger. Keeps the application in memory for the request and runs for one hour, never on disk; logs counts only.

  • vLLM:8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).

Hardware

  • Any CPU (code checks only) Fits

    Measured: the parser and all code checks for a 20-claim application take well under a second on one core; no GPU.

  • 1x RTX 5090 32 GB Fits

    Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for a 48,000-character specification (about 12k tokens). Estimate: same model and prompts as the measured card, not run here on a 5090.

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the eval, the hosted demo and the self-host check ran on this card, shared with other services the whole time.

Latency per lane

  • 13 elements (the planted demo sample), hosted gateway route7.8 s

    Measuredmeasured on our server 2026-09-25: 5.9-9.2 s over 6 runs of the planted sample with the gateway lightly loaded; 36 s once while other evaluation jobs loaded the shared gateway

  • 13 elements, self-hosted (clean clone, container build, direct route)3.4 s

    Measuredmeasured on our server 2026-09-25, one run

  • A whole granted patent (17-50 elements), hosted gateway route18.8 s

    Measuredmeasured on our server 2026-09-25: 15-88 s over the 6 test patents, 6 element calls in flight, shared gateway

Notes

  • The claims never count as their own support: the model reads only the description. Original claims are part of the disclosure as filed, so an element supported only by an original claim is shown as unsupported for the attorney to judge.
  • The antecedent check is deliberately lenient: shortened references ("the beads" after "a plurality of polymer beads"), verb forms ("the comparison" after "comparing"), inherent properties ("the surface of"), idioms, superlatives, spelling slips and defined acronyms become notes, not issues. It still misses nothing we planted, and about half of what it flags on granted claims is a genuine slip.
  • A line that only opens a list ("the laser source further comprises:") is shown but not judged; its sub-elements are.
  • Specifications up to 48,000 characters go to the model whole; longer ones are read as the best-matching paragraphs and their neighbours (BM25 from the grounding checker), which can miss support far from the element's words.
  • Not included yet: drafting a missing paragraph from an invention disclosure, reference numerals against the figures, and DOCX or PDF input (paste text).
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

patent-claim-support/assemble-prompt.md130 lines
# Assemble the Decosa patent claim-support checker on this machine

You are setting up a claim-support checker for a patent practice. It takes a specification (paragraphs numbered like
[0012]) and its claims, and returns a claim chart that maps every claim element to the paragraphs that describe it,
with a support strength and the model's reason, plus an issue list: elements with no support, "the X" with no "a X"
(antecedent basis), dependent claims that point at a missing or later claim, and claim words the specification never
uses. It seals the result in a record signed by this box's own key. Work step by step, show me each command before you
run anything with `sudo`, and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/patent-claim-support.zip (15 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py patent-claim-support` (the api image carries the same bundle under /app/rehearsal/patent-claim-support/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py patent-claim-support --bundle patent-claim-support.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "planted: claim 3 depends on a later claim (12)", "planted: claim 7 depends on a claim that does not exist (30)", "planted: "the clock signal" in claim 5 has no antecedent"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Model: Qwen3.8-27B (Apache-2.0). Everything else is CPU code in decosa-api (AGPL-3.0-or-later): the claim parser, the
  antecedent, reference and term checks, and the grounding checker it builds on.
- An unfiled invention is confidential, and its technical data may be export-controlled. Bind every port to
  127.0.0.1, and do not send the text to any hosted API. The service
  keeps the application in memory for the request and the run for one hour, writes nothing to disk and logs counts
  only; keep it that way.
- It is evidence for a registered practitioner, not a legal opinion. Written description (35 U.S.C. 112(a)) is a
  legal judgment, and whoever signs the filing must have reviewed it (USPTO guidance on AI tools, 89 FR 25609,
  11 Apr 2024). Say so wherever you show results.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 is about 20 GB of weights plus KV cache; a
   48,000-character specification goes to the model whole, about 12k tokens). Driver 570 or newer; Blackwell cards run
   NVFP4, older cards use the FP8 weights.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
   from the official Docker and NVIDIA repositories after asking me, then run
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 30 GB free for the model and images.

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a release that contains
  `decosa_api/verticals/patent/`, and run `docker build -f docker/api/Dockerfile -t decosa-api:local .`
- `vllm/vllm-openai:v0.29.0` for the model; weights `nvidia/Qwen3.8-27B-NVFP4` (or `Qwen/Qwen3.8-27B-FP8`).

## 3. docker-compose.yml
Write this in `~/decosa/patent/`:

```yaml
name: decosa-patent
services:
  llm:
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "65536",
              "--enable-prefix-caching"]
    ports: ["127.0.0.1:8114:8000"]
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>     # or decosa-api:local
    ports: ["127.0.0.1:8445:8445"]
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_BUDGET_LLM_TOKENS: "200000"        # per session; an element uses about 70 generated tokens
      DECOSA_PATENT_MAX_ELEMENTS: "400"         # per request on your own box (the hosted limit is 120)
      DECOSA_PATENT_FULL_CHARS: "48000"         # longer specifications are read as the best-matching paragraphs
    volumes: ["decosa-data:/data"]                # a named volume: the image runs as uid 10001, so a root-owned bind mount fails
    depends_on: { llm: { condition: service_healthy } }
    healthcheck:
      test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/patent/info', timeout=4)"]
      interval: 30s
      retries: 10
volumes:
  decosa-data: {}
```

Prefix caching matters: every element's prompt starts with the same specification, so the model server computes it
once per application. On first start the api service creates this box's Ed25519 key in the `decosa-data` volume
under `attest/` (mode 0600). Back the volume up and never print the key. Records and model calls are signed with it:
an attestation by me, the operator, not a proof.

## 4. Smoke test
1. `curl -s localhost:8445/patent/info | jq '{checks: (.checks | keys), route: .model.route, limits}'` shows route `direct`.
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"patent-claim-support"}' | jq -r .token)`.
3. Code checks only, instant: `curl -s -XPOST localhost:8445/patent/parse -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample":"ladar-planted"}' | jq '.issues[] | {kind, claim, message}'`.
   Expect antecedent issues in claims 1, 5 and 6, a `term` issue for "highway" (claim 6), and `dependency` errors
   for claim 3 (depends on the later claim 12) and claim 7 (claim 30 does not exist).
4. The full check: `curl -s -XPOST localhost:8445/patent/check -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample":"ladar-planted","stream":false}' > run.json`.
   Expect 13 elements, `jq '[.elements[] | {id, strength}]' run.json`: 2.1 and 5.1 are `none` (their paragraphs
   [0023]-[0024] were deleted from this sample), the rest `strong` with paragraphs cited, and every element has a
   receipt with status `attested`.
5. Stream it with `-H 'accept: text/event-stream' -N` and `"stream": true`: `ready`, a `receipt` and an `element`
   event per element, then `report`, `done` and `budget`.
6. `jq '{record: .report.record}' run.json | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @-`
   must say `ok: true`. Change one `strength` in the record's `support` entries and verify again: it must fail.
7. Exports: `curl -s "localhost:8445/patent/runs/$(jq -r .run_id run.json)/export?format=csv" -H "authorization: Bearer $T"`
   is the claim chart; `format=md` is the issue list and chart.
8. Time it and tell me. On our RTX PRO 6000, shared with other work, this sample took 5 s on the direct route (13
   calls); through the hosted gateway it took about 8 s.

## 5. Point your drafting workflow at the local API
Send each application to `POST /patent/check` (`{spec, claims}` or `{application}` with the claims after "What is
claimed is:"; `only_claims` for a subset). Put the red and amber elements and the error issues in front of the
drafting attorney first, record their calls with `POST /patent/runs/{id}/decisions`
(`{"decisions": [{"target": "2.1" or "I3", "decision": "agree"|"disagree", "reviewer", "note"?}]}`), and keep the
signed ledger (`export?format=ledger`) in the file. For the site, set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445`
in `.env.local`. Contract: `API_CONTRACT.md`, section "Patent claim-support checker".

Off, and it should stay off on a box that holds unfiled inventions: joining serves other people's requests on this
GPU. Only on a separate machine, and only with my explicit yes, follow the provider guide at `/provide` on the site.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

8 laws, rules and guidance pages cited; 8 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Every claim element mapped to the specification paragraphs that describe it, with antecedent-basis, claim-reference and claim-term checks, one receipt per element and a signed record.
Who it's for
Teams in legal.
Where it runs
Self-host for unfiled inventions; hosted for published patents only
Key numbers
  • 55% (36 of 66) Antecedent-basis flags that are genuine, random granted claims (held out, n = 66)
  • 93.4% Supported-or-not call agrees with the labels (test split, n = 166)
  • 100% Supported elements where a cited paragraph is one the labels list (test split, n = 150)
  • 7.8 s Median end-to-end run, hosted (QA sweep 2026-09-25)
All results, datasets and caveats
Models
Qwen3.8-27B
Where
Self-host for unfiled inventions; hosted for published patents only
Checks
Receipt per element; signed record and practitioner ledger
Industry
Legal
Output
Structured data · Signed record or verdict
Data
Privileged or legal · Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host
Built from
Signed record

Questions people ask

What does the checker look at?

Code splits the claims into elements and checks that every 'the' or 'said' has an antecedent, that every dependent claim points at an earlier claim that exists, and that the claims use no word the specification never uses. An open model then reads each element against the description and names the paragraphs that describe it, with a strength: supported, partly supported or no support found.

Does 'supported' mean the claim is allowable?

No. It means the model found paragraphs that describe the element, not that there is no 35 U.S.C. 112(a) problem. It is a review aid for a registered practitioner, who decides. It misses partial support often (5 of 13 held-out elements), so treat 'supported' as 'here is where to look'.

How accurate is it?

On 166 held-out elements from 6 granted US patents, the supported-or-not call agreed with the labels 93.4% of the time, and every supported element cited at least one paragraph the labels list. About half of the antecedent-basis flags on randomly drawn granted claims were genuine slips (36 of 66). Labels are AI agents', not a practitioner's.

Can I run it on an unfiled application?

Yes, on your own machine: run unfiled applications self-hosted. The hosted demo takes published patents and applications only. The USPTO's April 2024 guidance reminds practitioners to prevent disclosure of client information through AI tools.

What does it cost?

About $0.053 of model time per application at the gateway's list price (measured on 6 granted patents with 17-50 elements each); GPU time only when self-hosted.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Patent claim-support checker

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.