Skip to content
decosa
LiveHostedSelf-hostMacSelf-host first for real data

Redline from firm precedents

A Word redline with real tracked changes and a margin comment on each, using wording taken only from the firm's own precedents.

Measured0.82 / 0.85Clauses found vs expert labels (CUAD): precision / recall
On production28 smedian on production (2026-09-25); slower when the service is busy
List price~$0.89 per 100 contractsmeasured, at list price

Built on: Grounding, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended
  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; the DOCX writer, search and checks run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa
  • Call the privileged drafting editor API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real client material belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
legal-drafting-editor

Use the hosted API

# Decosa Privileged drafting editor: use the hosted API

You are wiring Decosa's drafting editor into this project. It takes a contract (a .docx, plain text or a sample) and a
law firm's playbook, and returns a Word file with real tracked changes (w:ins / w:del) and a margin comment on each
change naming the playbook rule, the position and the precedent it came from. New wording comes only from the firm's
precedent clauses: code traces every inserted word to one, and checks that reject-all gives back the input and
accept-all gives exactly the proposal. Each model call has its own signed receipt. Use only what is listed below. If
you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- The hosted API is for public contracts (the CUAD samples), fictional ones, and testing your integration. Client
  contracts are confidential and often privileged: for real matters use the self-host prompt instead. Never send client
  documents here.
- It proposes; a lawyer decides. Show every change as a proposal to accept or reject, never as final.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "legal-drafting-editor"}` returns `{"token", "expires_at", "budget"}`.
   a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); each session has a token allowance (its `budget`) (one review of the full ten-rule playbook uses about
   2,000). Over a limit: HTTP 429 with `Retry-After`; not enough budget: HTTP 402.
3. One review at a time per demo token (409 otherwise).

## Endpoints
- `POST /drafting/review` (token). Body: exactly one of `{"docx_b64": "<base64 .docx>"}`, `{"text": "..."}` or
  `{"sample": "northwind-msa"}`, plus optional `"rules": ["GL", "CAP", ...]`, `"parties": {"client": "<defined term>", "counterparty": "<defined term>"}`,
  `"playbook": {...}`, `"precedents": {...}` (the shapes of `GET /drafting/playbook`), `"title"`, `"stream"`.
  - Limits: .docx up to 2 MB, text up to 200,000 characters, 3,000 paragraphs, 20 rules, 80 precedent clauses.
  - With `"stream": true` (or `Accept: text/event-stream`): `ready`, `parties`, a `receipt` after each model call, a
    `finding` per rule `{rule, position: preferred|fallback|outside|walk_away|absent|error, paragraphs, quote, reason, action}`,
    an `edit` per change `{id, rule, action: edit|replace_clause|delete|insert|comment|flag, status: tracked|comment_only|flag, ops, before, after, precedent, trace, checks, comment}`,
    `validation`, `report` (totals, usage, signed record), `done` (run id and export links), `budget`.
    With `"stream": false`: one JSON object `{run_id, export, findings, edits, validation, report, ...}`.
- `GET /drafting/runs/{run_id}/export?format=docx` (the redline), `md` (review memo), `record` (signed), `ledger` (signed, with decisions).
- `POST /drafting/runs/{run_id}/reconcile` `{"docx_b64", "reviewer"}`: the file back from Word; which changes were
  accepted, rejected or left open, sealed into the ledger. `POST /drafting/runs/{run_id}/decisions` records them by hand.
- `POST /drafting/validate` `{"docx_b64"}` (token, no model): structural check of any DOCX with tracked changes.
- `GET /drafting/info`, `/drafting/playbook`, `/drafting/samples`, `/drafting/samples/{id}`, `/drafting/samples/{id}/docx` (no token).
- `POST /record/verify` (no token) `{"record": {...}}` → `{ok, summary, bad}`.

## Example: redline a contract and save the .docx (Python, `pip install httpx`)
```python
import base64, httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
docx = base64.b64encode(pathlib.Path("contract.docx").read_bytes()).decode()   # a public or test contract
r = httpx.post(f"{API}/drafting/review", headers=H, timeout=900, json={"docx_b64": docx, "stream": False})
r.raise_for_status()
run = r.json()
for f in run["findings"]:
    print(f["rule"], f["position"], f["action"])
assert run["validation"]["ok"] and run["report"]["totals"]["untraced_insertions"] == 0
pathlib.Path("redline.docx").write_bytes(httpx.get(f"{API}{run['export']['docx']}", headers=H).content)
```

## Honest limits
- Positions come from an open model and can be wrong; see the measured rates on the Stack tab. The lawyer reviews every change.
- It edits paragraphs of plain text runs in place; a paragraph with fields, hyperlinks or earlier tracked changes gets a
  comment with the proposed wording instead. Numbering of inserted clauses is left to the lawyer.
- It only uses the precedents you give it. With no precedent for a rule, it comments or flags; it never writes new clause text.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Privileged drafting editor: run it yourself (containers)

You are setting up the Decosa drafting editor on this machine, so client contracts never leave it. It applies the
firm's playbook to a contract and returns a .docx with Word tracked changes and a margin comment per change, using only
the firm's precedent clauses for new wording, and signs a record. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/legal-drafting-editor.zip (11 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py legal-drafting-editor` (the api image carries the same bundle under /app/rehearsal/legal-drafting-editor/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py legal-drafting-editor --bundle legal-drafting-editor.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "Texas governing law is placed outside the playbook", "the governing-law clause is redlined from Texas to New York as a tracked change", "the 30-day warranty is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
   check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`, use a named volume
   for `/data`, and bind every port to 127.0.0.1. Nothing in this tool needs the internet after the weights are
   downloaded.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/drafting/info` shows `"route": "direct"`; `GET /attest/signing-key` shows
   this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"legal-drafting-editor"}` and run
   `POST /drafting/review {"sample": "northwind-msa", "stream": false}`. Expect governing law (Texas) outside and
   redlined to New York, the one-month liability cap and the 30-day warranty at walk-away, confidentiality and the IP
   indemnity inserted from precedents, `report.totals.untraced_insertions` 0, and `validation.ok`,
   `reject_all_is_input` and `accept_all_is_proposal` all true. Download `export.docx` and open it in Word. Then
   `POST /record/verify` with `report.record`: `ok` must be true.
6. Load the firm's playbook and precedent library (`GET /drafting/playbook` shows the shape) and send them with each
   review as `playbook` and `precedents`.
7. Report back: the public key and key id, the findings, the file checks and how long the run took.

Off, and keep it off on a box that holds client contracts: joining serves other people's requests on this GPU.
If I ask for it later, on a separate machine, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/legal-drafting-editor-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (32 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (32 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"legal-drafting-editor"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py legal-drafting-editor

Download the mock-data bundle (11 KB, 12 checks)expected.json

A two-page synthetic services agreement (Northwind / Brightwell) as a .docx, reviewed against three rules of a fictional firm's playbook and precedent library, both sent as files. Texas law must be found outside the playbook and redlined to New York, the 30-day warranty flagged, and the mutual exclusion of indirect loss found; the redline must pass its checks (reject-all gives the input, accept-all the proposal) with no untraced words, and the signed record must verify.

What the rehearsal checks
  • Texas governing law is placed outside the playbook
  • the governing-law clause is redlined from Texas to New York as a tracked change
  • the 30-day warranty is flagged
  • the exclusion of indirect loss (clause 4.2) is found and quoted
  • no inserted word is untraced
  • the redline passes its package checks
  • rejecting every change gives back the input
  • accepting every change gives the proposal
  • the review memo (Markdown) has the governing-law finding
  • the signed record verifies
  • a record with its title edited no longer verifies
  • every model call has a signed receipt

Licence: Synthetic: Northwind, Brightwell and Halvorsen & Pike LLP are fictional; the contract, playbook and precedent clauses were written for the Decosa demo. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa Privileged drafting editor: run it yourself (containers)

You are setting up the Decosa drafting editor on this machine, so client contracts never leave it. It applies the
firm's playbook to a contract and returns a .docx with Word tracked changes and a margin comment per change, using only
the firm's precedent clauses for new wording, and signs a record. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/legal-drafting-editor.zip (11 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py legal-drafting-editor` (the api image carries the same bundle under /app/rehearsal/legal-drafting-editor/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py legal-drafting-editor --bundle legal-drafting-editor.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "Texas governing law is placed outside the playbook", "the governing-law clause is redlined from Texas to New York as a tracked change", "the 30-day warranty is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
   check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`, use a named volume
   for `/data`, and bind every port to 127.0.0.1. Nothing in this tool needs the internet after the weights are
   downloaded.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/drafting/info` shows `"route": "direct"`; `GET /attest/signing-key` shows
   this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"legal-drafting-editor"}` and run
   `POST /drafting/review {"sample": "northwind-msa", "stream": false}`. Expect governing law (Texas) outside and
   redlined to New York, the one-month liability cap and the 30-day warranty at walk-away, confidentiality and the IP
   indemnity inserted from precedents, `report.totals.untraced_insertions` 0, and `validation.ok`,
   `reject_all_is_input` and `accept_all_is_proposal` all true. Download `export.docx` and open it in Word. Then
   `POST /record/verify` with `report.record`: `ok` must be true.
6. Load the firm's playbook and precedent library (`GET /drafting/playbook` shows the shape) and send them with each
   review as `playbook` and `precedents`.
7. Report back: the public key and key id, the findings, the file checks and how long the run took.

Off, and keep it off on a box that holds client contracts: joining serves other people's requests on this GPU.
If I ask for it later, on a separate machine, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/legal-drafting-editor-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsPrivileged drafting editor on GeForce RTX 5090: use the Standard · one 96 GB card (measured; hosted demo) tier

The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB (fits): Qwen3.8-27B NVFP4 needs about 20 GB of weights; the longest prompt (six candidate paragraphs and a rule) is under 4k tokens. Estimate: same model and prompts as the measured card, not run here on a 5090.

Standard · one 96 GB card (measured; hosted demo): what changesuses estimates

  • Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Editor: decosa-api drafting module (decosa_api/verticals/drafting). CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Model: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Privileged drafting editor, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Privileged drafting editor on my hardware

Fetch https://decosa.ai/prompts/legal-drafting-editor-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=legal-drafting-editor)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · one 96 GB card (measured; hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Editor: decosa-api drafting module (decosa_api/verticals/drafting), CPU
- Model: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/legal-drafting-editor-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh

Mac prompt for your coding agent

# Decosa Privileged drafting editor: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Privileged drafting editor on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/legal-drafting-editor.zip (11 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py legal-drafting-editor` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "Texas governing law is placed outside the playbook", "the governing-law clause is redlined from Texas to New York as a tracked change", "the 30-day warranty is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Editor: DOCX reading and the tracked-change writer, candidate search, traceability check, file checks, signed record and ledger (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured |
| Model: places each clause against the playbook, proposes the find-and-replace edits, judges the grounding of each changed sentence, re-reads the edited clause | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py legal-drafting-editor`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 28 s · ~$0.007 per run · 20 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the assemble prompt's api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.

Measured cost to run: about $0.89 per 100 contracts (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Verified on 2026-09-25: the image builds, the service starts healthy, the Northwind sample passes end to end (9.7 s, 20 attested calls, 7 tracked changes, 0 untraced insertions, the file's checks pass), the signed record verifies and fails when one finding is changed, and the .docx export downloads. The model server's own startup was not re-verified (no new GPU load).

Known limits (5)
  • The playbook and precedent library in the demo are synthetic; a firm's own need to be loaded (as JSON) and were not tested.
  • Paragraphs with fields, hyperlinks, drawings or earlier tracked changes get a comment, not an in-place edit; inserted clauses are not numbered.
  • Checked against the OOXML rules and LibreOffice; not yet opened in Microsoft Word by us.
  • Position errors: 2% of planted clauses got the wrong one of five positions (v2 held-out run) and no planted deviation was missed; like-for-like clause detection against CUAD is 0.77 / 0.79, and warranty-duration and insurance clauses are the ones most often missed.
  • The client-side playbook needs a buyer side: contracts between equals (joint ventures, cooperation agreements) get findings but party-specific precedents stay in comments.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended
  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; the DOCX writer, search and checks run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa
  • Call the privileged drafting editor API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real client material belongs on your own hardware.
The open stack

A firm's playbook applied to a contract: a Word redline with real tracked changes, a margin comment on each, and new wording only from the firm's own precedents, traced in code.

Send a contract (.docx or text) and the firm's playbook: preferred, fallback and walk-away positions per clause, and a library of the firm's precedent clauses. For each rule a search picks the candidate paragraphs and an open model places the clause (preferred, fallback, outside, walk-away or missing) and quotes the words that decide it. Where the playbook says to act, the model proposes the smallest find-and-replace using only the precedent's words; code applies it, diffs it word by word and traces every inserted chunk to the precedent, then the grounding checker and a second read of the rule confirm the edit, and anything that fails falls back to the precedent's own sentence or a comment. Missing clauses are inserted word for word from the library. The result is a .docx with w:ins / w:del revisions and a comment per change naming the rule, the precedent and the receipts; reject-all must give back the input and accept-all exactly the proposal. Upload the file back from Word and the lawyer's accept or reject decisions go into a signed ledger. It proposes; a lawyer decides.

Deployment
Self-host first
Regulatory
A drafting aid for lawyers, not legal advice; the lawyer reviews and adopts every change. Under ABA Model Rule 1.1 (comment 8) a lawyer using it must understand what it does and check its output, and Rules 5.1 and 5.3 make supervision of the tool the lawyer's duty: the signed record and the ledger of accepted and rejected changes are evidence of that review. Rule 1.6(c) requires reasonable efforts to prevent unauthorised disclosure of client information, and ABA Formal Opinion 512 (29 Jul 2024) asks lawyers to understand how a generative-AI tool uses what they put in and to get informed consent before putting client information into a self-learning tool (boilerplate in an engagement letter is not enough); United States v. Heppner (S.D.N.Y., Feb 2026) held chats with a consumer AI service not privileged, partly because its terms allowed disclosure. Self-hosted, the contract never leaves the firm's machine and nothing is trained on it; the hosted demo is for public (CUAD) and fictional contracts only. Bar guidance varies by state (California's 2026 practical guidance, for example, requires AI costs to be billed at cost without markup unless agreed in writing); check the rules where you practise. The playbook and precedents in the demo are synthetic. CUAD v1 (The Atticus Project) is CC BY 4.0; its contracts are public SEC EDGAR filings. Model licence: Apache-2.0 (Qwen3.8-27B). Checked 25 Sep 2026.
Architecture
Text description

A contract (DOCX or text) and the firm's playbook and precedent library enter the drafting module on the firm's machine. A search picks candidate paragraphs per rule; Qwen3.8-27B (Apache-2.0) places each clause and proposes edits using only precedent words. Code traces every inserted word to a precedent; the grounding checker and a second read of the rule confirm each edit. The writer produces a DOCX with Word tracked changes and comments, checked so that reject-all gives the input and accept-all the proposal. Each model call's receipt is signed by our gateway (hosted) or the box's own key (self-host) and listed in a signed, hash-chained record; the file returned from Word feeds a signed decision ledger.

Architecture

At a glance

Data retention
The contract is held in memory for the request; the run (findings, changes, the redline) for one hour, for the token or key that made it. Nothing is written to disk; logs carry counts only.
What leaves the box
Self-hosted: nothing. Hosted demo: the contract's candidate paragraphs go to the model through our gateway, which is why the hosted demo is for public and fictional contracts only.
Where new wording comes from
Only the firm's precedent clauses, with the parties' defined terms filled in. Every inserted chunk is traced to its precedent in code; if a check fails the change falls back to the precedent's own sentence or to a comment.
Input
A .docx up to 2 MB or text up to 200,000 characters; a playbook of up to 20 rules and a library of up to 80 precedent clauses as JSON (the default is a synthetic one).
Output
A .docx with Word tracked changes and a comment per change; a review memo; a signed record; and, from the file returned from Word, a signed ledger of which changes the lawyer accepted or rejected.
Model cost per contract
About a cent or less per contract at the gateway's list price (median, measured on CUAD runs); GPU time only when self-hosted.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 32 GB card, self-hosted

    The same model and prompts on a single RTX 5090. The grounding check of each edit can be switched off (DECOSA_DRAFTING_GROUNDING=0) to save a call or two per change; the code trace still guarantees no invented words.

    Models
    • decosa-api drafting module (decosa_api/verticals/drafting)
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX 5090 32 GB
    Quality evidence
    • Findings and redlines on CUADnot measured yetnot measured yet
    Latency
    estimate: not timed on a 5090.
    Verification
    Proof: partialSelf-host onlySelf-hosted: calls and records are signed by your own box, not countersigned by the gateway.
  • In the hosted demo

    Standard

    one 96 GB card (measured; hosted demo)

    Qwen3.8-27B on an RTX PRO 6000: one call per playbook rule (with the closest CUAD train-split examples of the rule's clause), one per targeted edit, plus the grounding check and the re-read of each edit.

    Models
    • decosa-api drafting module (decosa_api/verticals/drafting)
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX PRO 6000 96 GB
    Quality evidence
    • Clause detection vs CUAD expert labels, 8 units × 30 held-out test contracts: precision / recall0.82 / 0.85 under the v2 mapping (v1: 0.81 / 0.79, re-scored); like for like with v1's mapping 0.77 / 0.79 (v1: 0.78 / 0.82), so no gain theredocs/evals/legal-drafting-editor.md (v2). The v2 gain is the new CONSEQ rule, which finds the exclusions of indirect damages CUAD files under "Cap On Liability" (9 of 9); warranty duration 0 of 3 and insurance 2 of 5 labelled contracts found
    • Planted deviations flagged (102 planted clauses, 17 CUAD test contracts): precision / recall0.98 / 1.00 (52 of 52 deviations flagged, 1 false flag); exact position (of 5) 0.98docs/evals/legal-drafting-editor.md; v1 was 0.89 / 0.95, exact 0.85, on 13 contracts, and 0.95 / 0.97, exact 0.92, once 6 items a harness bug had left out of the file are removed
    • Contracts where the buyer side was named (needed for party-specific precedents)17 of 30 (v1: 13); both names among CUAD's labelled parties in 17 of 17docs/evals/legal-drafting-editor.md
    • Redlines that are valid, reject-all = input, accept-all = proposal47 of 47docs/evals/legal-drafting-editor.md, every output file of the v2 held-out run
    • Redlines LibreOffice opens and whose Accept All / Reject All match47 of 47 open, Reject All 47 of 47, Accept All 46 of 47; 284 of 284 comments importeddocs/evals/legal-drafting-editor.md: the mismatch was a deleted last paragraph (LibreOffice and Word keep an empty one); fixed after the run and tested, not re-run on the test set
    • Inserted text not traceable to a firm precedent0 in 136 tracked changesdocs/evals/legal-drafting-editor.md; enforced in code: an untraced chunk sends the edit back to the precedent's sentence
    • Targeted edits the checks sent back to the precedent's own sentence26 of 77docs/evals/legal-drafting-editor.md (v2 held-out run)
    • Model cost per contract, eleven rules (median)US$0.0089 at list price ($0.30 / $1.50 per million tokens), 16 calls (v1: US$0.0065, 14 calls)docs/evals/legal-drafting-editor.md, 47 runs
    • Levers tried on 40 dev contracts and left offSelf-consistency (3 readings): 2.6x the cost, no detection gain. Topic gate: +0.05 precision, -0.06 recall. Embedding retrieval (bge-small, CPU): +0 to +1 point of labelled text shown. Log-probabilities: the gateway does not return them.docs/evals/legal-drafting-editor.md, lever study
    Latency
    measured: seconds per contract in the eval, up to about a minute per demo run on the shared gateway, seconds self-hosted.
    Verification
    Proof: strongEvery model call has a gateway-signed receipt; the record and ledger are signed by the instance key.
  • Needs more compute

    Wanted: the best setup

    a larger second reader on your own hardware

    GLM-5.3-Flash reads each clause beside Qwen3.8-27B, and a rule fires only when both agree or the reviewer decides. Privileged drafts stay on your hardware, never on community providers. Not served yet.

    Models
    • decosa-api drafting module (decosa_api/verticals/drafting)
    • Qwen3.8-27B (NVFP4)
    • GLM-5.3-Flash
    Hardware
    Your own hardware: 2x 96 GB cards (NVFP4, about 170-186 GB, unconfirmed) or a Mac with 192 GB or more (MLX 4-bit, 165 GB). Estimate; GLM's SGLang SM120 build hung on our server.
    Quality evidence
    • CUAD precision and recall on the same 8 rulesnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.

Also runs on

  • Fast reader for long contractsGemma-4-26B-A4B-itnot servedGemma-4-26B-A4B for the clause-by-clause reading at volume, keeping Qwen3.8-27B for edits (page 28's plan; not run). Page 32 suggests the dense Qwen3.5-9B (Apache-2.0) instead. Hardware: 1x RTX PRO 6000 96 GB (BF16 weights are 49 GB).

We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.

Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Editor: DOCX reading and the tracked-change writer, candidate search, traceability check, file checks, signed record and ledger (no model; CPU)decosa-api drafting module (decosa_api/verticals/drafting)
0 GBProof: partial
Model: places each clause against the playbook, proposes the find-and-replace edits, judges the grounding of each changed sentence, re-reads the edited clauseQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Faster model for the clause-by-clause reading at volume (alternate)Gemma-4-26B-A4B-itgoogle/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
26B (4B active) · 49 GBNo proof yet
Second reader for clause detectionGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

  • CUAD v1 (Contract Understanding Atticus Dataset) (opens in a new tab)CC BY 4.0 (The Atticus Project); contracts are public SEC EDGAR filings

    The eval's contracts and expert clause labels (30 held-out test-split contracts; 40 train-split dev contracts for tuning), the two public samples in the demo (train split), and the few-shot examples in the finding prompts (clause spans from the rest of the train split only; about 200 KB, data/exemplars.json).

  • The eval's independent reader: opens every redline and applies its own Accept All and Reject All (scripts/drafting/lo_roundtrip.py). Optional on a self-hosted box (DECOSA_DRAFTING_SOFFICE) to check every run.

  • Synthetic playbook and precedent librarySynthetic, written for Decosa (CC0); a fictional firm

    Eleven rules (governing law, liability cap amount, exclusion of indirect damages, IP indemnity, confidentiality, work-product ownership, assignment, termination for convenience, non-compete, warranty, insurance) and fourteen precedent clauses. Replace them with the firm's own.

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /drafting/info, /drafting/playbook, /drafting/samples; POST /drafting/review (SSE or JSON); GET /drafting/runs/{id}/export?format=docx|md|record|ledger; POST /drafting/runs/{id}/reconcile and /decisions; POST /drafting/validate. Contracts are held in memory for the request and runs for one hour, never on disk; logs carry counts only.

  • vLLM:8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).

Hardware

  • 1x RTX 5090 32 GB Fits

    Qwen3.8-27B NVFP4 needs about 20 GB of weights; the longest prompt (six candidate paragraphs and a rule) is under 4k tokens. Estimate: same model and prompts as the measured card, not run here on a 5090.

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the eval, the hosted demo and the self-host check ran on this card, shared with other services the whole time.

Latency per lane

  • one contract, ten rules, hosted gateway route28.0 s

    Measuredmeasured on our server 2026-09-25: 10 to 58 s over 8 runs of the three samples (15 to 22 calls each) on the shared gateway; v1 code. v2 (eleven rules): 40 to 60 s for the two samples on 2026-09-25 with 2 requests in flight on a saturated GPU

  • one contract, ten rules, eval (v1: CUAD test contracts through the gateway, 3 contracts at a time)12.4 s

    Measuredmeasured on our server 2026-09-25: median over 43 runs, 14 calls each

  • synthetic sample, self-hosted (clean clone, container build, direct route)9.7 s

    Measuredmeasured on our server 2026-09-25: 20 attested calls, one run

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

legal-drafting-editor/assemble-prompt.md124 lines
# Assemble the Decosa drafting editor on this machine

You are setting up a contract redlining assistant for a law firm. It takes a contract (.docx or text) and the firm's
playbook (preferred, fallback and walk-away positions per clause) with its library of precedent clauses, and returns a
.docx with real Word tracked changes and a margin comment on each change naming the rule, the position and the
precedent. New wording comes only from the precedents: code traces every inserted word to one. Everything is sealed in
a record signed by this box's own key. Work step by step, show me each command before you run anything with `sudo`,
and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/legal-drafting-editor.zip (11 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py legal-drafting-editor` (the api image carries the same bundle under /app/rehearsal/legal-drafting-editor/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py legal-drafting-editor --bundle legal-drafting-editor.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "Texas governing law is placed outside the playbook", "the governing-law clause is redlined from Texas to New York as a tracked change", "the 30-day warranty is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Model: Qwen3.8-27B (Apache-2.0). Everything else is CPU code in decosa-api (AGPL-3.0-or-later): the DOCX writer (Python
  standard library), the search, the traceability check and the grounding checker. LibreOffice (MPL-2.0) is optional.
- Client contracts are confidential and often privileged. Bind every port to 127.0.0.1 and do not send
  contracts to any hosted API. The service keeps contracts in memory for the request
  and runs for one hour, writes nothing to disk and logs counts only; keep it that way.
- It proposes; a lawyer decides (ABA Model Rules 1.1, 5.1 and 5.3; Formal Opinion 512). Say so wherever you show results.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 is about 20 GB of weights; the longest prompt is under
   4k tokens). Driver 570 or newer; Blackwell cards run NVFP4, older cards use the FP8 weights.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
   from the official Docker and NVIDIA repositories after asking me, then run
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 30 GB free for the model and images.

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a release that contains
  `decosa_api/verticals/drafting/`, and run `docker build -f docker/api/Dockerfile -t decosa-api:local .`
- `vllm/vllm-openai:v0.29.0` for the model; weights `nvidia/Qwen3.8-27B-NVFP4` (or `Qwen/Qwen3.8-27B-FP8`).

## 3. docker-compose.yml
Write this in `~/decosa/drafting/`:

```yaml
name: decosa-drafting
services:
  llm:
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
              "--enable-prefix-caching"]
    ports: ["127.0.0.1:8114:8000"]
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>     # or decosa-api:local
    ports: ["127.0.0.1:8445:8445"]
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_BUDGET_LLM_TOKENS: "400000"        # per session; a ten-rule review uses about 2,000 generated tokens
      DECOSA_DRAFTING_MAX_CONCURRENT: "2"
    volumes: ["decosa-data:/data"]                # a named volume: the image runs as uid 10001, so a root-owned bind mount fails
    depends_on: { llm: { condition: service_healthy } }
    healthcheck:
      test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/drafting/info', timeout=4)"]
      interval: 30s
      retries: 10
volumes:
  decosa-data: {}
```

On first start the api service creates this box's Ed25519 key in the `decosa-data` volume under `attest/` (mode 0600).
Back the volume up and never print the key. Records and model calls are signed with it: an attestation by me, the
operator, not a proof. Optional: install LibreOffice in the api image and set `DECOSA_DRAFTING_SOFFICE=/usr/bin/soffice`
to have every redline opened by LibreOffice as part of its checks.

## 4. Smoke test
1. `curl -s localhost:8445/drafting/info | jq '{model: .model.route, rules: [.playbook.rules[].id]}'` shows `direct` and eleven rules.
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"legal-drafting-editor"}' | jq -r .token)`.
3. Run the fictional sample:
   `curl -s -XPOST localhost:8445/drafting/review -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample":"northwind-msa","stream":false}' > run.json`.
   Expect: `jq '[.findings[] | {rule, position, action}]' run.json` shows governing law (Texas) outside, the liability
   cap, assignment, non-compete and warranty at walk-away, confidentiality and IP indemnity absent (inserted), insurance
   at fallback (comment); `.report.totals.untraced_insertions` is 0; `.validation` has `ok`, `reject_all_is_input` and
   `accept_all_is_proposal` all true.
4. Save the redline: `curl -s localhost:8445$(jq -r .export.docx run.json) -H "authorization: Bearer $T" -o redline.docx`
   and open it in Word or LibreOffice: Review shows each change with Accept / Reject and a comment beside it.
5. `jq '{record: .report.record}' run.json | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @-`
   must say `ok: true`. Change one `position` in the record's entries and verify again: it must fail and name that entry.
6. Time it and tell me. On our RTX PRO 6000, shared with other work, the sample took 10 s self-hosted (20 calls).

## 5. Load the firm's own playbook and precedents
`GET /drafting/playbook` returns the default (synthetic) playbook and library in the exact shape the review accepts.
Replace them with the firm's: each rule `{id, title, topic, preferred, fallback, walk_away, on_outside, on_walk_away,
if_missing, target, precedents, terms, patterns}` and each precedent `{id, topic, position, title, source, text}`, with
only `{Client}` and `{Counterparty}` as placeholders. Precedents must be clauses the firm has approved; the editor never
writes wording of its own. Send them with each `POST /drafting/review` as `playbook` and `precedents`, and pass
`parties: {client, counterparty}` with the defined terms when you know them. After the lawyer works in Word, send the
saved file to `POST /drafting/runs/{id}/reconcile {docx_b64, reviewer}` and keep `export?format=ledger` with the matter.
For the site, set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in `.env.local`. Contract: `API_CONTRACT.md`,
section "Privileged drafting editor".

Off, and it should stay off on a box that holds client contracts: joining serves other people's requests on this GPU.
Only on a separate machine, and only with my explicit yes, follow the provider guide at `/provide` on the site.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A firm's playbook applied to a contract: a Word redline with real tracked changes, a margin comment on each, and new wording only from the firm's own precedents, traced in code.
Who it's for
Teams in legal.
Where it runs
Self-host for client contracts; hosted for public and fictional samples
Key numbers
  • 0.77 / 0.79 Clause detection vs CUAD labels, v1 mapping (like for like): precision / recall (held out, n = 30)
  • 0.82 / 0.85 Clause detection vs CUAD labels, v2 mapping: precision / recall (held out, n = 30)
  • 0.98 / 1.00 Planted deviations flagged: precision / recall (held out, n = 102)
  • 28.0 s Median end-to-end run, hosted (QA sweep 2026-09-25)
All results, datasets and caveats
Models
Qwen3.8-27B
Where
Self-host for client contracts; hosted for public and fictional samples
Checks
Receipt per call; every insertion traced to a precedent; validated redline; signed record and decision ledger
Industry
Legal
Output
Notes, reports and drafts · Signed record or verdict
Data
Privileged or legal
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

Where does the new contract wording come from?

Only from the firm's own precedent clauses, with the parties' defined terms filled in. Code traces every inserted chunk back to a precedent; a grounding check and a second read of the rule confirm each edit, and anything that fails falls back to the precedent's own sentence or a comment. On 30 held-out CUAD contracts there were 0 untraced insertions in 136 tracked changes.

Is the output a real Word redline?

Yes: a .docx with w:ins / w:del revisions and a margin comment per change naming the rule, the position and the precedent. Reject-all must give back the input and accept-all exactly the proposal; LibreOffice Accept All matched 46 of 47 held-out files. We have checked it against the OOXML rules and LibreOffice, not yet in Microsoft Word itself.

How accurate is the clause review?

On held-out CUAD contracts, like-for-like clause detection was precision 0.77 / recall 0.79 against expert labels, and planted deviations were flagged with precision 0.98 and recall 1.00 (52 of 52). Warranty-duration and insurance clauses are missed most often, and the playbook and precedents were synthetic, so a lawyer reviews every finding.

Does the client's contract leave the firm?

Not when self-hosted: the model (Qwen3.8-27B, Apache-2.0) and the code run on your machine and nothing is trained on the contract. The hosted demo sends candidate paragraphs through our gateway, so it is for public (CUAD) and fictional contracts only.

How is the lawyer's review recorded?

Upload the file back from Word and the lawyer's accept or reject decisions go into a signed ledger, alongside a signed record of what the tool proposed. It proposes; a lawyer decides.

What does it cost?

A median US$0.0089 of model time per contract at the gateway's list price (measured on CUAD runs); GPU time only when self-hosted.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Privileged drafting editor

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.