Build a privilege log
A privilege call for every document with the reason checked against it, plus draft log entries that describe the document without giving the advice away.
Built on: Typed judgment, Signed record
1. Pick a sample
Your own documents: run the tool on your own hardware, or request confidential access. The demo takes samples or made-up data only.
2. Run it
On production the sample took 2.0 min (median, 2026-09-25). Slower when the service is busy.
Result
The answer appears here first, then what it found, the draft, and how long it took. Sample: Harborline Freight: 12 made-up emails.
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; the people map, grouping and leak rules run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the privilege review and privilege log API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real client material belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- privilege-log
Use the hosted API
# Decosa Privilege review and privilege log: use the hosted API
You are wiring Decosa's privilege reviewer into this project. It takes emails and memos from a production set and the
client's lawyers, and returns a typed call per document (attorney-client, work product, both, not privileged, or needs
attorney review) with the lawyers involved, a reason checked against the document and a calibrated probability. It
drafts privilege-log entries that pass a leak check, gives exact copies the same call, flags outsiders who could waive
privilege, and ends with a signed record. Each model call has its own signed receipt. Use only what is listed below. If
you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- The hosted API is for the public and synthetic sample sets and for testing your integration. Privileged documents
are the most sensitive a client has, and sending them to a third party's service can put privilege at risk: for real
matters use the self-host prompt instead. Never send client documents here.
- It drafts; a lawyer decides. Show every call as a draft for attorney review, never as a decision.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
`DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "privilege-log"}` returns `{"token", "expires_at", "budget"}`.
a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); each session has a token allowance (its `budget`) (about 20 documents). Over a limit: HTTP 429 with `Retry-After`.
3. One review at a time per demo token (409 otherwise).
## Endpoints
- `POST /privilege/review` (token). Body:
`{"documents": [{"id", "date"?, "from"?, "to"?: [...], "cc"?: [...], "bcc"?, "subject"?, "body", "thread_id"?, "family_id"?, "parent_id"?, "bates"?}], "people"?: {"attorneys": [{"name"?, "email"?, "role": "in-house"|"outside", "firm"?}], "client_domains": [...], "counsel_domains": [...], "agents"?: [{"email"}]}, "samples"?: 4, "title"?: "...", "stream"?: true}`
or `{"sample": "harborline-dispute"}` to run a sample set.
- Limits: 25 documents per request, 20,000 characters each, 150,000 in all. `samples`: 4 (default), 2, or 0 (one call
per question, stated confidence only).
- With `"stream": true` (or `Accept: text/event-stream`) it streams `ready` (people map, flags, groups), a `receipt`
after each model call, a `document` event per document (replace by `id`), then `consistency`, `log`, `report`,
`done` and `budget`. With `"stream": false`: one JSON object `{run_id, totals, documents, people, groups, consistency, log, report, budget}`.
- Each document: `{id, final, model_call, probability, reasons, legal_advice, litigation, flags, basis: {lawyers, reason, grounding}, log?: {description, fallback, attempts}, copy_of?}`.
`final` is `needs_review` whenever `reasons` is not empty; treat it as withheld until a lawyer decides.
- `POST /privilege/runs/{run_id}/decisions` `{"decisions": [{"doc", "decision": "confirm"|"change", "call"?, "reviewer", "note"?}]}`: the lawyer's calls.
- `GET /privilege/runs/{run_id}/export?format=csv` (the log), `md` (review memo), `record` (signed), `ledger` (signed, with the decisions).
- `POST /record/verify` (no token) `{"record": {...}}` → `{ok, summary, bad}`.
- `GET /privilege/info`, `GET /privilege/samples`, `GET /privilege/samples/{id}`, `GET /attest/signing-key` (no token).
## Example: review a batch and write the draft log (Python, `pip install httpx`)
```python
import httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
s = httpx.get(f"{API}/privilege/samples/harborline-dispute").json() # your own batch has the same shape
r = httpx.post(f"{API}/privilege/review", headers=H, timeout=900,
json={"documents": s["documents"], "people": s["people"], "stream": False})
r.raise_for_status()
run = r.json()
for d in run["documents"]:
print(d["id"], d["final"], d.get("probability"), "; ".join(d["reasons"]))
pathlib.Path("privilege-log.csv").write_text(httpx.get(f"{API}{run['export']['csv']}", headers=H).text)
```
## Honest limits
- The calls are drafts from an open model and code rules. On labelled public email they miss some privileged documents
and send many to review; see the measured rates on the Stack tab before relying on them.
- It reads what you send: attachments must be sent as their own documents, and the people map is only as good as the
lawyer list and domains you give it.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa Privilege review and privilege log: run it yourself (containers)
You are setting up the Decosa privilege reviewer on this machine, so privileged documents never leave it. It makes a
typed privilege call per email or memo, drafts privilege-log entries that pass a leak check, reconciles duplicates and
threads, and signs a record. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/privilege-log.zip (3 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py privilege-log` (the api image carries the same bundle under /app/rehearsal/privilege-log/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py privilege-log --bundle privilege-log.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the general counsel's legal advice (HL-002) is withheld as privileged or sent to review", "the ops update with a lawyer only copied (HL-021) is produced", "HL-021 is flagged: a lawyer only copied"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_PRIVILEGE_MAX_DOCS=200` and bind every port to 127.0.0.1. Nothing in this tool needs the internet after
the weights are downloaded.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/privilege/info` shows `"method": "logprobs (one call per question)"`;
`GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"privilege-log"}` and run
`POST /privilege/review {"sample": "harborline-dispute", "stream": false}`. Expect HL-004 to be a copy of HL-002 with
the same call, HL-003 (advice forwarded to an outside consultant) in `needs_review` with a third-party reason, and
HL-012 and HL-021 not privileged. Then `POST /record/verify` with `report.record`: `ok` must be true.
6. Report back: the public key and key id, the totals, the log entries and how long the run took.
Off, and keep it off on a box that holds privileged documents: joining serves other people's requests on this GPU.
If I ask for it later, on a separate machine, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/privilege-log-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVFP4) needs a GPU.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits (32 of 96 GB).
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits (32 of 64 GB).
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"privilege-log"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py privilege-log
Download the mock-data bundle (3 KB, 11 checks)expected.json
Four synthetic emails from the fictional Harborline Freight: the general counsel's legal advice on a lease, the same email kept by a second custodian, that advice forwarded to an outside consultant, and a weekly ops update with the general counsel only copied. The advice must be withheld (or sent to review), the ops update produced, the duplicate treated as one document, the forward flagged for its outside party, and the log must end in a signed record that verifies.
What the rehearsal checks
- the general counsel's legal advice (HL-002) is withheld as privileged or sent to review
- the ops update with a lawyer only copied (HL-021) is produced
- HL-021 is flagged: a lawyer only copied
- the forward to an outside consultant (HL-003) is flagged for its outside party
- the second custodian's copy (HL-004) is judged once, as a copy of HL-002
- every drafted log description passed the leak check
- the log (CSV) lists the withheld advice and leaves out the produced ops update
- the produced ops update is not on the privilege log
- the signed record verifies
- a record with its call counts changed no longer verifies
- every model call has a signed receipt
Licence: Synthetic, written for Decosa (CC0): Harborline Freight, Calder & Voss LLP and every person are fictional.
Prompt for your coding agent
# Decosa Privilege review and privilege log: run it yourself (containers)
You are setting up the Decosa privilege reviewer on this machine, so privileged documents never leave it. It makes a
typed privilege call per email or memo, drafts privilege-log entries that pass a leak check, reconciles duplicates and
threads, and signs a record. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/privilege-log.zip (3 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py privilege-log` (the api image carries the same bundle under /app/rehearsal/privilege-log/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py privilege-log --bundle privilege-log.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the general counsel's legal advice (HL-002) is withheld as privileged or sent to review", "the ops update with a lawyer only copied (HL-021) is produced", "HL-021 is flagged: a lawyer only copied"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_PRIVILEGE_MAX_DOCS=200` and bind every port to 127.0.0.1. Nothing in this tool needs the internet after
the weights are downloaded.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/privilege/info` shows `"method": "logprobs (one call per question)"`;
`GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"privilege-log"}` and run
`POST /privilege/review {"sample": "harborline-dispute", "stream": false}`. Expect HL-004 to be a copy of HL-002 with
the same call, HL-003 (advice forwarded to an outside consultant) in `needs_review` with a third-party reason, and
HL-012 and HL-021 not privileged. Then `POST /record/verify` with `report.record`: `ok` must be true.
6. Report back: the public key and key id, the totals, the log entries and how long the run took.
Off, and keep it off on a box that holds privileged documents: joining serves other people's requests on this GPU.
If I ask for it later, on a separate machine, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/privilege-log-mac.md instead.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsPrivilege review and privilege log on GeForce RTX 5090: use the Standard · one 96 GB card (measured; hosted demo) tier
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
What this tool's stack says about this hardware:
- 1x RTX 5090 32 GB (fits): Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for a 12,000-character email (about 4k tokens). Estimate: same model and prompts as the measured card, not run here on a 5090.
Standard · one 96 GB card (measured; hosted demo): what changesuses estimates
- Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Reviewer: decosa-api privilege module (decosa_api/verticals/privilege). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Model: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Privilege review and privilege log, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Privilege review and privilege log on my hardware Fetch https://decosa.ai/prompts/privilege-log-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=privilege-log) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · one 96 GB card (measured; hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Reviewer: decosa-api privilege module (decosa_api/verticals/privilege), CPU - Model: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/privilege-log-assemble.md
Or on a Mac Studio
No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where
From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh
Mac prompt for your coding agent
# Decosa Privilege review and privilege log: run it on this Mac (Apple Silicon, no NVIDIA GPU) You are setting up the Decosa Privilege review and privilege log on this Mac, natively on Apple Silicon. The models run on the Mac's GPU through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API. Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more. Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop or kill a process this setup did not start; if a port is taken, pick another one. ## Step 0: set up with a coding agent, rehearse on mock data, then go private This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works. Work in this order: 1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to "test with something realistic". 2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool, https://decosa.ai/samples/privilege-log.zip (3 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json), show me what is in it, and run the rehearsal against the local API: `.venv/bin/python scripts/rehearse.py privilege-log` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key). It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the general counsel's legal advice (HL-002) is withheld as privileged or sent to review", "the ops update with a lawyer only copied (HL-021) is produced", "HL-021 is flagged: a lawyer only copied"). Show me the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json` to make a check pass. 3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this machine. For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent can read. Switch to your own data only after the rehearsal has passed and the agent's work is done. ## What runs where | Part | On an NVIDIA GPU | On this Mac | Status | |---|---|---|---| | Reviewer: people map, waiver flags, duplicate and thread grouping, review rules, leak rules, consistency, signed record and ledger (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured | | Model: the typed privilege call, the two element questions, the grounding check of the reason, the log description and the leak judge | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured | ## Steps 1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory. 2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`. 3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`. Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me. 4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api) and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them. If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`. 5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key: show it to me, because it is what others pin to check the receipts and records this Mac signs. 6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py privilege-log`. It runs the tool's own sample end to end against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts. `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found. 7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`, the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`. 8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of `scripts/mac/setup.sh status`. ## Good to know - Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a self-hosted Mac. - The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published evals use. Expect small differences in wording and scores. - Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel calls; typed judgments then use sampling because oMLX returns no log-probabilities). - mlx_lm.server returns 10 log-probabilities per token instead of 20; the setup script sets DECOSA_JUDGMENT_TOP_LOGPROBS=10. - Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details: `docs/self-host-mac.md` in the checkout.
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 121 s · ~$0.031 per run · 100 receipts
Loading the nightly status…
Self-host: verified 25 Sep 2026 · Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the assemble prompt's api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.
Measured cost to run: about $1.49 per 1,000 documents (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
Verified on 2026-09-25: the image builds, the service starts healthy, info reports logprobs, the harborline-dispute sample passes end to end (copy of HL-002 gets its call, HL-003 goes to review, HL-012 and HL-021 produced, 17 s, 58 attested calls), the signed record verifies and fails when one call is changed, and the CSV export works. The model server's own startup was not re-verified (no new GPU load).
Known limits (5)
- Labels are one AI reviewer's (Claude's), not a lawyer's; hard judgment calls are where it errs: 4 of 34 hard privileged held-out emails would have been produced.
- About 38% of held-out real email goes to attorney review (26% where the label is clear).
- The hosted gateway route slows sharply when the shared gateway is loaded (one 12-email run took 938 s).
- Attachments must be sent as separate documents; no OCR, PDF or native file parsing in this version.
- Partial privilege (redacting part of a document) is not proposed; such documents go to review.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; the people map, grouping and leak rules run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the privilege review and privilege log API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real client material belongs on your own hardware.
A typed privilege call for every email or memo, with the lawyer involved, a grounded reason and a calibrated probability; log entries that describe without disclosing; the same call for every copy; one signed record.
Send the documents of a production set and say who the lawyers are. Code maps every address to a role and flags outsiders, personal accounts and lawyers who are only copied; it groups exact copies, near-duplicate drafts, threads and emails that quote each other. An open model answers three typed questions per document (the privilege call, whether it asks for or gives legal advice, whether it was prepared because of litigation) through the typed-judgment engine, and code turns them into withhold, produce or needs attorney review. The reason is checked against the document by the grounding checker. For each withheld document the model drafts a Rule 26(b)(5)(A) log description, and a leak check (code rules plus a typed judge) rejects drafts that give the advice away. Everything is sealed in a signed hash-chained record, and the lawyer's confirm-or-change decisions go into a signed ledger. It drafts; a lawyer decides.
- Deployment
- Self-host first
- Regulatory
- A review aid for lawyers, not legal advice and not a privilege determination. Fed. R. Civ. P. 26(b)(5)(A) requires a party withholding documents as privileged or work product to expressly claim it and describe each document in a way that lets others assess the claim without revealing the protected content; the attorney who signs discovery responses certifies them under Rule 26(g), so every call and log entry here is a draft for that attorney. A wrong 'not privileged' can waive privilege; Fed. R. Evid. 502(b) protects inadvertent disclosure only where reasonable steps were taken, and a Rule 502(d) order (non-waiver regardless of care) is the safer footing for any AI-assisted review. Courts have judged generative-AI review by the same reasonableness and proportionality standard as earlier technology-assisted review (Schulte v. LinkedIn, N.D. Cal., 30 Jun 2026); the signed record and ledger document which model, prompt and lawyer produced each call. Confidentiality: ABA Formal Opinion 512 (29 Jul 2024) asks lawyers to understand how an AI tool uses what they put in, to protect client information, and to get informed consent before putting it into a self-learning tool; United States v. Heppner (S.D.N.Y., Feb 2026) held a defendant's chats with a consumer AI service not privileged, partly because its terms allowed disclosure to third parties. Run real matters on your own machine; the hosted demo is for the fictional and public sample sets only. Model licence: Apache-2.0 (Qwen3.8-27B). The Enron emails are public records released by FERC (2003), used here as a small research sample. Checked 25 Sep 2026.
Text description
Emails and memos and the people map (lawyers, client and outside counsel domains) go to the reviewer, which runs on your own machine. Code maps each address to a role, flags outsiders, personal accounts and lawyers only copied, and groups copies, drafts, threads, families and quoting emails. Qwen3.8-27B (Apache-2.0) answers three typed questions per document through the typed-judgment engine; code turns them into withhold, produce or needs attorney review; the grounding checker reads the reason against the document; for each withheld document the model drafts a log description that must pass a leak check. Outputs: calls for review, a draft privilege log and a signed record and reviewer ledger with hashes, calls and receipt ids but no document text. Each hosted model call gets a receipt that our gateway countersigns.
At a glance
- Data retention
- Documents are held in memory for the request; the run (calls, log, no document text) for one hour, for the token or key that made it. Nothing is written to disk; logs carry counts only.
- What leaves the box
- Self-hosted: nothing. Hosted demo: the documents go to the model through our gateway, which is why the hosted demo is for the sample sets only.
- Model cost per 1,000 documents
- A dollar or two per thousand documents on the direct route at the gateway's list price (measured); GPU time only when self-hosted.
- Input
- Emails or memos as JSON: id, date, from, to, cc, bcc, subject, body, thread_id, family_id, parent_id, bates. Up to 25 per request hosted, 12,000 characters each; set DECOSA_PRIVILEGE_MAX_DOCS on your own box.
- Output
- Per document: withhold / produce / needs review, p(privileged), the lawyers, a grounded reason and flags. A Rule 26(b)(5)(A) log as CSV, a review memo, a signed record and a signed ledger of each lawyer's decision.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
one 32 GB card, self-hosted
The same model and prompts on a single RTX 5090, one call per question with log-probabilities. The grounding check of the reason can be switched off (DECOSA_PRIVILEGE_GROUNDING=0) to save one call per document.
- Models
- decosa-api privilege module (decosa_api/verticals/privilege)
- Qwen3.8-27B (NVFP4)
- Hardware
- 1x RTX 5090 32 GB
- Quality evidence
- Calls on held-out Enron emailnot measured yetnot measured yet
- Latency
- estimate: not timed on a 5090.
- Verification
- Proof: partialSelf-host onlySelf-hosted: calls and records are signed by your own box, not countersigned by the gateway.
- In the hosted demo
Standard
one 96 GB card (measured; hosted demo)
Qwen3.8-27B on an RTX PRO 6000. Self-hosted it uses log-probabilities (about 5 calls per document); the hosted demo, whose gateway does not pass log-probabilities through yet, samples the privilege call 5 times instead.
- Models
- decosa-api privilege module (decosa_api/verticals/privilege)
- Qwen3.8-27B (NVFP4)
- Hardware
- 1x RTX PRO 6000 96 GB
- Quality evidence
- Privileged vs not, model call, 149 held-out Enron emails (64 privileged by our labels): accuracy / precision / recall / AUROC of p(privileged)83.9% / 82.3% / 79.7% / 0.908docs/evals/privilege-log.md, direct route with logprobs; labels written by Claude (an AI agent), not a lawyer
- After the review rules: privileged emails the tool would have produced / non-privileged it would have withheld / sent to attorney review4 of 64 (6.2%) / 3 of 85 / 57 of 149 (38%); decided without review: 92, of which 92.4% agree with the labelsdocs/evals/privilege-log.md, same run
- The same, on the 92 held-out emails whose label we marked clear (not a judgment call)accuracy 95.7%, 0 privileged produced, 0 wrongly withheld, 26% to reviewdocs/evals/privilege-log.md; all 4 misses and 3 over-withholds were on the 57 emails we marked as hard calls
- Four-way call (attorney-client / work product / both / not privileged), held out80.5% exact, Cohen's kappa 0.65docs/evals/privilege-log.md
- Planted leaky log descriptions caught (21 leaky, 16 clean, synthetic documents)19 of 21 caught, 0 of 16 false alarms (code rules alone 12 of 21; the judge catches paraphrases)docs/evals/privilege-log.md, leak-check stage, direct route
- Drafted log descriptions that leaked (87 on held-out Enron, 18 on the synthetic set)1 first draft flagged (a subject phrase copied from the email), rewritten; 0 in the final log by the code rulesdocs/evals/privilege-log.md
- Synthetic set (28 emails): labels met / expected flags and groups met / consistency across copies, drafts and threads / same decision on a re-run0 privileged produced, 0 wrongly withheld, 3 to review / 12 of 12 / 7 of 7 groups consistent / 28 of 28 (27 of 28 same privilege type)docs/evals/privilege-log.md
- Hosted gateway route on the synthetic set (sampled call): privileged produced / wrongly withheld / to review / expectations met0 / 0 / 3 / 12 of 12; same privileged-vs-not accuracy as the direct route (96.4%), AUROC 0.974docs/evals/privilege-log.md, hosted route stage
- Latency
- measured on our server, shared card: about a second per document on the direct route; the hosted gateway route took a couple of minutes for the demo when the gateway was quiet.
- Verification
- Proof: strongHosted: every call has its own gateway-signed receipt, listed in the signed record.
- Needs more compute
Wanted: the best setup
a larger second judge on your own hardware
GLM-5.3-Flash re-reads the calls Qwen3.8-27B is least sure of, from a different model family. Privileged documents stay on your hardware, never on community providers. Not served yet.
- Models
- decosa-api privilege module (decosa_api/verticals/privilege)
- Qwen3.8-27B (NVFP4)
- GLM-5.3-Flash
- Hardware
- Your own hardware: 2x 96 GB cards (NVFP4, about 170-186 GB, unconfirmed) or a Mac with 192 GB or more (MLX 4-bit, 165 GB). Estimate; GLM's SGLang SM120 build hung on our server.
- Quality evidence
- accuracy and AUROC, same protocol as standardnot measured yet
- Latency
- not measured yet
- Verification
- No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Also runs on
- Fast option-token judge for volumeGemma-4-26B-A4B-itnot servedA 4B-active judge read through option-token probabilities: cheaper per call, not better. Page 32 suggests the dense Qwen3.5-9B (Apache-2.0) instead. The real blocker is logprob passthrough on the gateway. Hardware: 1x RTX PRO 6000 96 GB (BF16 weights are 49 GB).
We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Reviewer: people map, waiver flags, duplicate and thread grouping, review rules, leak rules, consistency, signed record and ledger (no model; CPU)decosa-api privilege module (decosa_api/verticals/privilege) 0 GBProof: partial | LiteStandardWanted | 0 GB | Proof: partial | |
| ||||
Model: the typed privilege call, the two element questions, the grounding check of the reason, the log description and the leak judgeQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | LiteStandardWanted | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Faster model for the typed calls at volume (alternate)Gemma-4-26B-A4B-itgoogle/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab) 26B (4B active) · 49 GBNo proof yetSelf-host only | Alternate | 26B (4B active) · 49 GB | No proof yetSelf-host only | |
| ||||
Second judge on the calls the 27B is least sure ofGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab) 321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only | Wanted | 321B (18B active) · about 170 GB (estimate) | No proof yetSelf-host only | |
| ||||
Tools, services and hardware
Tools
- Enron email corpus (CMU copy, via the Hugging Face dataset corbt/enron-emails) (opens in a new tab)Public record released by FERC in 2003; distributed by CMU as a research resource, with a request to respect the privacy of the people in it
The eval's real email: 200 messages sampled from Enron lawyers' mailboxes and labelled by hand (50 dev, 150 held out). Eight of them are the public sample set in the demo.
- Harborline synthetic setSynthetic, written for Decosa (CC0); fictional people and .example domains
28 emails with labels and expected flags: an exact duplicate, a near-duplicate draft, advice forwarded to an outside consultant and to a personal account, a lawyer only copied, the other side's settlement letter. Twelve are the demo sample.
- scripts/privilege_eval.py and docs/evals/privilege-log.mdApache-2.0
Rebuilds the split, runs both sets and the planted-leak set through the API, and writes the metrics, the repeat check and the cost per 1,000 documents.
- POST /record/verifyApache-2.0
Checks the signed record or ledger and names the first entry that was changed. The console also verifies it in your browser.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /privilege/info, /privilege/samples; POST /privilege/review (SSE or JSON); POST /privilege/runs/{id}/decisions; GET /privilege/runs/{id}/export?format=csv|md|record|ledger. Keeps documents in memory for the request and runs for one hour, never on disk; logs counts only.
- vLLM:8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly with logprobs (self-host).
Hardware
- 1x RTX 5090 32 GB Fits
Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for a 12,000-character email (about 4k tokens). Estimate: same model and prompts as the measured card, not run here on a 5090.
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server: the eval and the hosted demo ran on this card, shared with other services the whole time.
Latency per lane
- one document, direct route with logprobs (self-host configuration), 6 calls in flight1.4 s
Measuredmeasured on our server 2026-09-25: 149 held-out Enron emails in 215 s on a shared card (about 5 calls per document)
- 12-email demo sample, hosted gateway route (privilege call sampled 5 times, about 9 calls per document)121.0 s
Measuredmeasured on our server 2026-09-25: 85-135 s over 5 runs with the gateway quiet; 938 s once while other evaluation jobs loaded the shared gateway
- 12-email demo sample, self-hosted (clean clone, container build, direct route with logprobs)17.0 s
Measuredmeasured on our server 2026-09-25: 58 calls, one run
Notes
- Withhold or produce is decided on p(privileged) = 1 - p(not privileged), not on the most likely option: attorney-client and 'both' overlap, so a document the model is sure is privileged can split its probability between them. Withhold at 0.7 or more, produce at 0.3 or less, and a lawyer decides in between. The band was set on the 50-email dev split and not changed after the held-out run.
- A document also goes to review when an outsider or a personal account is on a possibly privileged document (waiver), when the element answers disagree with the call, when a privileged call has no confirmed lawyer on it, when a near-duplicate was called differently, or when a document that would be produced contains most of a withheld one's words.
- Exact copies are judged once and get the same call; each copy's entry says so. Threads with both withheld and produced messages are fine unless a produced message quotes a withheld one, which sends it to review.
- The leak check runs on every drafted description: a run of 6 or more words copied from the document, a figure from it, or wording that reports the advice ('advising that ...') is a leak, and a typed judge catches paraphrases. A leaky draft is rewritten once with the findings; if that one leaks too, a fixed generic description is used and the entry says so.
- The people map is only as good as the lawyer list and domains you give it: a lawyer missing from the list is treated as an employee, and signature blocks that look like a lawyer's are shown as suggestions to confirm, never used on their own.
- Attachments are judged as their own documents; send them with parent_id or family_id so the family check can compare them.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa privilege reviewer on this machine
You are setting up a privilege review assistant for e-discovery. It takes emails and memos from a production set and a
list of the client's lawyers, and returns a typed call per document (attorney-client, work product, both, not
privileged, or needs attorney review) with the lawyers involved, a reason checked against the document and a
probability. It drafts privilege-log entries that pass a leak check, gives exact copies the same call, flags outsiders
who could waive privilege and produced emails that quote withheld ones, and seals everything in a record signed by this
box's own key. Work step by step, show me each command before you run anything with `sudo`, and stop to ask if a check
fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/privilege-log.zip (3 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py privilege-log` (the api image carries the same bundle under /app/rehearsal/privilege-log/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py privilege-log --bundle privilege-log.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the general counsel's legal advice (HL-002) is withheld as privileged or sent to review", "the ops update with a lawyer only copied (HL-021) is produced", "HL-021 is flagged: a lawyer only copied"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- Model: Qwen3.8-27B (Apache-2.0). Everything else is CPU code in decosa-api (AGPL-3.0-or-later): the people map, grouping,
the leak rules, the typed-judgment engine and the grounding checker.
- These are privileged documents. Bind every port to 127.0.0.1 and do not
send documents to any hosted API. The service keeps documents in memory for the request and runs for one hour,
writes nothing to disk and logs counts only; keep it that way.
- It drafts; a lawyer decides. Every call and every log entry is for attorney review, and the lawyer who serves the log
certifies it (Fed. R. Civ. P. 26(g)). Say so wherever you show results.
## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 is about 20 GB of weights plus KV cache; an email up to
12,000 characters goes to the model whole, about 4k tokens). Driver 570 or newer; Blackwell cards run NVFP4, older
cards use the FP8 weights.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
from the official Docker and NVIDIA repositories after asking me, then run
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 30 GB free for the model and images.
## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
`git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a release that contains
`decosa_api/verticals/privilege/`, and run `docker build -f docker/api/Dockerfile -t decosa-api:local .`
- `vllm/vllm-openai:v0.29.0` for the model; weights `nvidia/Qwen3.8-27B-NVFP4` (or `Qwen/Qwen3.8-27B-FP8`).
## 3. docker-compose.yml
Write this in `~/decosa/privilege/`:
```yaml
name: decosa-privilege
services:
llm:
image: vllm/vllm-openai:v0.29.0
command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
"--enable-prefix-caching", "--max-logprobs", "20"]
ports: ["127.0.0.1:8114:8000"]
volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
api:
image: ${DECOSA_REGISTRY}/decosa-api:<tag> # or decosa-api:local
ports: ["127.0.0.1:8445:8445"]
environment:
DECOSA_HOST: 0.0.0.0
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_BUDGET_LLM_TOKENS: "400000" # per session; a document uses about 800 generated tokens
DECOSA_PRIVILEGE_MAX_DOCS: "200" # per request on your own box (the hosted limit is 25)
DECOSA_PRIVILEGE_MAX_CONCURRENT: "2"
volumes: ["decosa-data:/data"] # a named volume: the image runs as uid 10001, so a root-owned bind mount fails
depends_on: { llm: { condition: service_healthy } }
healthcheck:
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/privilege/info', timeout=4)"]
interval: 30s
retries: 10
volumes:
decosa-data: {}
```
On the direct route the model server returns log-probabilities, so each question is one call with the best-ranked
probability (the hosted demo, without them, samples each question five times). Prefix caching matters: the three
questions for one document share it as a prefix. On first start the api service creates this box's Ed25519 key in
the `decosa-data` volume under `attest/` (mode 0600). Back the volume up and never print the key. Records and model calls are signed with it: an
attestation by me, the operator, not a proof.
## 4. Smoke test
1. `curl -s localhost:8445/privilege/info | jq '{method, calls, limits: .limits.documents}'` shows `logprobs` as the method.
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"privilege-log"}' | jq -r .token)`.
3. Run the fictional sample:
`curl -s -XPOST localhost:8445/privilege/review -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample":"harborline-dispute","stream":false}' > run.json`.
Expect: HL-004 is a copy of HL-002 with the same call; HL-003 (the general counsel's advice forwarded to an outside
consultant) is "needs_review" with a third-party reason; HL-012 (the other side's settlement letter) and HL-021 (an
ops update with a lawyer copied) are "not_privileged"; HL-013 and HL-015 are near-duplicates with the same call.
`jq '.log.entries[] | {doc, privilege, description}' run.json`: no description quotes the advice or its figures.
4. Stream it with `-H 'accept: text/event-stream' -N` and `"stream": true`: `ready`, `receipt` events, `document`
events, then `consistency`, `log`, `report`, `done` and `budget`.
5. `jq '{record: .report.record}' run.json | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @-`
must say `ok: true`. Change one `final` in the record's entries and verify again: it must fail and name that entry.
6. Exports: `curl -s localhost:8445/privilege/runs/$(jq -r .run_id run.json)/export?format=csv -H "authorization: Bearer $T"`
is the log; `format=md` is the review memo.
7. Time it and tell me. On our RTX PRO 6000, shared with other work, the 12-email sample took 17 s on the direct route
(58 calls, one per question); through the hosted gateway, with 5 samples of the privilege call, it takes about 2 minutes.
## 5. Point your review workflow at the local API
Send each batch (up to `DECOSA_PRIVILEGE_MAX_DOCS`) with your lawyer list and domains to `POST /privilege/review`,
put the `needs_review` documents in front of a lawyer first, record their decisions with
`POST /privilege/runs/{id}/decisions` (`{"decisions": [{"doc", "decision": "confirm"|"change", "call"?, "reviewer", "note"?}]}`),
and keep the signed ledger (`export?format=ledger`) with the matter. For the site, set
`NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in `.env.local`. Contract: `API_CONTRACT.md`, section "Privilege review".
Off, and it should stay off on a box that holds privileged documents: joining serves other people's requests on this
GPU. Only on a separate machine, and only with my explicit yes, follow the provider guide at `/provide` on the site.Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- A privilege log drafted on your own GPU: a typed privilege call for every email or memo, with the lawyer involved, a grounded reason and a calibrated probability; Rule 26(b)(5)(A) log entries that describe without disclosing; the same call for every copy; and one signed record. It drafts; a lawyer decides.
- Who it's for
- Teams in legal.
- Where it runs
- Self-host for real matters; hosted for the demo sets only
- Key numbers
On 149 held-out emails from Enron lawyers' mailboxes the privileged-or-not call was 83.9% accurate (AUROC 0.908), with 38% sent to attorney review. The labels are one AI reviewer's, not a lawyer's, and on hard calls the model alone is barely better than chance, so the review queue is the safeguard.
- 83.9% / 82.3% / 79.7% Privileged vs not, model call: accuracy / precision / recall (test split, n = 149)
- 0.908 / 0.058 AUROC of p(privileged) / ECE (test split, n = 149)
- 4 (6.2% of privileged) Privileged emails the tool would have produced (waiver risk) (test split, n = 64)
- 121.0 s Median end-to-end run, hosted (QA sweep 2026-09-25)
- Models
- Qwen3.8-27B
- Where
- Self-host for real matters; hosted for the demo sets only
- Checks
- Receipt per call; leak check; signed record and reviewer ledger
- Industry
- Legal
- Runs
- Self-host
- Output
- Structured data · Signed record or verdict
- Data
- Privileged or legal
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Built from
- Typed judgment · Signed record
Questions people ask
Does it decide what is privileged?
No. It drafts; a lawyer decides. Each document gets withhold, produce or needs attorney review from a calibrated probability: withhold at 0.7 or more, produce at 0.3 or less, and a lawyer decides everything in between. The lawyer's confirm-or-change decisions go into a signed ledger.
What goes in a privilege log?
An express claim of privilege or work product for each withheld document, and a description that lets the other side assess the claim without revealing the protected content (Fed. R. Civ. P. 26(b)(5)(A)). Here the model drafts that description for each withheld document, the lawyers on it are marked, a leak check rejects drafts that give the advice away, and the log comes out as CSV for the attorney who signs it.
What federal rule requires a privilege log?
Federal Rule of Civil Procedure 26(b)(5)(A). The attorney who signs the discovery responses certifies them under Rule 26(g), so every call and log entry drafted here is a draft for that attorney. The drafts follow the federal rule; this is not legal advice.
How accurate is it?
On 149 held-out emails from Enron lawyers' mailboxes the privileged-or-not call was 83.9% accurate (precision 82.3%, recall 79.7%), with AUROC 0.908. 57 documents (38%) went to attorney review. 4 privileged emails (6.2%, all hard judgment calls) would have been produced. The labels are one AI reviewer's (Claude's), not a lawyer's, so measure it on a labelled sample of your own matter first.
How does it keep privilege log descriptions from giving the advice away?
Every drafted Rule 26(b)(5)(A) description goes through a leak check: code rules (a run of six or more copied words, a figure from the document, wording that reports the advice) plus a typed judge for paraphrases. On 37 planted descriptions it caught 19 of 21 leaky ones with no false alarms on 16 clean ones. A draft that leaks twice is replaced by a fixed generic description, and the entry says so.
Where do the documents go?
Self-hosted, nowhere: the model (Qwen3.8-27B, Apache-2.0) and the reviewer run on your machine. The hosted demo sends documents through our gateway, which is why it takes only the fictional and public sample sets. Run real matters on your own box.
What about duplicates, threads and outsiders on the email?
Exact copies are judged once and get the same call. A document goes to review when an outsider or a personal account is on a possibly privileged document, when a near-duplicate was called differently, or when a produced message quotes a withheld one.
What does it cost to run?
$1.49 per 1,000 documents in model calls on the direct route at the gateway's list price (measured on 149 emails); GPU time only when self-hosted.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Privilege review and privilege log
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…