Skip to content
decosa
LabsHostedSelf-hostMacSelf-host first for real data

Redact a public-records release

A release PDF with the withheld text truly removed, each redaction marked with its exemption and reason, plus the index and a draft response letter.

Held-out test70 of 71 (98.6%)Planted sensitive passages redacted (test sets re-run after fixes, so not a clean held-out number)
On production40 smedian on production (2026-09-25); slower when the service is busy
List price~$0.77 per 1,000 recordsmeasured, at list price

Built on: Typed judgment, Grounding, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; pattern finders, the release PDF and its check run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the public-records desk API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
foia-desk

Use the hosted API

# Decosa Public-records desk: use the hosted API

You are wiring Decosa's public-records desk into this project. It takes a FOIA or state public-records request and the
agency's records, and returns a search scope, a responsiveness call per record, proposed redactions (each with its
exemption, category and a reason that does not repeat the withheld words), a release PDF with the withheld text removed
and checked, an index of withheld information, a response letter draft and a signed record. Each model call has its own
signed receipt. Use only what is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- The hosted API is for the fictional and public sample sets and for testing your integration. Unredacted records carry
  personal data: for real requests use the self-host prompt instead. Never send real records here.
- It proposes; the records officer decides. Show every redaction, exemption and letter as a draft, never as a decision.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "foia-desk"}` returns `{"token", "expires_at", "budget"}`.
   a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); each session has a token allowance (its `budget`) (about 15 records). Over a limit: HTTP 429 with `Retry-After`.
3. One review at a time per demo token (409 otherwise).

## Endpoints
- `POST /foia/scope` (token) `{"request": {"text", "id"?, "received"?, "requester"?, "agency"?}}` → `{scope, text, receipt}`:
  `scope = {summary, subjects, keywords, custodians, date_from, date_to, record_types, examples, exclusions, clarify}`.
  Show it to the officer to edit. `record_types` is a limit only when the officer sets it; `examples` never excludes anything.
- `POST /foia/review` (token). Body:
  `{"request": {...}, "scope"?: {...}, "documents": [{"id", "date"?, "from"?, "to"?: [...], "cc"?: [...], "subject"?, "custodian"?, "body"}], "people"?: {"attorneys": [{"name", "email"}], "client_domains": ["agency.gov"]}, "agency_domains"?: ["agency.gov"], "jurisdiction"?: "federal" | "ca-cpra" | {"name", "items": [{"code", "label", "categories"}]}, "names"?: "private" | "none", "stream"?: true}`
  or `{"sample": "coastal-permits", "use_sample_scope": true}` to run a sample set.
  - Limits: 25 records per request, 20,000 characters each, 150,000 in all; the request text up to 6,000 characters.
  - With `"stream": true` (or `Accept: text/event-stream`) it streams `scope`, `ready`, a `receipt` after each model call,
    a `document` event per record (replace by `id`), then `consistency`, `release`, `index`, `letter`, `report`, `done`
    and `budget`. With `"stream": false`: one JSON object `{run_id, scope, documents, exemptions, consistency, release, index, letter, report, export, budget}`.
  - Each record: `{id, text, disposition: release_in_full|release_in_part|withhold_in_full|not_responsive, responsive: {answer, p, by, reason}, redactions: [{id, start, end, category, exemption, source, what, reason}], kept, withheld?, privilege?, review_reasons, copy_of?}`.
    Offsets point into `text` (header lines, a blank line, the body). A record with `review_reasons` needs the officer first.
  - `release.integrity.ok` must be true before any release goes out.
- `POST /foia/runs/{run_id}/decisions` `{"decisions": [{"redaction", "decision": "accept"|"reject"|"change", "category"?, "reviewer"}, {"doc", "decision": "confirm"|"change", "disposition"?, "reviewer"}, {"add": {"doc", "text", "category"}, "reviewer"}]}`:
  the officer's calls; the release is rebuilt and checked again.
- `GET /foia/runs/{run_id}/export?format=pdf` (the release), `csv` (index), `letter`, `md` (review memo), `record` (signed), `ledger` (signed, with the decisions).
- `POST /record/verify` (no token) `{"record": {...}}` → `{ok, summary, bad}`.
- `GET /foia/info`, `GET /foia/samples`, `GET /foia/samples/{id}`, `GET /attest/signing-key` (no token).

## Example: answer a request and save the release (Python, `pip install httpx`)
```python
import httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
s = httpx.get(f"{API}/foia/samples/coastal-permits").json()          # your own request has the same shape
scope = httpx.post(f"{API}/foia/scope", headers=H, json={"request": s["request"]}, timeout=120).json()["scope"]
# ... let the officer edit `scope` here ...
r = httpx.post(f"{API}/foia/review", headers=H, timeout=900,
               json={"request": s["request"], "scope": scope, "documents": s["documents"], "people": s["people"],
                     "agency_domains": s["agency_domains"], "jurisdiction": s["jurisdiction"], "stream": False})
r.raise_for_status()
run = r.json()
assert run["release"]["integrity"]["ok"]
for d in run["documents"]:
    print(d["id"], d["disposition"], [(x["id"], x["exemption"]) for x in d["redactions"]], d["review_reasons"])
pathlib.Path("release.pdf").write_bytes(httpx.get(f"{API}{run['export']['pdf']}", headers=H).content)
```

## Honest limits
- The calls are drafts from an open model and code rules. On labelled synthetic records they miss some personal details
  (most often part of a medical detail) and some responsive records; see the measured rates on the Stack tab.
- The release is typeset from the records' text: it does not keep the original layout, and it cannot read scanned
  images. Exemption lists: federal FOIA and California CPRA are checked; a custom list is yours to check.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Public-records desk: run it yourself (containers)

You are setting up the Decosa public-records desk on this machine, so unredacted agency records never leave it. It
scopes a FOIA or state public-records request, finds the responsive records, proposes redactions with exemptions and
reasons, builds a release PDF with the withheld text removed and checked, drafts the index and the letter, and signs a
record. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/foia-desk.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py foia-desk` (the api image carries the same bundle under /app/rehearsal/foia-desk/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py foia-desk --bundle foia-desk.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the resident's comment (CP-002) is released in part", "her phone number and email address are redacted under (b)(6)", "both Social Security numbers on the dive roster (CP-007) are redacted"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
   check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
   `DECOSA_FOIA_MAX_DOCS=200`, use a named volume for `/data`, and bind every port to 127.0.0.1. Nothing in this use
   case needs the internet after the weights are downloaded.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/foia/info` shows `"method": "logprobs (one call per question)"` and the
   `federal` and `ca-cpra` exemption lists; `GET /attest/signing-key` shows this box's public key. Show me the key: it is
   what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"foia-desk"}` and run
   `POST /foia/review {"sample": "coastal-permits", "use_sample_scope": true, "stream": false}`. Expect CP-010 not
   responsive (dated before the range), CP-004 and CP-005 withheld in full under (b)(5), the resident's details in
   CP-002 redacted under (b)(6) and `release.integrity.ok` true. Download `export?format=pdf` and confirm that copying the
   text out of it gives none of the withheld words. Then `POST /record/verify` with `report.record`: `ok` must be true.
6. Report back: the public key and key id, the totals, the integrity result and how long the run took.

Off, and keep it off on a box that holds unredacted records: joining serves other people's requests on this GPU.
If I ask for it later, on a separate machine, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/foia-desk-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (32 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (32 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"foia-desk"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py foia-desk

Download the mock-data bundle (4 KB, 11 checks)expected.json

A federal FOIA request about a dredging permit, its confirmed search scope and three synthetic records from a fictional Bureau of Coastal Permits: a resident's comment with her phone and email, a dive-team roster with two Social Security numbers, and an email from before the requested date range. The comment must be released in part with the phone and email redacted under (b)(6), both SSNs redacted, the early record put out of scope by code, the release PDF must pass its recoverable-text check and the signed record must verify.

What the rehearsal checks
  • the resident's comment (CP-002) is released in part
  • her phone number and email address are redacted under (b)(6)
  • both Social Security numbers on the dive roster (CP-007) are redacted
  • both dates of birth on the dive roster are redacted
  • the record from before the requested range (CP-010) is not responsive, decided by code with no model call
  • nothing redacted can be recovered from the release PDF
  • the Vaughn index (CSV) cites (b)(6) for CP-002
  • the Vaughn index does not repeat a redacted SSN
  • the signed record verifies
  • a record with its redaction count changed no longer verifies
  • every model call has a signed receipt

Licence: Synthetic, written for Decosa (CC0): the Bureau of Coastal Permits, Tidewater Dredge Co. and every person, address, phone number and SSN are fictional.

Prompt for your coding agent

# Decosa Public-records desk: run it yourself (containers)

You are setting up the Decosa public-records desk on this machine, so unredacted agency records never leave it. It
scopes a FOIA or state public-records request, finds the responsive records, proposes redactions with exemptions and
reasons, builds a release PDF with the withheld text removed and checked, drafts the index and the letter, and signs a
record. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/foia-desk.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py foia-desk` (the api image carries the same bundle under /app/rehearsal/foia-desk/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py foia-desk --bundle foia-desk.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the resident's comment (CP-002) is released in part", "her phone number and email address are redacted under (b)(6)", "both Social Security numbers on the dive roster (CP-007) are redacted"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
   check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
   `DECOSA_FOIA_MAX_DOCS=200`, use a named volume for `/data`, and bind every port to 127.0.0.1. Nothing in this use
   case needs the internet after the weights are downloaded.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/foia/info` shows `"method": "logprobs (one call per question)"` and the
   `federal` and `ca-cpra` exemption lists; `GET /attest/signing-key` shows this box's public key. Show me the key: it is
   what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"foia-desk"}` and run
   `POST /foia/review {"sample": "coastal-permits", "use_sample_scope": true, "stream": false}`. Expect CP-010 not
   responsive (dated before the range), CP-004 and CP-005 withheld in full under (b)(5), the resident's details in
   CP-002 redacted under (b)(6) and `release.integrity.ok` true. Download `export?format=pdf` and confirm that copying the
   text out of it gives none of the withheld words. Then `POST /record/verify` with `report.record`: `ok` must be true.
6. Report back: the public key and key id, the totals, the integrity result and how long the run took.

Off, and keep it off on a box that holds unredacted records: joining serves other people's requests on this GPU.
If I ask for it later, on a separate machine, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/foia-desk-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsPublic-records desk on GeForce RTX 5090: use the Standard · one 96 GB card (measured; hosted demo) tier

The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB (fits): Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for a 12,000-character record. Estimate: same model and prompts as the measured card, not run here on a 5090.

Standard · one 96 GB card (measured; hosted demo): what changesuses estimates

  • Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Records desk: decosa-api foia module (decosa_api/verticals/foia). CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Model: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Public-records desk, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Public-records desk on my hardware

Fetch https://decosa.ai/prompts/foia-desk-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=foia-desk)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · one 96 GB card (measured; hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Records desk: decosa-api foia module (decosa_api/verticals/foia), CPU
- Model: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/foia-desk-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh

Mac prompt for your coding agent

# Decosa Public-records desk: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Public-records desk on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/foia-desk.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py foia-desk` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the resident's comment (CP-002) is released in part", "her phone number and email address are redacted under (b)(6)", "both Social Security numbers on the dive roster (CP-007) are redacted"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Records desk: date filter, pattern finders, exemption lists and reasons, reason leak check, consistency, release PDF writer and its integrity check, index, letter, signed record and ledger (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured |
| Model: the request scope, the responsiveness and deliberative calls, the redaction spans with their category, description and harm, and the privilege engine's calls | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py foia-desk`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 40 s · ~$0.012 per run · 41 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the assemble prompt's api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.

Measured cost to run: about $0.77 per 1,000 records (hosted, 30 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Verified on 2026-09-25: the image builds, the service starts healthy, info reports logprobs and both exemption lists, scoping works, the coastal-permits sample passes end to end (17 s, 41 calls, every disposition as expected, release check ok on 28 spans), the signed record verifies and fails at the changed entry, the PDF export carries no withheld text, one officer decision rebuilds the release and the ledger verifies. The model server's own startup was not re-verified (no new GPU load).

Known limits (6)
  • Labels are one AI reviewer's (Claude's), not a records officer's; the held-out sets are small and synthetic.
  • Medical details are the weakest kind: both held-out misses were medical.
  • It over-redacts company officials' names that the model reads as private, and agency staff named only in a record's text (not in its mail headers); they go to review or to the officer.
  • Runs are not identical: the same records can come back with a few redactions more or fewer from one run to the next (temperature 0 on a shared server is not bit-for-bit repeatable). The pattern finders, the date filter and the last contact check are code and do not vary; the officer decides every proposal.
  • The release is re-typeset from text: no original layout, no PDF, scan or email-archive intake, no audio or video.
  • The hosted route slows when the shared gateway is loaded.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; pattern finders, the release PDF and its check run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the public-records desk API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.
The open stack

Scope a FOIA or public-records request, find the responsive records, and propose each redaction with its exemption and a reason; release a PDF with the withheld text really removed, checked, indexed and signed.

Paste the request and send the agency's records. An open model turns the request into a search scope the officer edits; code drops records dated outside it; a typed judgment calls each remaining record responsive or not, with a calibrated probability. For responsive records, pattern finders catch identifiers (SSNs, dates of birth, licence, account and card numbers; phones, emails and home addresses unless they are official contacts), and the model proposes the rest: names of private individuals, medical details, staff recommendations written before a decision, advice from agency counsel. Code maps each proposal's category to the exemption of the chosen list (federal FOIA or California CPRA), so the same kind of text always gets the same label, and checks that no reason repeats the withheld words. Records with agency counsel on them go through the privilege engine. A detail redacted anywhere is redacted everywhere. The release PDF is typeset without the withheld text, with a box and the exemption code where each deletion was, and read back by three readers to prove nothing can be recovered. It drafts the index of withheld information and the response letter, and every call and officer decision is sealed in a signed record. It proposes; the records officer decides.

Deployment
Self-host first
Regulatory
A drafting aid for records officers, not legal advice and not a determination under any public-records law. Federal: 5 U.S.C. § 552(b)(1)-(9) lists the exemptions; an agency may withhold only where it reasonably foresees harm to an interest an exemption protects or where law prohibits disclosure (§ 552(a)(8)(A), FOIA Improvement Act of 2016); it must release reasonably segregable portions and mark the exemption at the place of each deletion where technically feasible (§ 552(b)); the deliberative process privilege does not apply to records created 25 years or more before the request (§ 552(b)(5)); an adverse determination must give at least 90 days to appeal and point to the FOIA Public Liaison and OGIS (§ 552(a)(6)(A)). California: the Public Records Act, Gov. Code § 7920.000 et seq. (recodified by AB 473, operative 1 Jan 2023), with § 7927.700 (personnel, medical or similar files), § 7927.705 (records exempt under other law, including Evidence Code privileges), § 7922.000 (the public-interest balancing test, under which California courts recognise the deliberative process privilege), § 7922.525(b) (segregable portions) and § 7922.540 (a written denial naming each person responsible). A custom list for another state is the agency's to check. Vaughn indexes (Vaughn v. Rosen, 484 F.2d 820 (D.C. Cir. 1973)) are filed in litigation; the index here is a draft for that use. Real requests involve personal data: run them on the agency's own machine; the hosted demo is for the fictional and public sample sets only. Model licence: Apache-2.0 (Qwen3.8-27B). The Enron emails are public records released by FERC (2003), used as a small research sample. Statutes checked against the official texts on 25 Sep 2026.
Architecture
Text description

A public-records request and the agency's records, with its counsel and mail domains, go to the records desk on the agency's own machine. Qwen3.8-27B (Apache-2.0) scopes the request, judges responsiveness and deliberative content through the typed-judgment engine and proposes redaction spans; records with counsel go through the privilege engine. Code drops records outside the dates, finds identifiers by pattern, maps each category to the exemption of the chosen list, checks reasons for leaks, redacts a detail everywhere, and writes a release PDF without the withheld text that three readers check. Outputs: the release PDF, an index and a letter, and a signed record and decision ledger with hashes, exemptions and receipt ids but no withheld text. Each hosted model call gets a receipt that our gateway countersigns.

Architecture

At a glance

Data retention
Records are held in memory for the request and the run (one hour, so the release can be rebuilt after the officer's decisions), for the token or key that made it. Nothing is written to disk; logs carry counts only.
What leaves the box
Self-hosted: nothing. Hosted demo: the records go to the model through our gateway, which is why the hosted demo is for the sample sets only.
Model cost per 1,000 records
Measured on the held-out test run at the gateway's list price (see Cost at list price in the measured results); GPU time only when self-hosted.
Exemption lists
Federal FOIA (5 U.S.C. § 552(b)(1)-(9)) and the California Public Records Act, checked against the statutes on 25 Sep 2026; any other state as a custom list the agency checks.
Output
Per record: responsive or not, disposition, each redaction with its exemption, category and reason. A release PDF with a passed integrity check, the index as CSV, a response letter draft, a review memo, a signed record and a signed ledger of each officer decision.
Input
The request text and the records as JSON text (id, date, from, to, cc, subject, custodian, body). Up to 25 records per request hosted; set DECOSA_FOIA_MAX_DOCS on your own box. No PDF, scan or email-archive intake in this version.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 32 GB card, self-hosted

    The same model and prompts on a single RTX 5090, one call per question with log-probabilities. The privilege engine's reason check can be switched off (DECOSA_FOIA_GROUNDING=0) to save a call per record with counsel on it.

    Models
    • decosa-api foia module (decosa_api/verticals/foia)
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX 5090 32 GB
    Quality evidence
    • Held-out test setsnot measured yetnot measured yet
    Latency
    estimate: not timed on a 5090.
    Verification
    Proof: partialSelf-host onlySelf-hosted: calls and records are signed by your own box, not countersigned by the gateway.
  • In the hosted demo

    Standard

    one 96 GB card (measured; hosted demo)

    Qwen3.8-27B on an RTX PRO 6000, through our gateway with one call per question on the hosted demo (log-probabilities when self-hosted).

    Models
    • decosa-api foia module (decosa_api/verticals/foia)
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX PRO 6000 96 GB
    Quality evidence
    • Planted personal details redacted, 48 held-out synthetic records (60 labelled: names, phones, addresses, emails, SSN, dates of birth, licence and account numbers, medical details)59 of 60 (98.3%); the pattern finders alone would catch 46.5% of all planted spansdocs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline); labels written by Claude (an AI agent), not a records officer
    • What it missedThe misses are listed per record in the eval write-up; medical details are the weakest kind.docs/evals/foia-desk.md
    • Responsiveness, 77 held-out records (48 synthetic, 29 public FERC-released Enron emails): precision / recall96.2% / 100.0%docs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline)
    • Exemption label on redacted spans / records with agency counsel withheld in full with the right exemption70 of 70 / 6 of 6 (federal (b)(5), (b)(6), (b)(4); CPRA § 7927.700, § 7927.705, § 7922.000)docs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline)
    • Redaction precision (redactions that overlap a labelled span) / text that had to be released but was boxed83.1% / 3 spans: mostly company officials' names and emails, boxed everywhere by the privacy-protective consistency step and sent to reviewdocs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline)
    • Release check: withheld spans recoverable from the PDF (built-in reader, pdfplumber, filing-preflight black-box check)0 of 95 on the held-out releases; 0 characters under any boxdocs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline)
    • How the numbers moved as design faults were fixed (the test sets were run five times; no threshold tuned on them)personal details redacted 81.0% → 78.8% → 91.7% → 96.7% → 98.3%; responsiveness recall 91.4% → 86.0% → 98.0% → 100.0% → 100.0%docs/evals/foia-desk.md, History; two fresh held-out sets were written before runs 2 and 3
    Latency
    measured on our server, shared card and gateway; the run time and the model cost per 1,000 records are in the measured results below.
    Verification
    Proof: strongHosted: every call has its own gateway-signed receipt, listed in the signed record.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Records desk: date filter, pattern finders, exemption lists and reasons, reason leak check, consistency, release PDF writer and its integrity check, index, letter, signed record and ledger (no model; CPU)decosa-api foia module (decosa_api/verticals/foia)
0 GBProof: partial
Model: the request scope, the responsiveness and deliberative calls, the redaction spans with their category, description and harm, and the privilege engine's callsQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo

Around the models

Tools, services and hardware

Tools

  • Synthetic request sets (6 sets, 69 records)Synthetic, written for Decosa (CC0); fictional agencies, people, numbers (555-01xx phones) and .example domains

    Two development sets (the console samples, 21 records: a federal permit office and a California city) and four held-out test sets (48 records) (a federal grants office, a California water district, a federal parks office, a California school district), with planted personal data, medical details, deliberative text, advice from agency counsel, official contacts that must be released, copies and out-of-scope records.

  • Enron email corpus (CMU copy, via the Hugging Face dataset corbt/enron-emails) (opens in a new tab)Public record released by FERC in 2003; distributed by CMU as a research resource, with a request to respect the privacy of the people in it

    A stand-in for agency email: 39 messages from May-August 2001 labelled for a hypothetical request about price caps (10 in the demo sample, 29 held out). Responsiveness only; no data planted.

  • scripts/foia_eval.py and docs/evals/foia-desk.mdApache-2.0

    Runs the sets through the API and scores responsiveness, planted-span recall by kind, exemption labels, redaction precision, the release check and cost.

  • POST /record/verifyApache-2.0

    Checks the signed record or ledger and names the first entry that was changed. The console also verifies it in your browser.

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /foia/info, /foia/samples; POST /foia/scope, /foia/review (SSE or JSON); POST /foia/runs/{id}/decisions; GET /foia/runs/{id}/export?format=pdf|csv|letter|md|record|ledger. Keeps records in memory for the request and the run (one hour), never on disk; logs counts only.

  • vLLM:8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly with logprobs (self-host).

Hardware

  • 1x RTX 5090 32 GB Fits

    Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for a 12,000-character record. Estimate: same model and prompts as the measured card, not run here on a 5090.

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the eval and the hosted demo ran on this card, shared with other services the whole time.

Latency per lane

  • 12-record demo sample, hosted gateway route (about 42 calls)40.0 s

    Measuredmeasured on our server 2026-09-25: 13-76 s over six runs (41-42 calls each), on a gateway shared with other evaluation jobs; 17 s self-hosted on the direct route

  • one record, hosted gateway route, 3 records in flight4.0 s

    Measuredmeasured on our server 2026-09-25: 77 held-out records in 310 s (about 2.8 calls per record)

Notes

  • Responsiveness uses a typed yes/no question: responsive at p ≥ 0.7, not responsive at ≤ 0.3, the officer decides in between (such records still get redaction proposals). Records dated outside the scope's range are dropped by code, with the reason. The question says outright that an exempt or privileged record is still responsive; before that line the model dropped counsel's emails as 'exempt'.
  • Kinds of records the request lists ('including emails and memos') are kept as examples, never as a limit, unless the officer sets them as one: the model read them as limits and dropped press releases and letters from the public.
  • A model mistake on contact details can only over-redact: phones, emails and addresses are redacted unless the model names them as official or they are on the agency's own mail domain.
  • Names of agency staff and counsel from the mail headers are not redacted even if the model proposes it, and words the request itself uses are not withheld as commercial or other non-privacy information.
  • A personal detail redacted in one record is redacted wherever it appears in the release; when one record released it as official and another redacted it, both go to the officer and the redaction wins until they decide.
  • The release PDF is re-typeset from the records' text: box widths are rounded to 4 characters so a box does not give away a name's length; the original layout is not kept, and scanned images are not read.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

foia-desk/assemble-prompt.md134 lines
# Assemble the Decosa public-records desk on this machine

You are setting up a public-records assistant for an agency's records officer. It takes a FOIA or state public-records
request and the agency's records, and returns a search scope, a responsiveness call for each record, proposed
redactions (each with its exemption and a reason), a release PDF with the withheld text removed and checked, an index
of withheld information, a response letter draft, and a record signed by this box's own key. Work step by step, show me
each command before you run anything with `sudo`, and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/foia-desk.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py foia-desk` (the api image carries the same bundle under /app/rehearsal/foia-desk/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py foia-desk --bundle foia-desk.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the resident's comment (CP-002) is released in part", "her phone number and email address are redacted under (b)(6)", "both Social Security numbers on the dive roster (CP-007) are redacted"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Model: Qwen3.8-27B (Apache-2.0). Everything else is CPU code in decosa-api (AGPL-3.0-or-later): the pattern finders, the
  exemption lists, the release PDF writer and its check, the typed-judgment engine and the privilege engine. The PDF
  check also uses pdfplumber (MIT) when it is installed in the image.
- Unredacted records carry personal data. Bind every port to 127.0.0.1 and do not send records to any hosted API. The service keeps records in memory for the request and the run (one hour,
  so the release can be rebuilt after the officer's decisions), writes nothing to disk and logs counts only; keep it that way.
- It proposes; the records officer decides. Every redaction, exemption and letter is a draft. Say so wherever you show
  results.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 is about 20 GB of weights plus KV cache; a record up to
   12,000 characters goes to the model whole). Driver 570 or newer; Blackwell cards run NVFP4, older cards use the FP8 weights.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
   from the official Docker and NVIDIA repositories after asking me, then run
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 30 GB free for the model and images.

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a release that contains
  `decosa_api/verticals/foia/`, and run `docker build -f docker/api/Dockerfile -t decosa-api:local .`
- `vllm/vllm-openai:v0.29.0` for the model; weights `nvidia/Qwen3.8-27B-NVFP4` (or `Qwen/Qwen3.8-27B-FP8`).

## 3. docker-compose.yml
Write this in `~/decosa/foia/`:

```yaml
name: decosa-foia
services:
  llm:
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
              "--enable-prefix-caching", "--max-logprobs", "20"]
    ports: ["127.0.0.1:8114:8000"]
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>     # or decosa-api:local
    ports: ["127.0.0.1:8445:8445"]
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_BUDGET_LLM_TOKENS: "400000"        # per session; a record uses about 1,100 generated tokens at most
      DECOSA_FOIA_MAX_DOCS: "200"               # records per request on your own box (the hosted limit is 25)
      DECOSA_FOIA_MAX_CONCURRENT: "2"
    volumes: ["decosa-data:/data"]                # a named volume: the image runs as uid 10001, so a root-owned bind mount fails
    depends_on: { llm: { condition: service_healthy } }
    healthcheck:
      test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/foia/info', timeout=4)"]
      interval: 30s
      retries: 10
volumes:
  decosa-data: {}
```

On the direct route the model server returns log-probabilities, so each typed question is one call with the
best-ranked probability. On first start the api service creates this box's Ed25519 key in the `decosa-data` volume under
`attest/` (mode 0600). Back the volume up and never print the key. Records and model calls are signed with it: an
attestation by me, the operator, not a proof.

## 4. Smoke test
1. `curl -s localhost:8445/foia/info | jq '{method, limits: .limits.documents, lists: (.exemption_lists | keys)}'` shows
   `logprobs` as the method and the `federal` and `ca-cpra` lists.
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"foia-desk"}' | jq -r .token)`.
3. Scope a request: `curl -s -XPOST localhost:8445/foia/scope -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample":"coastal-permits"}' | jq .scope`
   gives subjects, keywords and the dates 2026-01-01 to 2026-03-31.
4. Run the fictional sample with its confirmed scope:
   `curl -s -XPOST localhost:8445/foia/review -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample":"coastal-permits","use_sample_scope":true,"stream":false}' > run.json`.
   Expect: CP-010 (dated November 2025) "not_responsive" by code; CP-011 (another permit) "not_responsive"; CP-004 and
   CP-005 (with agency counsel) "withhold_in_full" under (b)(5); CP-002 and CP-008 have the resident's name, address,
   phone and email redacted under (b)(6); CP-007 has both SSNs and dates of birth redacted; CP-001 and CP-009 keep the
   company's and the agency's office phones and addresses; CP-012 is a copy of CP-003 with the same redactions.
   `jq '.release.integrity' run.json` must say `"ok": true` with `recovered: []` and `hidden_chars_under_boxes: 0`.
5. Stream it with `-H 'accept: text/event-stream' -N` and `"stream": true`: `scope`, `ready`, `receipt` events,
   `document` events, then `consistency`, `release`, `index`, `letter`, `report`, `done` and `budget`.
6. `jq '{record: .report.record}' run.json | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @-`
   must say `ok: true`. Change one `exemption` in the record's entries and verify again: it must fail and name that entry.
7. Exports: `R=$(jq -r .run_id run.json)`; `curl -s localhost:8445/foia/runs/$R/export?format=pdf -H "authorization: Bearer $T" -o release.pdf`
   (open it: boxes with exemption codes; select all and copy: none of the withheld words come out); `format=csv` is the
   index, `format=letter` the letter draft.
8. Record one officer decision and check the ledger:
   `curl -s -XPOST localhost:8445/foia/runs/$R/decisions -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"decisions":[{"redaction":"CP-002-R1","decision":"accept","reviewer":"Test Officer"}]}' | jq .release.integrity.ok`
   then `export?format=ledger` verifies at `/record/verify`.
9. Time it and tell me. On our RTX PRO 6000, shared with other work, the 12-record sample took about 40 s on the hosted
   gateway route (42 calls); the direct route is faster.

## 5. Point your workflow at the local API
Per request: `POST /foia/scope` to draft the scope, let the officer edit it, then `POST /foia/review` with the scope,
the records (up to `DECOSA_FOIA_MAX_DOCS`), agency counsel and mail domains, and the exemption list (`"jurisdiction":
"federal"`, `"ca-cpra"`, or your own `{"name", "items": [{"code", "label", "categories"}]}` for another state; check a
custom list against your statute yourself). Put the flagged records in front of the officer first, record decisions with
`POST /foia/runs/{id}/decisions`, and keep the release PDF, index, letter and signed ledger with the request file. For
the site, set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in `.env.local`. Contract: `API_CONTRACT.md`, section
"Public-records desk".

Off, and it should stay off on a box that holds unredacted records: joining serves other people's requests on this
GPU. Only on a separate machine, and only with my explicit yes, follow the provider guide at `/provide` on the site.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

1 law, rule and guidance page cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
FOIA software for records officers that runs on the agency's own machine: it scopes a FOIA or public-records request, finds the responsive records, proposes each redaction with its exemption and a reason, and releases a PDF with the withheld text really removed, checked, indexed and signed. It proposes; the officer decides.
Who it's for
Teams in public sector and legal.
Where it runs
Self-host for real requests; hosted for the demo sets only
Key numbers

On the current pipeline it redacted 59 of 60 planted personal details (98.3%) at 83.1% redaction precision, and 0 of 95 withheld spans were recoverable from the release PDF. The test sets were run five times with fixes in between and labelled by an AI reviewer, so these are not clean held-out numbers.

  • 5/8 PII recall on a set new to the pipeline (bus-contract, run 3) (held out, n = 8)
  • 59 of 60 (98.3%) Personal-data spans redacted (PII recall), final pipeline (test split, n = 60)
  • 70 of 71 (98.6%) Planted spans redacted, final pipeline (test split, n = 71)
  • 40.0 s Median end-to-end run, hosted (QA sweep 2026-09-25)
All results, datasets and caveats
Models
Qwen3.8-27B
Where
Self-host for real requests; hosted for the demo sets only
Checks
Receipt per call; release check; signed record and decision ledger
Industry
Public sector · Legal
Output
Signed record or verdict · Structured data
Data
Personal data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

What does the FOIA software do with a request?

It turns the request into a search scope the officer edits, calls each record responsive or not with a calibrated probability, proposes each redaction with its exemption (federal FOIA or the California CPRA) and a reason that names the harm without repeating the withheld words, and drafts the index and the response letter. It proposes; the records officer decides.

Can a FOIA response be redacted?

Yes. Under federal FOIA an agency may withhold information only under the nine exemptions in 5 U.S.C. § 552(b), and only where it foresees harm to an interest an exemption protects or the law prohibits disclosure. It must release reasonably segregable portions and mark the exemption at each deletion where technically feasible. Here each redaction is proposed with its exemption and a reason, and the records officer decides. Not legal advice.

How do you know the redacted text is really gone?

The release PDF is typeset without the withheld text, with a box and the exemption code where each deletion was, then read back by three readers. On the test sets, 0 of 95 withheld spans were recoverable. Box widths are rounded so a box does not give away a name's length.

How accurate are the redactions?

On the current pipeline it redacted 59 of 60 planted personal details (98.3%), and redaction precision was 83.1%, mostly from over-redacting company officials' names. Medical details are the weakest kind. The test sets were run five times with fixes in between and labelled by an AI reviewer, so these are not clean held-out numbers; an officer still reads every released record.

Does it work for state public-records laws?

It ships federal FOIA (5 U.S.C. § 552(b)(1)-(9)) and the California Public Records Act, checked against the statutes on 25 Sep 2026. Any other state works as a custom list the agency checks.

Where do the records go?

Self-hosted, nowhere: the model is Apache-2.0 and runs on the agency's machine. The hosted demo sends records through our gateway, so it is for the fictional and public sample sets only. There is no PDF, scan or email-archive intake in this version.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Public-records desk

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.