Digest a deposition
A topic-by-topic digest where every sentence links to the page:line it came from, plus where the witnesses contradict each other.
Built on: Grounding, Signed record
1. Pick a sample
Your own documents: run the tool on your own hardware, or request confidential access. The demo takes samples or made-up data only.
2. Run it
On production the sample took 40 s (median of 5 runs, 2026-09-30; slowest 44 s). Slower when the service is busy.
Result
The answer appears here first, then what it found, the draft, and how long it took. Sample: Mock civil deposition: Alvarez v. Northgate, two witnesses (fictional).
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB); no GPU for the parser, flags and export.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the deposition and hearing digest API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real client material belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- deposition
Use the hosted API
# Decosa Deposition and hearing digest: use the hosted API
You are adding Decosa's deposition and hearing digest to this project. Decosa runs an open model (Qwen3.8-27B) behind
the Decosa API; every model call comes with a signed receipt. Use only the endpoints below. If you
need something that is not listed, stop and ask me; do not guess endpoints.
**The hosted API is for public-record and fictional transcripts only.** Deposition transcripts are often under a
protective order and may be privileged. Before sending anything else, stop and tell me to self-host instead (the
"Self-host" tab). Not a certified transcript and not legal advice.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`, created on the tool page, shown once). Keep it in `DECOSA_API_KEY`, never in code.
Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "deposition"}` returns `{"token", "expires_at", "budget"}`.
A demo session has 20,000 generated tokens: enough for about 400 transcript lines with the contradiction check.
3. Over a rate limit the API answers 429 with `Retry-After` (seconds).
## Endpoints
- `GET /deposition/samples` (no token): public-domain hearing excerpts and a fictional deposition pair, with sources and
`sets` of ids that go together. `GET /deposition/samples/{id}` returns one with its `text`.
- `GET /deposition/policy` (no token): what it is and is not, limits, privilege notes.
- `POST /deposition/parse` (token): body is `text/plain`, `application/json {"text", "cite_as"?, "witnesses"?}` or
`application/pdf` (up to 12 MB). Returns `{layout, pages, lines, range, witnesses, speakers, warnings, sha256, text, preview}`
without any model call. `text` is the transcript in the canonical layout (`Page N` then ` n line`). Layouts:
`numbered` (the reporter's own page:line), `page:line` (condensed `12:05 text`), `printed-pages` (page numbers only;
lines are counted per page) and `plain` (no numbers; paginated 25 lines a page by Decosa, with a warning).
- `POST /deposition/digest` (token): `{"transcripts": [{"text" | "sample_id", "cite_as"?, "title"?, "witnesses"?: [..]}], "contradictions"?: true}`,
up to 4 transcripts and 4,000 lines in all. Returns Server-Sent Events:
- `ready` `{docs, estimate_tokens}`; then `lane` events `{lane, title, body, data}`. Each lane event is that lane's full
current state: replace, don't append.
- `transcript`: `data.docs` with each file's sha256 (every cite is anchored to it).
- `digest`: `data.docs[].topics[].sentences[]` = `{i, text, cite, cite_as, speaker, label, reason, quote}`; `cite` is a
page:line range such as `22:3–9`; `label` is `null` until checked.
- `verifier`: `data` = `{counts, checked, total, flagged, final}`. Labels: `SUPPORTED`, `PARTIAL` (a detail such as a
hedge or a number is not in the cited lines), `UNSUPPORTED` (the cited lines do not say it, or the cite does not exist),
`ERROR` (not checked).
- `flags`: non-answers, exhibits, objections and instructions not to answer, found by fixed patterns (no model).
- `contradictions` (2+ transcripts): `data.pairs[]` = `{a, b, verdict: INCONSISTENT|CONSISTENT|NOT_COMPARABLE, why}`, each side
with `cite_as`, `cite`, `speaker`, `text` and the cited `lines`.
- `record`: a signed `decosa.record.v1` record; check it with `POST /record/verify`.
- `receipt` after every model call (`status: "signed"`, look it up at `GET /receipts/{id}`); `done`
`{digest_id, summary, export}`; a final `budget`.
- Errors: 400 bad body, 402 when the token budget is below the estimate (the body has `estimate`), 413 over a limit,
422 a transcript with no lines.
- `GET /deposition/digests/{id}` and `GET /deposition/digests/{id}/export?format=md|docx` (same token or key only): the
digest, or a Markdown or Word file. Exports hold back `UNSUPPORTED` sentences and list them at the end; `PARTIAL` ones are
marked `[check]`. Digests stay in server memory for one hour, never on disk.
## Test
Get a token, `POST /deposition/digest` with `{"transcripts": [{"sample_id": "fictional-pruitt"}, {"sample_id": "fictional-okafor"}]}`.
Pass if you see `ready`, every lane above, `receipt` events and `done`, and the contradictions lane has the time of the
fall (2:15 pm against 11:30 am) as `INCONSISTENT`. Then download the Word export.
## Rules
- Show each sentence with its cite and its check label; never show an `UNSUPPORTED` sentence as if it were checked.
- Say what the check proves: each sentence is supported by the lines it cites. It does not prove the digest covers
everything important; a lawyer still reads the transcript.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa Deposition and hearing digest: run it yourself (containers)
You are setting up Decosa's deposition and hearing digest on this machine: a page:line digest where every sentence is
checked against the lines it cites, a cross-witness contradiction finder and Word or Markdown export, on Qwen3.8-27B.
Nothing is sent to Decosa's hosted API. The local service speaks the same API as the hosted one. This is the right
route for real matters: transcripts under a protective order or with privileged content stay on this box.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Hardware: 1x RTX PRO 6000 96 GB (measured), or a 48 GB card with the FP8 model (not measured). About 80 GB of disk.
Linux x86_64. Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/deposition.zip (4 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py deposition` (the api image carries the same bundle under /app/rehearsal/deposition/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py deposition --bundle deposition.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the transcript parses into 3 pages of numbered page:line lines", "the parser recognises the numbered deposition layout", "the digest finishes without errors"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Rules you must keep
- Keep `DECOSA_LLM_ROUTE=direct`. Never point this box at a hosted gateway while it holds client
transcripts. Keep the API on 127.0.0.1.
- Do not remove the cite check or the export's held-back section.
- Not a certified transcript and not legal advice. The audio path gives an uncertified rough transcript.
## Steps
1. Docker and the NVIDIA Container Toolkit: if `docker compose version` or
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi` fails, install them from the official
Docker and NVIDIA instructions.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Read it.
It needs the `api` and `llm` services; `diarize` only if I want the audio path.
3. Start: `docker compose pull && docker compose up -d`. Wait for `curl -fsS http://localhost:8445/healthz` to show `"llm": true`.
4. Smoke test: get a token with `POST /demo/session {"vertical":"deposition"}`, then `POST /deposition/digest` with the
two fictional samples (`fictional-pruitt`, `fictional-okafor`). Check that every lane streams, each model call has a
`receipt` event with status `attested` (signed by this box), and the contradictions lane flags the time of the fall.
5. Download `GET /deposition/digests/{id}/export?format=docx` and open it.
6. Report back: health, the digest id, the cite-check counts and the contradictions found.
Off by default. Leave it off for client work. Only public-record transcripts could ever go to the network.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/deposition-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVIDIA NVFP4) needs a GPU.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (61.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns
The standard tier fits (61.6 of 192 GB). The best tier fits too.
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits (32 of 96 GB).
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits (32 of 64 GB).
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API, the speech model and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with both models loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "asr": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"deposition"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py deposition
Download the mock-data bundle (4 KB, 9 checks)expected.json
Two short fictional deposition transcripts (page:line layout) from the same fictional case. The witnesses disagree about when the fall happened (2:15 pm against 11:30 am). The digest must cite page:line, find that contradiction, export to Markdown and end in a signed record that verifies.
What the rehearsal checks
- the transcript parses into 3 pages of numbered page:line lines
- the parser recognises the numbered deposition layout
- the digest finishes without errors
- the digest has every lane: transcript, flags, digest, verifier, contradictions, record
- the time-of-fall contradiction (2:15 against 11:30) is found
- the Markdown export cites the depositions by page and line
- the signed record verifies
- a record with one entry edited no longer verifies
- every model call has a signed receipt
Licence: Fictional: written for Decosa, no real case or people. CC0.
Prompt for your coding agent
# Decosa Deposition and hearing digest: run it yourself (containers)
You are setting up Decosa's deposition and hearing digest on this machine: a page:line digest where every sentence is
checked against the lines it cites, a cross-witness contradiction finder and Word or Markdown export, on Qwen3.8-27B.
Nothing is sent to Decosa's hosted API. The local service speaks the same API as the hosted one. This is the right
route for real matters: transcripts under a protective order or with privileged content stay on this box.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Hardware: 1x RTX PRO 6000 96 GB (measured), or a 48 GB card with the FP8 model (not measured). About 80 GB of disk.
Linux x86_64. Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/deposition.zip (4 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py deposition` (the api image carries the same bundle under /app/rehearsal/deposition/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py deposition --bundle deposition.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the transcript parses into 3 pages of numbered page:line lines", "the parser recognises the numbered deposition layout", "the digest finishes without errors"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Rules you must keep
- Keep `DECOSA_LLM_ROUTE=direct`. Never point this box at a hosted gateway while it holds client
transcripts. Keep the API on 127.0.0.1.
- Do not remove the cite check or the export's held-back section.
- Not a certified transcript and not legal advice. The audio path gives an uncertified rough transcript.
## Steps
1. Docker and the NVIDIA Container Toolkit: if `docker compose version` or
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi` fails, install them from the official
Docker and NVIDIA instructions.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Read it.
It needs the `api` and `llm` services; `diarize` only if I want the audio path.
3. Start: `docker compose pull && docker compose up -d`. Wait for `curl -fsS http://localhost:8445/healthz` to show `"llm": true`.
4. Smoke test: get a token with `POST /demo/session {"vertical":"deposition"}`, then `POST /deposition/digest` with the
two fictional samples (`fictional-pruitt`, `fictional-okafor`). Check that every lane streams, each model call has a
`receipt` event with status `attested` (signed by this box), and the contradictions lane flags the time of the fall.
5. Download `GET /deposition/digests/{id}/export?format=docx` and open it.
6. Report back: health, the digest id, the cite-check counts and the contradictions found.
Off by default. Leave it off for client work. Only public-record transcripts could ever go to the network.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/deposition-mac.md instead.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsDeposition and hearing digest on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Standard · the hosted demo, one 96 GB card: what changesuses estimates
- Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Digest writer, cite checker and contradiction...: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)
- Self-host only: MOSS-Transcribe-Diarize 0.9B. ~4 GB, weights 1.8 GB (estimate). MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate.
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Deposition and hearing digest, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Deposition and hearing digest on my hardware Fetch https://decosa.ai/prompts/deposition-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=deposition) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 32 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Digest writer, cite checker and contradiction...: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. - Self-host only: MOSS-Transcribe-Diarize 0.9B (OpenMOSS-Team/MOSS-Transcribe-Diarize), 4 GB GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%), MOSS-Transcribe-Diarize 0.9B ~4 GB (13%); about 0 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/deposition-assemble.md
Or on a Mac Studio
No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where
From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh
Mac prompt for your coding agent
# Decosa Deposition and hearing digest: run it on this Mac (Apple Silicon, no NVIDIA GPU) You are setting up the Decosa Deposition and hearing digest on this Mac, natively on Apple Silicon. The models run on the Mac's GPU through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API. Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more. Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop or kill a process this setup did not start; if a port is taken, pick another one. ## Step 0: set up with a coding agent, rehearse on mock data, then go private This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works. Work in this order: 1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to "test with something realistic". 2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool, https://decosa.ai/samples/deposition.zip (4 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json), show me what is in it, and run the rehearsal against the local API: `.venv/bin/python scripts/rehearse.py deposition` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key). It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the transcript parses into 3 pages of numbered page:line lines", "the parser recognises the numbered deposition layout", "the digest finishes without errors"). Show me the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json` to make a check pass. 3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this machine. For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent can read. Switch to your own data only after the rehearsal has passed and the agent's work is done. ## What runs where | Part | On an NVIDIA GPU | On this Mac | Status | |---|---|---|---| | Digest writer, cite checker and contradiction judge | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured | | Self-host only: a recording to an uncertified rough transcript (POST /deposition/rough) | transformers on CUDA | MLX 8-bit (vanch007/mlx-MOSS-Transcribe-Diarize-8bit) on mlx-audio, scripts/mac/diarize_server.py | Runs, measured | ## Steps 1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory. 2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`. 3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`. Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me. 4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api) and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them. If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`. 5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key: show it to me, because it is what others pin to check the receipts and records this Mac signs. 6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py deposition`. It runs the tool's own sample end to end against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts. `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found. 7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`, the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`. 8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of `scripts/mac/setup.sh status`. ## Good to know - Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a self-hosted Mac. - The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published evals use. Expect small differences in wording and scores. - Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel calls; typed judgments then use sampling because oMLX returns no log-probabilities). - Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details: `docs/self-host-mac.md` in the checkout.
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 30 Sep 2026 · measured 30 Sep 2026: · p50 40 s · p95 44 s (5 runs) · ~$0.014 per run · 48 receipts
Loading the nightly status…
Self-host: verified 25 Sep 2026 · fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers
Measured cost to run: about $0.014 per deposition (hosted, 30 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The step 5 smoke passed as written in 23 s: every lane, 54 attested receipts, the 2:15 pm / 11:30 am conflict marked INCONSISTENT, the record verified and the Word export opened. A PDF transcript parsed as numbered (4:1-6:14), and with the diarize service /deposition/rough returned an uncertified rough transcript. Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified.
Known limits (3)
- Hosted numbers are the production smoke check of the sample, run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
- The claim check shows each sentence is supported by the lines it cites; it does not show the digest covers everything important.
- Hosted is for public-record and fictional transcripts only.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB); no GPU for the parser, flags and export.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the deposition and hearing digest API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real client material belongs on your own hardware.
A page:line digest where every sentence is checked against the lines it cites, plus the places where witnesses contradict each other.
For litigation associates, paralegals and litigation-support teams. Give it one to four transcripts with page:line numbering (text or PDF). It writes a topic digest where every sentence cites the page:line range it relies on, and a second model call checks each sentence against exactly those lines; unsupported sentences are struck through and held back from the Word or Markdown export. It lists non-answers and exhibits, and with two or more transcripts it pairs up testimony from different witnesses and flags answers that cannot both be true, with both cites. Transcripts are often privileged or under a protective order, so the real product runs on your own GPU; the hosted demo takes public-domain and fictional transcripts only.
- Deployment
- Self-host first
- Regulatory
- Not a certified transcript and not legal advice; the digest is a work aid and a lawyer still reads the transcript. The check shows each sentence is supported by the lines it cites; it does not show the digest covers everything important. ABA Formal Opinion 512 (29 Jul 2024): lawyers must understand how a generative AI tool uses client information, get informed consent before putting it into a self-learning tool (boilerplate in an engagement letter is not enough), and bill only for time actually spent. United States v. Heppner (S.D.N.Y., ruled 10 Feb 2026, opinion 17 Feb 2026) held a defendant's exchanges with a consumer AI service not privileged, partly because the provider was a third party whose terms allowed disclosure; a self-hosted tool used at counsel's direction avoids that third-party problem but does not by itself make anything privileged. Check any protective order before hosted use. The audio path gives an uncertified rough transcript only; whether AI-assisted transcripts of recorded depositions are admissible is being litigated (In re Hughey, Tex., pending as of Sep 2026).
Text description
A transcript arrives as text or PDF, or on a self-hosted box as a recording that MOSS-Transcribe-Diarize turns into an uncertified rough transcript. decosa-api parses it into page:line lines and hashes it. Qwen3.8-27B writes a topic digest where every sentence cites page:line ranges, checks each sentence against the lines it cites, and judges candidate pairs of testimony from different witnesses. Fixed patterns list non-answers and exhibits. Outputs: the checked digest, a contradiction table, a Word or Markdown export that holds back unsupported sentences, and a signed hash-chain record. Hosted calls get gateway-signed receipts; a self-hosted box signs with its own key.
At a glance
- Measured latency
- Hosted p50 is the whole digest of the fictional two-deposition pair (117 lines), over 3 runs on 2026-09-25, gateway route.
- Typical run cost
- A few cents or less for the fictional pair (a few dozen model calls) at the gateway list price. Each run shows its own measured cost.
- Data retention
- Transcripts and digests live in server memory for one hour (for export) and are never written to disk or logs.
- What leaves the box (self-host)
- Nothing: keep DECOSA_LLM_ROUTE=direct.
- Input
- Text or PDF transcripts with page:line numbers (numbered, condensed or plain); up to 4 transcripts and 4,000 lines per digest (8,000 in the self-host compose).
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
one 48 GB card
The same pipeline on a smaller mixture-of-experts model. Faster and cheaper; digest and check quality unknown.
- Models
- Gemma 4 26B A4B (instruction-tuned)
- Hardware
- 1x L40S or RTX 6000 Ada 48 GB (not measured)
- Quality evidence
- cite-check accuracy on transcriptsnot measured yet
- Latency
- not measured yet
- Verification
- Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
- In the hosted demo
Standard
the hosted demo, one 96 GB card
Qwen3.8-27B writes the digest, checks every cite and judges contradiction pairs; every call receipted.
- Models
- Qwen3.8-27B (NVIDIA NVFP4)
- MOSS-Transcribe-Diarize 0.9B
- Hardware
- 1x RTX PRO 6000 Blackwell 96 GB
- Quality evidence
- hand check: sentences fully supported by their cited lines33/40decosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route; checked by the building agent, not a lawyer
- cite checker: wrong cites / changed facts flagged59/59 / 49 of 54decosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route
- contradictions, held-out fictional set: recall / precision9/9 / 0.82-0.90 (two runs)decosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route
- cite check on real deposition transcripts (by a lawyer)not measured yet
- Latency
- measured on our server: under a minute to over a minute for a two-transcript sample, depending on length
- Verification
- Proof: strongHosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.
Best
DeepSeek-V4-Flash on two more cards
A larger writer and checker for many-witness matters.
- Models
- DeepSeek-V4-Flash (NVIDIA NVFP4)
- MOSS-Transcribe-Diarize 0.9B
- Hardware
- 2x RTX PRO 6000 96 GB
- Quality evidence
- digest and check qualitynot measured yet
- Latency
- not measured yet
- Verification
- Proof: partialSelf-host onlyNot a hosted model: calls are attested by the box's key only.
- Needs more compute
Wanted: the best setup
GLM-5.3-Flash on your own hardware
A larger model from a different family for the cross-witness contradiction search and long records. Transcripts are privileged, so it runs on your hardware, never on community providers. Not served yet.
- Models
- GLM-5.3-Flash
- MOSS-Transcribe-Diarize 0.9B
- Hardware
- Your own hardware: 2x 96 GB cards (NVFP4, about 170-186 GB, unconfirmed) or a Mac with 192 GB or more (MLX 4-bit, 165 GB). Estimate; GLM's SGLang SM120 build hung on our server.
- Quality evidence
- digest qualitynot measured yet
- Latency
- not measured yet
- Verification
- No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
| After the session | ||||
Digest writer, cite checker and contradiction judgeQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 57 GBProof: strongIn the hosted demo | Standard | 27.8B · 57 GB | Proof: strongIn the hosted demo | |
| ||||
Lite tier: digest, check and pairs on a 48 GB cardGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab) 25.2B (3.8B active)No proof yetSelf-host only | Lite | 25.2B (3.8B active) | No proof yetSelf-host only | |
| ||||
Best tier: digest writer and checker for many-witness mattersDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab) 284B (13B active) · 192 GBNo proof yetSelf-host only | Best | 284B (13B active) · 192 GB | No proof yetSelf-host only | |
| ||||
Self-host only: a recording to an uncertified rough transcript (POST /deposition/rough)MOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab) 0.9BNo proof yetSelf-host only | StandardBestWanted | 0.9B | No proof yetSelf-host only | |
| ||||
| Other | ||||
Multi-witness matters, on your own bigger box (wanted)GLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab) about 170 GB (estimate)No proof yet | Wanted | about 170 GB (estimate) | No proof yet | |
| ||||
Tools, services and hardware
Tools
- poppler pdftotext (opens in a new tab)GPL-2.0-or-later (run as a separate program)
PDF to text with -layout, one printed line per text line, so page:line cites match the PDF.
- pypdf (opens in a new tab)BSD-3-Clause
PDF fallback when pdftotext is missing; can join two printed lines, so it adds a warning.
- govinfo (U.S. GPO) (opens in a new tab)U.S. Government works, public domain (17 U.S.C. 105)
Source of the public-domain hearing transcripts in the demo (S. Hrg. 118-216; House Financial Services Serial No. 118-12).
- cryptography (Python) (opens in a new tab)Apache-2.0 OR BSD-3-Clause
Ed25519 signature on the digest's hash-chain record.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:0.1.0Parser, digest pipeline, exports and HTTP API (/deposition/*). No GPU. Binds 127.0.0.1 by default.
- decosa-llm:8000
${DECOSA_REGISTRY}/decosa-llm:0.1.0vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.
- decosa-diarize:8092
Optional, self-host only: MOSS-Transcribe-Diarize for rough transcripts from recordings. No published image yet; built from services/diarize.
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server.
- 1x L40S / RTX 6000 Ada 48 GB
Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite).
- CPU only Fits
Parsing, the non-answer and exhibit patterns, the number guard and the exports need no GPU; the digest itself needs the model.
Latency per lane
- first digest topics, two-transcript sample9.3 s
Measureddecosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route
- whole run, fictional deposition pair (117 lines, 37-38 sentences, 12 pairs)34.0 s
Measureddecosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route; 33.9 s and 76.0 s in two runs, depending on load
- whole run, two hearing excerpts (429 lines, 34-43 sentences, 12 pairs)55.0 s
Measureddecosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route; 54.6 s and 73.5 s
- 300-page depositionn/a
Measurednot measured yet
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa deposition and hearing digest on this machine
You are setting up a self-hosted deposition and hearing digest on this Linux machine: it reads transcripts with page:line numbering (text or PDF), writes a topic digest where every sentence cites page:line ranges, checks each sentence against the lines it cites, lists non-answers and exhibits, finds testimony that conflicts across witnesses, and exports to Word or Markdown. Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.
**Before anything else, remind me:** transcripts under a protective order or with privileged content must stay on this machine; keep the model route local (`direct`) and never send them to a hosted service. The digest is a work aid, not a certified transcript, and not legal advice. ABA Formal Opinion 512 (29 Jul 2024) expects lawyers to understand how a tool uses client information and to bill only for time actually spent. Repeat these points in your final summary.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/deposition.zip (4 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py deposition` (the api image carries the same bundle under /app/rehearsal/deposition/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py deposition --bundle deposition.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the transcript parses into 3 pages of numbered page:line lines", "the parser recognises the numbered deposition layout", "the digest finishes without errors"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## What you are building
| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU; includes poppler `pdftotext` for PDF transcripts, pypdf as a fallback) | none | `127.0.0.1:8445` |
| `diarize` (optional, audio only) | built from the decosa-api source (`services/diarize`) | `OpenMOSS-Team/MOSS-Transcribe-Diarize` @ `704aa4a9c304e8520be88901e0d1960158ef5b15`, Apache-2.0 | internal 8092 |
No speech model is needed for transcripts. `diarize` only turns a recording into an uncertified rough transcript.
## 1. Check the GPU, driver and Docker
1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
- Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
- Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main`; on 48 GB also `LLM_MAX_LEN=32768`. Not measured.
- Under 48 GB: stop and tell me it will not fit.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories (`sudo nvidia-ctk runtime configure --runtime=docker`, then restart Docker).
3. Confirm about 60 GB of free disk.
## 2. Get the images
The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`. If a pull fails, build from source once the `decosa-api` source is published: clone it and run `docker compose build llm api`. If neither works, stop and tell me.
## 3. Write the compose file
Create `~/decosa-deposition/.env`:
```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these digests, e.g. Firm litigation support>"
```
Create `~/decosa-deposition/docker-compose.yml` with exactly these services:
```yaml
name: decosa-deposition
x-health: &health
interval: 15s
timeout: 5s
retries: 5
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
"--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
"--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
"--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
api:
image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
restart: unless-stopped
depends_on: { llm: { condition: service_healthy } }
environment:
DECOSA_LLM_ROUTE: direct # local model only; receipts are signed by this box's key ("attested")
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LOCAL_SIGNING: "on" # Ed25519 key created at /data/attest/ed25519.pem on first start
DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_LLM_TOKENS: "400000" # per session; a 300-page deposition needs roughly 100k-350k (estimate)
DECOSA_SESSION_TTL_S: "28800"
DECOSA_DEPOSITION_MAX_LINES: "8000" # about 320 deposition pages across all transcripts in one digest
DECOSA_DEPOSITION_AUDIO: "0" # "1" only with the diarize service
DECOSA_DIARIZE_URL: "" # step 4 sets this for recordings
DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
ports: ["127.0.0.1:8445:8445"]
volumes: [decosa-data:/data]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```
Run `docker compose up -d` and poll `docker compose ps` until both are healthy (the LLM takes 5-10 minutes the first time). `curl -s localhost:8445/healthz` should show `"llm": true` (`asr` is false here and that is fine).
## 4. Optional: recordings (diarize)
Only if I want rough transcripts from recordings: run `services/diarize` from the decosa-api source on this box (install with the `uv` commands at the top of `services/diarize/requirements.txt`, then `DIARIZE_HOST=172.17.0.1 DIARIZE_PORT=8092 DIARIZE_DEVICE=cuda:0 .venv-diarize/bin/python services/diarize/server.py`; `172.17.0.1` is the docker bridge address from `ip -4 addr show docker0`, reachable from containers and not from the LAN). Add `extra_hosts: ["host.docker.internal:host-gateway"]` to `api` and set `DECOSA_DIARIZE_URL: http://host.docker.internal:8092`. Also set `DECOSA_DEPOSITION_AUDIO: "1"` and lower `LLM_GPU_UTIL` to 0.80. `POST /deposition/rough` then takes a 16 kHz mono 16-bit WAV and returns text headed "UNCERTIFIED ROUGH TRANSCRIPT" with `S01:`/`S02:` speaker labels. The fit beside the LLM is an estimate, not measured.
## 5. Smoke test
```bash
TOKEN=$(curl -s localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"deposition"}' | jq -r .token)
curl -sN localhost:8445/deposition/digest -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
-d '{"transcripts":[{"sample_id":"fictional-pruitt"},{"sample_id":"fictional-okafor"}]}' > /tmp/dep.sse
grep -oE '"(type|lane|status)": "[a-z_]+"' /tmp/dep.sse | sort | uniq -c
ID=$(grep '"type": "done"' /tmp/dep.sse | sed 's/^data: //' | jq -r .digest_id)
curl -s "localhost:8445/deposition/digests/$ID/export?format=docx" -H "authorization: Bearer $TOKEN" -o /tmp/digest.docx
grep '"lane": "record"' /tmp/dep.sse | sed 's/^data: //' | jq .data > /tmp/record.json
curl -s localhost:8445/record/verify -H 'content-type: application/json' --data-binary @/tmp/record.json | jq '{ok, summary}'
```
Pass if: lanes `transcript`, `flags`, `digest`, `verifier`, `contradictions` and `record` all appear; every `receipt` has `"status": "attested"`; the contradictions lane marks the time of the fall (2:15 pm against 11:30 am) `INCONSISTENT`; the verify prints `ok: true`; and `/tmp/digest.docx` opens in Word. On the measured card the run takes about 75 s.
Then test a PDF: `curl -s localhost:8445/deposition/parse -H "authorization: Bearer $TOKEN" -H 'content-type: application/pdf' --data-binary @some-transcript.pdf | jq '{layout, pages, lines, range, warnings}'`. `layout: numbered` means the reporter's own page:line numbers were found; `plain` means the file had none and the cites are Decosa's own pagination.
## 6. Point the app at the local API
- Base URL `http://localhost:8445` (web app: `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`). Add other origins to `DECOSA_CORS_ORIGINS`.
- `POST /deposition/digest` streams Server-Sent Events; each `lane` event replaces that lane. Exports: `GET /deposition/digests/{id}/export?format=md|docx` with the same token, for one hour. Unsupported sentences are held back from exports and listed at the end.
- Keep the API on 127.0.0.1. For other users on the LAN, put a TLS reverse proxy with authentication in front and set `DECOSA_TRUSTED_PROXIES`.
## 7. Keep the direct route
`DECOSA_LLM_ROUTE=gateway` would send prompts, which contain the transcript, to the hosted Decosa
API. Do not use it for client matters. At most, public-record transcripts.
Finish with a summary: what is running, the health output, the smoke-test results, and the privilege, not-certified and Opinion 512 reminders.Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
2 laws, rules and guidance pages cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- A page:line digest where every sentence is checked against the lines it cites, plus the places where witnesses contradict each other.
- Who it's for
- Litigation associates, paralegals and litigation-support teams.
- Where it runs
- Self-host (transcripts stay with the matter)
- Key numbers
- 9 / 9 in both runs Contradiction finder: planted conflicts found (held-out set) (test split, n = 9)
- 2 / 1 (precision 0.82 / 0.90) Contradiction finder: false-positive flags, run 1 / run 2 (test split, n = 5)
- 33 / 40 Digest sentences fully supported by their cited lines (hand-checked) (dev (tuned on), n = 40)
- 40.1 s Median end-to-end run, hosted (QA sweep 2026-09-30)
- Models
- Qwen3.8-27B
- Where
- Self-host (transcripts stay with the matter)
- Checks
- Every cite checked; receipted calls; signed record
- Industry
- Legal
- Runs
- Self-host
- Output
- Notes, reports and drafts · Signed record or verdict
- Data
- Privileged or legal · Personal data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Built from
- Grounding · Signed record
Questions people ask
Can AI summarize a deposition transcript?
The Decosa deposition and hearing digest takes one to four transcripts with page:line numbering, as text or PDF, and writes a topic digest in which every sentence cites the page:line range it relies on. It also lists non-answers and exhibits. A second model call checks each sentence against exactly those lines, and unsupported sentences are struck through and held back from the Word or Markdown export. It is a work aid, not a certified transcript; a lawyer still reads the transcript.
How accurate is the AI deposition summary?
In a hand check of the Decosa deposition digest, 33 of 40 digest sentences were fully supported by their cited lines, 7 were partly supported (mostly a dropped hedge) and 0 were not supported. On a synthetic set the cite checker flagged 59 of 59 wrong cites and 49 of 54 changed facts. One annotator, not a lawyer, did the check, no real deposition was used, and whether the digest covers everything important is not measured.
Can it find contradictions between deposition witnesses?
With two or more transcripts, the Decosa deposition digest pairs testimony from different witnesses and flags answers that cannot both be true, with both page:line cites. On a held-out fictional set it found 9 of 9 planted conflicts in both runs, with 2 and 1 false-positive flags (precision 0.82 / 0.90). The set is very small and fictional, and it was written by the same agent that wrote the prompts.
Can privileged deposition transcripts stay in-house?
Yes. The Decosa deposition digest is self-host-first: it runs on your own GPU, and with DECOSA_LLM_ROUTE=direct nothing leaves the box. Transcripts and digests live in server memory for one hour for export and are never written to disk or logs. The hosted demo takes public-domain and fictional transcripts only. Self-hosting avoids the third-party problem in United States v. Heppner but does not by itself make anything privileged; check any protective order.
What does an AI deposition summary cost to run?
Measured on 2026-09-25, a hosted Decosa deposition digest of a fictional two-deposition pair (117 lines) cost about $0.0169 at the gateway list price, using about 36k tokens and 48 model calls, with a p50 of 29,600 ms over 3 runs. A 300-page deposition has not been measured, so cost and time on a real matter are not yet known.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Deposition and hearing digest
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…