Make and do
Make a settlement video
A narrated settlement video built from the chronology, the bills and the client's own words and photos, with every fact on screen cited to its exhibit page and the bills tied out in code.
104 s typical (median) on the sample; slowest 1 in 20: 110 s~$0.014 per video in model and GPU time, measured, at list price376 / 376 Cites on the right page (held-out synthetic matters)Details, API and self-host
The matter: loading…
How it works
- It starts from the matter's chronology (every entry already cited to a page and box), the itemized bills and the client's signed statement.
- An open model drafts seven scenes. It picks which entries each line rests on; the cites are attached in code, never written by the model.
- Every line is checked against what it cites, and every date, count and amount against the record. The bills scene is added up in code.
- You edit, reorder and approve each scene, and pick the narration: your own recording, your voice with your recorded consent, or a house voice.
- Frames are drawn from the exhibits and the script, then encoded to an MP4 with content credentials and a cite sheet PDF.
What it never does
- No video-generation model, no reenactments, no generated faces, no “what the crash looked like”.
- The client's photos and statement appear only with the client's recorded consent, and are shown as supplied, never edited.
- No demand figure, valuation or future-care number: those are yours to write.
- Nothing renders until you approve every scene.
Today, and with this
Today most cases go to mediation with a PDF demand the adjuster skims; a settlement video from a vendor costs thousands of dollars and takes weeks. Here the first cut comes from the file you already have, in minutes, with every fact on screen tied to its exhibit page.
Real client files are medical records and work product: run it on your own hardware or request confidential access. The hosted demo takes synthetic matters only. Self-host is in early access: the code and images are not public yet, so ask us for access. The eval write-up behind the figures above is not public yet either.
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B and the document reader (as for the medical chronology); the voices, the checks and the renderer run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the settlement video from the case file API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real patient data belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- settlement-video
Use the hosted API
# Decosa settlement video: use the hosted API
You are wiring Decosa's settlement-video maker into this project (a case-management tool, a demand-package workflow or a
script). From a matter's medical chronology (made by `POST /chronology/run`), itemized bills, the client's signed
statement and photos, it drafts a 7-scene script where every line carries its exhibit, page and box, checks each line
against what it cites, ties the bills out in code, and, once the attorney has approved every scene, renders an MP4 with
content credentials, a cite sheet PDF and a signed record. No video-generation model is used. Use only what is listed
below; if you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes synthetic matters only** (`"synthetic": true` is required). Real client files, medical records
and photos belong on the self-hosted version or confidential access.
- It is a drafting aid for lawyers, not legal advice. It never writes a demand amount, and nothing renders until every
scene is approved.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool's page, kept in `DECOSA_API_KEY`, never in code.
Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "settlement-video"}` returns `{"token", "expires_at", "budget"}`.
429 with `Retry-After` over a limit, 402 when the budget left is too small, 409 when a demo token already has a run going.
## Draft, edit, render
- `POST /settlement/draft` (token) `{"sample": "whitlock"}` (the bundled synthetic matter), or
`{"synthetic": true, "chronology": <the /chronology/run report>, "files": [{"id": "F1", "title": "...", "role": "record|bills|statement",
"file_b64": "..."}], "photos": [{"file_b64": "...", "when": "before|after"}], "matter": {"client_name": "...", "incident": "YYYY-MM-DD"},
"consent": {"client": "<identity id>", "matter": "<matter id>"}}` (file ids as in the chronology; consent is needed with photos or a statement).
The answer: `draft_id`, `lines` (each with `scene`, `text`, `refs`, `kind`, `cites` [exhibit, page, box, quote], `status`
ok | check | unsupported | numbers | argument, and why), `tie` (every bill line, statement sums, printed totals, flags,
`specials`), `raise` (gaps, prior conditions, date conflicts and bill problems, for the lawyer only), `usage`.
With `Accept: text/event-stream` (or `"stream": true`): `ready`, `consent`, `bills`, `script`, a `line` per checked line,
`report`, `budget`, `done`.
- `POST /settlement/check {"draft_id", "line": {...}}` re-checks an edited line; `POST /settlement/bills {"draft_id", "include": [...],
"exclude": [...]}` redoes the tie-out and the bills scene in code.
- `POST /settlement/render {"draft_id", "edit": {"lines": [...], "approved": {"before": true, ...}}, "narration": {"mode": "house",
"voice": "am_michael"}}` (modes: house, own with one recording per scene, consented, none). 409 until every scene in the
script is approved and every line is `ok` or marked `argument`.
- `GET /settlement/jobs/{job_id}` until `status` is `done`; then `GET /settlement/jobs/{job_id}/video.mp4`, `/cite-sheet.pdf`
and `/record.json` (the token as a header or `?token=`). Everything is deleted after two hours; `DELETE` the job sooner.
- `POST /record/verify` (no token) `{"record": {...}}` re-checks the signed record.
## Example (Python, `pip install httpx`)
```python
import httpx, os, time, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
d = httpx.post(f"{API}/settlement/draft", json={"sample": "whitlock"}, headers=H, timeout=300).json()
print(d["tie"]["specials"], [f["kind"] for f in d["tie"]["flags"]])
lines = [l for l in d["lines"] if l["status"] in ("ok", "argument")] # the lawyer reviews and edits these first
edit = {"lines": lines, "approved": {l["scene"]: True for l in lines}}
job = httpx.post(f"{API}/settlement/render", json={"draft_id": d["draft_id"], "edit": edit, "narration": {"mode": "house"}}, headers=H).json()
while (j := httpx.get(f"{API}/settlement/jobs/{job['job_id']}", headers=H).json())["status"] not in ("done", "error"):
time.sleep(5)
pathlib.Path("settlement-video.mp4").write_bytes(httpx.get(f"{API}{j['urls']['video']}", headers=H).content)
pathlib.Path("cite-sheet.pdf").write_bytes(httpx.get(f"{API}{j['urls']['cite_sheet']}", headers=H).content)
```
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa settlement video: run it yourself (containers)
You are setting up Decosa's settlement-video maker on this machine, so medical records, bills and client photos never
leave it. It drafts a 7-scene script from the matter's chronology where every line carries its exhibit, page and box,
checks each line against what it cites, ties the bills out in code, and renders an MP4 drawn from the exhibits (no
video-generation model) with content credentials, a cite sheet and a signed record, after the attorney approves every
scene. It is a drafting aid for lawyers: an attorney decides.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/settlement-video.zip (2 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py settlement-video` (the api image carries the same bundle under /app/rehearsal/settlement-video/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py settlement-video --bundle settlement-video.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the bills tie to $24,331.00 in code", "the statement whose printed total is off its own lines is flagged", "the duplicate therapy charge is flagged and counted once"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker and the NVIDIA container toolkit, as in Docker's and NVIDIA's official instructions; check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the `llm`
service (Qwen3.8-27B on vLLM), the document reader (`services/docreader` with PaddleOCR-VL-1.6) and the `api`.
3. The api needs the Kokoro house voices for narration: build it with `docker/drama/Dockerfile` on top of the api image
and set `DECOSA_STUDIO_KOKORO_PYTHON=/opt/kokoro/bin/python`. Set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`,
`DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_DOCREADER_URL=http://docreader:8497`, `DECOSA_CHRONO_SYNTHETIC_ONLY=0` and
`DECOSA_SETTLE_SYNTHETIC_ONLY=0`, and bind every port to 127.0.0.1. Never set the gateway route on a box with client files.
4. `docker compose up -d` and wait for the health checks. `GET /settlement/info` shows `synthetic_only: false`.
5. Smoke test: a token from `POST /demo/session {"vertical":"settlement-video"}`, then `POST /settlement/draft {"sample":"whitlock"}`.
Expect `tie.specials` `$24,514.00` and the flags `mis_total`, `duplicate`, `pre_incident`, `no_record`. Render the lines with
status `ok` (all scenes approved, `narration: {"mode": "house"}`), wait for the job, download the MP4 and the cite sheet,
and send the record to `POST /record/verify`: `ok` must be true.
6. Real matters: record the client's consent with `POST /settlement/consent` (the client reads the sentence it shows), then
draft from your chronology. Report back: the public key, the tie-out, the render time.
## provider network (optional)
Off by default, and never on a box that holds client material.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVFP4) needs a GPU.
- GeForce RTX 4090Doesn't fit
Needs about 25.4 GB of GPU memory at the smallest settings; 24 GB available.
- GeForce RTX 5090Doesn't fit
Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (63 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (63 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBCan't tell
Memory not known for Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build.
- Apple M5 Max, 64 GBCan't tell
Memory not known for Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"settlement-video"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py settlement-video
Download the mock-data bundle (2 KB, 11 checks)expected.json
Dana Whitlock is fictional (the medical chronology's sample packet). The matter adds itemized bills from five providers and her signed statement, plus stand-in photos with no person in them. The run drafts the 7-scene script from the chronology (every line cited and checked), ties the bills out in code, then renders the approved lines with captions only (no narration) and signs a record. Planted in the bills: one statement's printed total is $18 off its own lines, one therapy charge is billed twice, a 2022 visit comes before the injury, and one therapy charge falls in the break in care with no record. It takes about a minute: one script call, one check per line, and about 3 minutes of frames drawn on CPU. The sample matter ships with the API (GET /settlement/samples/whitlock): the chronology report, Exhibits 1-7 and the stand-in photos, so the bundle needs no input files.
What the rehearsal checks
- the bills tie to $24,331.00 in code
- the statement whose printed total is off its own lines is flagged
- the duplicate therapy charge is flagged and counted once
- the 2022 charge before the injury is left out
- the therapy charge with no matching record is flagged
- most lines pass the checks and carry cites
- the other side's points are listed for the lawyer (the 138-day gap among them)
- the render finishes
- the cite sheet lists the cited lines
- the signed record verifies
- a changed record fails
Licence: Synthetic: the client, providers, dates, bills and statement are invented (decosa_api/verticals/chronology/synth.py, decosa_api/verticals/settlement/synth.py). Stand-in photos: public domain or CC0, no person shown (decosa_api/verticals/settlement/data/photos/licences.json). Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa settlement video: run it yourself (containers)
You are setting up Decosa's settlement-video maker on this machine, so medical records, bills and client photos never
leave it. It drafts a 7-scene script from the matter's chronology where every line carries its exhibit, page and box,
checks each line against what it cites, ties the bills out in code, and renders an MP4 drawn from the exhibits (no
video-generation model) with content credentials, a cite sheet and a signed record, after the attorney approves every
scene. It is a drafting aid for lawyers: an attorney decides.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/settlement-video.zip (2 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py settlement-video` (the api image carries the same bundle under /app/rehearsal/settlement-video/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py settlement-video --bundle settlement-video.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the bills tie to $24,331.00 in code", "the statement whose printed total is off its own lines is flagged", "the duplicate therapy charge is flagged and counted once"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker and the NVIDIA container toolkit, as in Docker's and NVIDIA's official instructions; check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the `llm`
service (Qwen3.8-27B on vLLM), the document reader (`services/docreader` with PaddleOCR-VL-1.6) and the `api`.
3. The api needs the Kokoro house voices for narration: build it with `docker/drama/Dockerfile` on top of the api image
and set `DECOSA_STUDIO_KOKORO_PYTHON=/opt/kokoro/bin/python`. Set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`,
`DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_DOCREADER_URL=http://docreader:8497`, `DECOSA_CHRONO_SYNTHETIC_ONLY=0` and
`DECOSA_SETTLE_SYNTHETIC_ONLY=0`, and bind every port to 127.0.0.1. Never set the gateway route on a box with client files.
4. `docker compose up -d` and wait for the health checks. `GET /settlement/info` shows `synthetic_only: false`.
5. Smoke test: a token from `POST /demo/session {"vertical":"settlement-video"}`, then `POST /settlement/draft {"sample":"whitlock"}`.
Expect `tie.specials` `$24,514.00` and the flags `mis_total`, `duplicate`, `pre_incident`, `no_record`. Render the lines with
status `ok` (all scenes approved, `narration: {"mode": "house"}`), wait for the job, download the MP4 and the cite sheet,
and send the record to `POST /record/verify`: `ok` must be true.
6. Real matters: record the client's consent with `POST /settlement/consent` (the client reads the sentence it shows), then
draft from your chronology. Report back: the public key, the tie-out, the render time.
## provider network (optional)
Off by default, and never on a box that holds client material.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Doesn't fitSettlement video from the case file on GeForce RTX 5090
Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available.
Lite · captions or your own recording, no voice models: what changesuses estimates
- Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available.
Memory per component
- The script checks: decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocks. CPU. Runs on CPU (vram_gb 0 in stack.json).
- The 7-scene script: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
- Finds the regions of each bill and statement...: Docling 2.130 with the Heron layout model (document reader block). ~1 GB (from stack.json). vram_gb 1 in stack.json.
- Reads the bill tables as cells and the statem...: PaddleOCR-VL-1.6 (0.9B, document reader block). ~4.4 GB (from stack.json). vram_gb 4.4 in stack.json.
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Settlement video from the case file, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Settlement video from the case file on my hardware Fetch https://decosa.ai/prompts/settlement-video-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=settlement-video) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Lite · captions or your own recording, no voice models (lite). Fit check: doesn't fit; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - The script checks: decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocks, CPU - The 7-scene script: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB - Finds the regions of each bill and statement...: Docling 2.130 with the Heron layout model (document reader block) (docling-project/docling-layout-heron), 1 GB - Reads the bill tables as cells and the statem...: PaddleOCR-VL-1.6 (0.9B, document reader block) (PaddlePaddle/PaddleOCR-VL-1.6), 4.4 GB Warning: the fit check says this tier does not fit: Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further. During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/settlement-video-assemble.md
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 30 Sep 2026 · measured 30 Sep 2026: · p50 104 s · p95 110 s (5 runs) · ~$0.010 per run · 23 receipts
Loading the nightly status…
Self-host: verified 29 Sep 2026 · fresh clone of the branch into a clean directory, run with the host's Python environment (no container build: the server's root disk was full at the time), direct route to the local Qwen3.8-27B, the running document reader, local signing, real-matters mode on
Measured cost to run: about $0.014 per video (hosted, 30 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The rehearsal bundle passed 11/11 in 38.5 s (draft, tie-out flags, render, record verified and failed when changed) and the smoke module passed in 19.9 s.
Known limits (6)
- Hosted numbers are the whole sample task measured on production (draft + one line re-checked + render with a house voice, through the production API), run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
- Measured on synthetic matters only; real records, bills and photos are not measured, and the lawyer's review time (10-20 minutes estimated) has not been timed with a lawyer.
- When the chronology misses a visit, a real charge on that day is flagged as having no matching visit and left out of the default total until the lawyer puts it back (4 of 11 eval matters, $186-594).
- A printed statement total on a degraded scan or fax is sometimes unreadable; the tool says so and the video uses the sum of the lines.
- A child's photo is refused until a guardian consent path for this purpose is approved.
- No causation, future-care, wage-loss or billed-versus-paid figures: only what the chronology and the itemized bills hold.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B and the document reader (as for the medical chronology); the voices, the checks and the renderer run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the settlement video from the case file API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real patient data belongs on your own hardware.
A narrated settlement video from the chronology, the bills and the client's own words, every fact cited to its exhibit page and the bills tied out in code.
For plaintiff's personal-injury lawyers taking a case to mediation. It starts from the matter's medical chronology, itemized bills, the client's signed statement and photos, drafts a 7-scene script where every line carries its exhibit, page and box, checks each line against what it cites, adds up the bills in code, and renders an MP4 drawn from the exhibits after the lawyer edits and approves every scene. No video-generation model is used.
- Deployment
- Self-host first
- Regulatory
- Written 29 Sep 2026. Settlement communications are generally inadmissible to prove or disprove a disputed claim's validity or amount under Federal Rule of Evidence 408 (https://www.law.cornell.edu/rules/fre/rule_408, read 29 Sep 2026); states have their own versions, and whether and how to label a video is the lawyer's call. Medical records are protected health information in a covered entity's hands (45 CFR Parts 160 and 164, https://www.hhs.gov/hipaa/for-professionals/privacy/index.html); a firm holds them under its own duties of confidentiality, so real matters run self-hosted or on the confidential API, never the hosted demo. The client's photos and statement are shown only with the client's recorded consent (consent ledger), and any generated narration voice is disclosed in the video and its content credential. It is a drafting aid, not legal advice, and it never writes a demand amount.
Text description
The matter's chronology, bills, the client's statement and photos go to decosa-api. The document reader reads the bills and statement. Qwen3.8-27B drafts a seven-scene script naming the entries each line rests on; code attaches the cites. Every line is checked by the grounding judge and the numeric block, and the bills are tied out in code. The lawyer edits and approves every scene and picks the narration; the consent ledger gates the client's photos and statement and any voice. A renderer draws the frames from the exhibits with Pillow and ffmpeg, with no video model. Out come an MP4 with content credentials, a cite sheet and a signed record.
At a glance
- Data retention
- Drafts, page images, photos and videos are kept for two hours on the server that made them, then deleted (Delete now removes a render at once). The signed record holds hashes and receipt ids, never record text, names or photos; logs carry counts, timings and ids only. The hosted demo takes synthetic matters only.
- What leaves the box
- Self-hosted on the direct route: nothing. The model, the document reader, the voices and the renderer run on the same machine. Hosted: record text and page images go through the Decosa API, and only synthetic matters are accepted.
- What every line on screen carries
- The exhibit, page and box it rests on, attached in code, and a check against the cited text; bill amounts are sums in code. Lines that fail are held back until the lawyer fixes or removes them; only the closing scenes may carry the lawyer's own words, labelled as such.
- What it will not do
- Generate a person, a face, a reenactment or 'what the crash looked like'; edit or generate from the client's photos; show a child's photo (refused for now); write a demand amount; render before every scene is approved.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
captions or your own recording, no voice models
The same script, checks, tie-out, renderer, cite sheet and record, narrated by the attorney's own recording or captions only; no Kokoro or Chatterbox runtime to install.
- Models
- decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocks
- Qwen3.8-27B (NVFP4)
- Docling 2.130 with the Heron layout model (document reader block)
- PaddleOCR-VL-1.6 (0.9B, document reader block)
- Hardware
- 1x GPU with about 26 GB free for Qwen3.8-27B and the reader, plus 8+ CPU cores for the frames
- Quality evidence
- Script, checks and tie-outsame as standard (the voice does not change them)decosa-api docs/evals/settlement-video.md
- Render time, captions onlynot measured yetdecosa-api docs/evals/settlement-video.md, render table
- Latency
- measured: frames and encode in under a minute for a short video on many CPU cores; narration with a house voice adds a little more
- Verification
- Proof: strongSelf-host onlyEvery model call is receipted; the record and the C2PA credential are signed by the instance.
- In the hosted demo
Standard
reader, model, house voices and the consented clone (hosted demo)
What the hosted demo runs on synthetic matters: the document reader for the bills and statement, Qwen3.8-27B for the script and the per-line checks, Kokoro-82M house voices and Chatterbox for the attorney's consented voice on CPU, and the renderer.
- Models
- decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocks
- Qwen3.8-27B (NVFP4)
- Docling 2.130 with the Heron layout model (document reader block)
- PaddleOCR-VL-1.6 (0.9B, document reader block)
- Kokoro-82M (stock voicepacks am_michael, af_heart)
- Chatterbox Multilingual
- Hardware
- 1x RTX PRO 6000 96 GB (measured on shared cards) and 32 CPU cores
- Quality evidence
- Cites on the right page, lines that render (6 held-out synthetic matters)376 / 376decosa-api docs/evals/settlement-video.md, held-out seeds, 29 Sep 2026
- Cites on the right page, 2 matters written blind by another agent143 / 144 on review (139 / 144 by the writer's key)decosa-api docs/evals/settlement-video.md, blind split
- Dates in rendered lines that match the answer key (all splits)100 / 100decosa-api docs/evals/settlement-video.md: 53 held out, 17 blind, 30 post-fix
- Every itemized charge read and summed to the cent11 / 11 mattersdecosa-api docs/evals/settlement-video.md: the default total also leaves out charges with no matching visit, which equals the key in 7 / 11 (the chronology missed 1-3 visits in the others)
- Planted bill problems flagged (before the injury, duplicate, no matching visit, statement total off)11/11, 11/11, 11/11, 10/11decosa-api docs/evals/settlement-video.md, all splits
- Rendered lines a blind judge found supported by their cited text77 / 80 (3 partly, 0 not supported)decosa-api docs/evals/settlement-video.md, Claude Code Opus 5.5 blind, held-out sample
- Latency
- measured on the pre-release server, gateway route: the draft in under a minute, the render in about a minute for a short video with a house voice
- Verification
- Proof: strongEvery model call goes through our gateway with its own signed receipt; the record carries the receipt ids.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
The script checks (refs, cites attached in code, numbers, spinal levels and doses), the bills tie-out, the scene plan, the frames and the cite sheet (no model; CPU)decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocks 0 GBProof: partial | LiteStandard | 0 GB | Proof: partial | |
| ||||
The 7-scene script (one call: which chronology entries and statement paragraphs each line rests on) and one grounding verdict per line against the cited record text; re-reads of bill and statement regions the parser was unsure ofQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | LiteStandard | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Finds the regions of each bill and statement page (tables, text) with their boxesDocling 2.130 with the Heron layout model (document reader block)docling-project/docling-layout-heron on Hugging Face (opens in a new tab) 1 GBProof: partialIn the hosted demo | LiteStandard | 1 GB | Proof: partialIn the hosted demo | |
| ||||
Reads the bill tables as cells and the statement's paragraphsPaddleOCR-VL-1.6 (0.9B, document reader block)PaddlePaddle/PaddleOCR-VL-1.6 on Hugging Face (opens in a new tab) 0.9B · 4.4 GBProof: partialIn the hosted demo | LiteStandard | 0.9B · 4.4 GB | Proof: partialIn the hosted demo | |
| ||||
The stock house voice that reads the approved script, when the lawyer picks itKokoro-82M (stock voicepacks am_michael, af_heart)hexgrad/Kokoro-82M on Hugging Face (opens in a new tab) 82M · 0 GBProof: partialIn the hosted demo | Standard | 82M · 0 GB | Proof: partialIn the hosted demo | |
| ||||
The attorney's own voice, generated from the attorney's consent recording, only with an active consent-ledger entry for this matterChatterbox MultilingualResembleAI/chatterbox on Hugging Face (opens in a new tab) 0 GBProof: partialIn the hosted demo | Standard | 0 GB | Proof: partialIn the hosted demo | |
| ||||
Tools, services and hardware
Tools
- Federal Rule of Evidence 408 (opens in a new tab)US government work
Settlement communications, cited in the regulatory note (read 29 Sep 2026).
- Stand-in photos (Wikimedia Commons) (opens in a new tab)Public domain and CC0 (per file)
The sample matter's photos: a trail, a garden, a guitar and a pill organizer, with no person in any of them (licences in data/photos/licences.json).
- Lato (opens in a new tab)SIL Open Font License 1.1
The video's typeface (the fonts-lato package in the api image; DejaVu otherwise).
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag> plus the Kokoro layer (docker/drama/Dockerfile)GET /settlement/info, /settlement/samples; POST /settlement/draft (SSE or JSON), /settlement/check, /settlement/bills, /settlement/render; GET /settlement/jobs/{id} and its files; POST /settlement/consent (self-host). Deletes drafts and videos after two hours.
- decosa-docreader (+ PaddleOCR-VL on vLLM):8497
built from services/docreader; parser on vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Bill and statement pages in, regions with boxes out.
- vLLM (Qwen3.8-27B):8000
${DECOSA_REGISTRY}/decosa-llm:0.1.0The script and the per-line grounding verdicts.
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server: the model through the shared gateway on one card, the document reader on another shared card; the voices, checks and frames on CPU (32 cores).
- 1x 48 GB card Fits
Estimate: Qwen3.8-27B NVFP4 (about 20 GB with a small KV cache) plus about 6 GB for the reader. Not measured.
- CPU only Does not fit
The model and the page parser need a GPU. The renderer, voices and checks run on CPU: about a minute of drawing and encoding for a 3-minute video on our server.
Latency per lane
- draft (script and per-line checks), held-out matters, gateway23.6 s
Measuredmeasured on our server 2026-09-29 (pre-release server): 16.6-31.0 s over 6 held-out synthetic matters, 15-18 model calls each
- render with a house voice, the Whitlock sample64.7 s
Measuredmeasured on our server 2026-09-29: 61.3 s and 68.1 s for 3:20 and 3:17 videos (narration about 17 s on CPU, frames and encoding about 42 s on 32 cores)
Notes
- On six held-out synthetic matters every cite of every line that would render pointed at a page where the answer key has that fact (376 of 376), every stated date matched (53 of 53), and every itemized charge was read and summed to the cent.
- On two matters written blind by another agent, 143 of 144 cites were on the right page on review; a blind judge found 77 of 80 sampled lines supported by their cited text and none unsupported.
- Planted bill problems were flagged 43 of 44 times; the flag for charges with no matching visit also fires when the chronology missed a visit (7 false flags), and those charges are left out of the default total until the lawyer puts them back.
- A blind plaintiff's-attorney persona would use it on mid-value cases after changes (most now made); a blind adjuster persona said it does not move the number but makes the specials easy to trust and verify.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa settlement-video maker on this machine
You are setting up a self-hosted tool for a plaintiff's personal-injury firm on this Linux machine. From a matter's
medical records, itemized bills, the client's signed statement and the client's own photos, it makes:
- a medical chronology where every line cites the file, page and box (the "Build a medical chronology" tool);
- a 7-scene settlement video script where every line names the chronology entries or statement paragraphs it rests on,
with the cites attached in code, each line checked against what it cites, every date, count and amount checked against
the record, and the bills added up in code (a statement's wrong total, duplicate charges, charges before the injury and
charges with no matching visit flagged);
- after the attorney edits and approves every scene: an MP4 drawn from the exhibits (no video-generation model), narrated
by the attorney's own recording, a stock house voice, or the attorney's voice with a consent-ledger entry, with content
credentials, a cite sheet PDF and a signed, hash-chained record.
Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.
**Before anything else, remind me:**
- Medical records, bills and client photos are protected health information and work product. On this box nothing leaves
the machine: the model, the document reader, the voices and the renderer run here. Never point it at a hosted gateway
while it holds client material.
- The client's photos and statement are used only with the client's recorded consent (POST /settlement/consent), and the
attorney approves every scene. It is a drafting aid, not legal advice; it never writes a demand figure.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/settlement-video.zip (2 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py settlement-video` (the api image carries the same bundle under /app/rehearsal/settlement-video/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py settlement-video --bundle settlement-video.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the bills tie to $24,331.00 in code", "the statement whose printed total is off its own lines is flagged", "the duplicate therapy charge is flagged and counted once"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## What you are building
| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `parser` | `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1` | `PaddlePaddle/PaddleOCR-VL-1.6` @ `c5630abae1d940eafe0697512a0325494b02ab42`, Apache-2.0 | internal 8498 |
| `docreader` | built from `decosa-api/services/docreader` | Docling layout `docling-project/docling-layout-heron` @ `8f39ad3c0b4c58e9c2d2c84a38465abf757272d8` (MIT + Apache-2.0) | internal 8497 |
| `api` | built: `decosa-api` + a Kokoro layer (`docker/drama/Dockerfile`), no GPU | Kokoro-82M `hexgrad/Kokoro-82M` @ `f3ff3571791e39611d31c381e3a41a3af07b4987`, Apache-2.0, CPU | `127.0.0.1:8445` |
The attorney's consented voice uses Chatterbox Multilingual (`ResembleAI/chatterbox` @ `5bb1f6ee58e50c3b8d408bc82a6d3740c2db6e18`,
MIT) in its own venv; it is optional (step 6). Without it, narrate with the attorney's own recording or a house voice.
## 1. Check the GPU, driver and Docker
1. `nvidia-smi`: one NVIDIA GPU with at least 48 GB (the model about 20 GB plus KV cache, the parser and layout model
about 6 GB). Driver 580 or newer. Blackwell (RTX PRO 6000) runs the NVFP4 defaults below; on Hopper or a 48 GB card use
`LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main`, `LLM_GPU_UTIL=0.62`, `LLM_MAX_LEN=32768` (not measured).
2. `docker --version`, `docker compose version`, `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the NVIDIA
Container Toolkit is missing, install them from the official repositories and run `sudo nvidia-ctk runtime configure --runtime=docker`.
3. About 80 GB of free disk; at least 8 CPU cores (frames are drawn on CPU, about a minute for a 3-minute video on 8 cores).
## 2. Get the source and build the api image
```bash
mkdir -p ~/decosa-settle && cd ~/decosa-settle
git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> # publishing soon; if it fails, stop and tell me
cd decosa-api
docker build -f docker/api/Dockerfile -t decosa-api:local .
docker build -f docker/drama/Dockerfile --build-arg BASE=decosa-api:local -t decosa-api:settle .
cd ..
```
The second image adds Kokoro-82M and its stock voicepacks on CPU (`/opt/kokoro`), so the house voices run offline.
## 3. Write the compose file
Create `~/decosa-settle/.env`:
```bash
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.62
PARSER_GPU_UTIL=0.05
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Injury Law records team>"
```
Create `~/decosa-settle/docker-compose.yml`:
```yaml
name: decosa-settle
x-gpu: &gpu { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
x-health: &health { interval: 15s, timeout: 5s, retries: 5 }
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:0.1.0
deploy: *gpu
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
"--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}", "--max-num-seqs", "16",
"--kv-cache-dtype", "fp8_e4m3", "--seed", "0", "--enable-force-include-usage", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
parser:
image: vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1
deploy: *gpu
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["PaddlePaddle/PaddleOCR-VL-1.6", "--revision", "c5630abae1d940eafe0697512a0325494b02ab42", "--trust-remote-code",
"--max-model-len", "8192", "--gpu-memory-utilization", "${PARSER_GPU_UTIL}", "--max-num-seqs", "16",
"--no-enable-prefix-caching", "--mm-processor-cache-gb", "0", "--generation-config", "vllm", "--host", "0.0.0.0", "--port", "8498"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8498/health', timeout=4)"], start_period: 600s }
docreader:
build: { context: ./decosa-api/services/docreader }
deploy: *gpu
restart: unless-stopped
depends_on: { parser: { condition: service_healthy } }
environment: { DOCREADER_PARSER_URL: "http://parser:8498/v1", DOCREADER_PARSER_MODEL: PaddlePaddle/PaddleOCR-VL-1.6, DOCLING_DEVICE: cuda }
volumes: [hf-cache:/models]
api:
image: decosa-api:settle
restart: unless-stopped
depends_on: { llm: { condition: service_healthy }, docreader: { condition: service_started } }
environment:
DECOSA_LLM_ROUTE: direct # local model only; receipts are signed by this box's key ("attested")
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LOCAL_SIGNING: "on"
DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
DECOSA_DOCREADER_URL: http://docreader:8497
DECOSA_CHRONO_RETRIEVAL: bm25 # the chronology's cite index without the retrieval service
DECOSA_CHRONO_SYNTHETIC_ONLY: "0" # this box accepts real records
DECOSA_SETTLE_SYNTHETIC_ONLY: "0" # and real matters, with recorded consent
DECOSA_CHRONO_MAX_PAGES: "600"
DECOSA_STUDIO_KOKORO_PYTHON: /opt/kokoro/bin/python
DECOSA_SETTLE_RENDER_WORKERS: "6"
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_LLM_TOKENS: "600000"
DECOSA_SESSION_TTL_S: "28800"
DECOSA_LLM_TIMEOUT_S: "300"
DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
ports: ["127.0.0.1:8445:8445"]
volumes: [decosa-data:/data]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```
Use the named volume `decosa-data` exactly as written (a root-owned bind mount breaks `/data`). Run `docker compose up -d --build`
and poll `docker compose ps` until healthy (the LLM takes 5 to 10 minutes the first time).
`curl -s localhost:8445/settlement/info | jq '{synthetic_only, narration}'` should show `synthetic_only: false`.
## 4. Smoke test on the bundled synthetic matter (fictional client, stand-in photos)
```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"settlement-video"}' | jq -r .token)
curl -s $API/settlement/draft -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample":"whitlock"}' > /tmp/draft.json
jq '{counts, specials: .tie.specials, flags: [.tie.flags[].kind]}' /tmp/draft.json
jq -r '.lines[] | "\(.id)\t\(.status)\t\(.scene)\t\(.text[0:70])"' /tmp/draft.json
jq '{draft_id, edit: {lines: [.lines[] | select(.status=="ok" or .status=="argument")], approved: {before:true, incident:true, timeline:true, injuries:true, treatment:true, bills:true, after:true}}, narration: {mode: "house", voice: "am_michael"}}' /tmp/draft.json \
| curl -s $API/settlement/render -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @- | tee /tmp/job.json
JOB=$(jq -r .job_id /tmp/job.json); until curl -s $API/settlement/jobs/$JOB -H "authorization: Bearer $TOKEN" | jq -e '.status=="done" or .status=="error"' >/dev/null; do sleep 5; done
curl -s $API/settlement/jobs/$JOB -H "authorization: Bearer $TOKEN" | jq '{status, error, outputs: .outputs | {seconds, cited_lines, credential, timings_s}}'
curl -s "$API/settlement/jobs/$JOB/video.mp4?token=$TOKEN" -o settlement-video.mp4
curl -s "$API/settlement/jobs/$JOB/cite-sheet.pdf?token=$TOKEN" -o cite-sheet.pdf
curl -s "$API/settlement/jobs/$JOB/record.json?token=$TOKEN" | jq '{record: .}' | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```
Pass if: the bills tie to `$24,514.00` with the flags `mis_total`, `duplicate`, `pre_incident` and `no_record`; most lines are
`ok`, and a line naming the MRI level "L4-15" (the scanned page's read slip) is held back as `numbers`; the job is `done`
with a video of about 3 minutes and `credential: "c2pa"`; and the record verifies.
## 5. A real matter
1. Build the chronology first: `POST /chronology/run` with the record files (see the medical chronology's own prompt).
2. Record the client's consent: show the client the sentence from `POST /settlement/consent` (it is in the 422 answer when the
read-back fails), record it on a phone, then
`POST /settlement/consent {"role":"client","name":"<client>","matter":"<matter id>","recording_b64":"<audio>"}`.
3. `POST /settlement/draft` with `{chronology, files: [{id, title, role: record|bills|statement, file_b64}], photos: [{file_b64, when}],
matter: {client_name, incident, id, caption, prepared_by}, consent: {client: <identity_id>, matter: <matter>}}`. File ids must
match the chronology's (F1, F2, ... in the order you sent them). Mark a photo of a child `minor: true`: it is refused for now.
4. Edit and approve in the web app, or send the final lines to `POST /settlement/render` as in step 4.
## 6. Optional: the attorney's consented voice
Create a venv with Chatterbox (`python -m venv /opt/chatterbox && /opt/chatterbox/bin/pip install chatterbox-tts==0.1.4`)
inside the api container or a derived image, set `DECOSA_DUB_TTS_PYTHON=/opt/chatterbox/bin/python`, enrol the attorney
with `POST /settlement/consent {"role":"attorney", ...}` and render with `{"mode":"consented","identity_id":"<id>","reference_b64":"<the same recording>"}`.
It runs on CPU here (`DECOSA_SETTLE_TTS_DEVICE=cpu`): several minutes for a 3-minute script. Generated voices are disclosed
on the end card and in the content credential.
## 7. Point the web app at it, and keep the network off
- Base URL `http://localhost:8445`; for the web app set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`.
- Drafts and renders are deleted after two hours (`DECOSA_SETTLE_TTL_S`); keep the MP4, the cite sheet and the record with the file.
- `DECOSA_LLM_ROUTE=gateway` would send record text to a the gateway: leave it off on a box with client material.
Finish with a summary: what is running, the health output, the smoke-test results and the reminders above.Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- A settlement video tool for personal-injury lawyers: it turns the matter's chronology, bills, the client's statement and photos into a narrated video where every fact on screen is cited to its exhibit page, and nothing renders until the lawyer approves each scene.
- Who it's for
- Plaintiff's personal-injury lawyers and paralegals preparing a case for mediation.
- Where it runs
- Self-host or the confidential API for real client files; the hosted demo takes synthetic matters only
- Key numbers
On six held-out synthetic matters, 376 of 376 cites of rendered lines pointed at the right page and the bills tied out to the cent on all six; real records are not measured yet.
- 143 / 144 Cites on the right page, matters written blind by another agent (held out, n = 144)
- 376 / 376 Cites on the right page, rendered lines (test split, n = 376)
- 53 / 53 Dates in rendered lines right (test split, n = 53)
- 103.5 s Median end-to-end run, hosted (QA sweep 2026-09-30)
- Models
- Qwen3.8-27B (script and per-line grounding) · document reader (bills and statement) · Kokoro-82M house voices or Chatterbox (the attorney's consented voice), CPU · Pillow + ffmpeg renderer (no generative video model)
- Where
- Self-host or the confidential API for real client files; the hosted demo takes synthetic matters only
- Checks
- Receipt per model call; every line's cites attached in code and checked; bills tied out in code; consent-ledger decisions for the voice, the photos and the statement; C2PA on the MP4; signed hash-chained record
- Industry
- Legal
- Input
- Files and media
- Runs
- Self-host
- Output
- Media · Signed record or verdict
- Data
- Patient data (PHI) · Privileged or legal
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Self-host
- Built from
- Document reader · Grounding · Numeric grounding · Consent gate · Content credentials · Signed record
Questions people ask
Does the settlement video tool generate footage of the client or the accident?
No. There is no video-generation model. The frames are drawn from the matter's own exhibit pages, the client's own photos (shown as supplied) and graphics made from the record: a timeline, record pages zoomed to the cited box and a tally of the bills.
How are the medical bills added up?
In code, from the bill tables the document reader reads: each charge keeps its exhibit, page and row. A statement whose printed total is off its own lines, duplicate charges, charges before the injury and charges with no matching visit are flagged for the lawyer.
Can the video use my own voice?
Yes: record the narration yourself, one file per scene, or record a consent sentence once and use a voice generated from it for that matter only (disclosed in the video). A stock house voice that is no person's is the third option.
Can I use real client files on the hosted demo?
No. The hosted demo takes synthetic matters only and deletes drafts and videos after two hours. Real matters run on your own hardware or with confidential access.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Settlement video from the case file
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…