Skip to content
decosa

Make and do

Make a settlement video

A narrated settlement video built from the chronology, the bills and the client's own words and photos, with every fact on screen cited to its exhibit page and the bills tied out in code.

104 s typical (median) on the sample; slowest 1 in 20: 110 s~$0.014 per video in model and GPU time, measured, at list price376 / 376 Cites on the right page (held-out synthetic matters)Details, API and self-host

The matter: loading…

How it works

  1. It starts from the matter's chronology (every entry already cited to a page and box), the itemized bills and the client's signed statement.
  2. An open model drafts seven scenes. It picks which entries each line rests on; the cites are attached in code, never written by the model.
  3. Every line is checked against what it cites, and every date, count and amount against the record. The bills scene is added up in code.
  4. You edit, reorder and approve each scene, and pick the narration: your own recording, your voice with your recorded consent, or a house voice.
  5. Frames are drawn from the exhibits and the script, then encoded to an MP4 with content credentials and a cite sheet PDF.

What it never does

  • No video-generation model, no reenactments, no generated faces, no “what the crash looked like”.
  • The client's photos and statement appear only with the client's recorded consent, and are shown as supplied, never edited.
  • No demand figure, valuation or future-care number: those are yours to write.
  • Nothing renders until you approve every scene.

Today, and with this

Today most cases go to mediation with a PDF demand the adjuster skims; a settlement video from a vendor costs thousands of dollars and takes weeks. Here the first cut comes from the file you already have, in minutes, with every fact on screen tied to its exhibit page.

Real client files are medical records and work product: run it on your own hardware or request confidential access. The hosted demo takes synthetic matters only. Self-host is in early access: the code and images are not public yet, so ask us for access. The eval write-up behind the figures above is not public yet either.

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B and the document reader (as for the medical chronology); the voices, the checks and the renderer run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the settlement video from the case file API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
settlement-video

Use the hosted API

# Decosa settlement video: use the hosted API

You are wiring Decosa's settlement-video maker into this project (a case-management tool, a demand-package workflow or a
script). From a matter's medical chronology (made by `POST /chronology/run`), itemized bills, the client's signed
statement and photos, it drafts a 7-scene script where every line carries its exhibit, page and box, checks each line
against what it cites, ties the bills out in code, and, once the attorney has approved every scene, renders an MP4 with
content credentials, a cite sheet PDF and a signed record. No video-generation model is used. Use only what is listed
below; if you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes synthetic matters only** (`"synthetic": true` is required). Real client files, medical records
  and photos belong on the self-hosted version or confidential access.
- It is a drafting aid for lawyers, not legal advice. It never writes a demand amount, and nothing renders until every
  scene is approved.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool's page, kept in `DECOSA_API_KEY`, never in code.
   Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "settlement-video"}` returns `{"token", "expires_at", "budget"}`.
   429 with `Retry-After` over a limit, 402 when the budget left is too small, 409 when a demo token already has a run going.

## Draft, edit, render
- `POST /settlement/draft` (token) `{"sample": "whitlock"}` (the bundled synthetic matter), or
  `{"synthetic": true, "chronology": <the /chronology/run report>, "files": [{"id": "F1", "title": "...", "role": "record|bills|statement",
  "file_b64": "..."}], "photos": [{"file_b64": "...", "when": "before|after"}], "matter": {"client_name": "...", "incident": "YYYY-MM-DD"},
  "consent": {"client": "<identity id>", "matter": "<matter id>"}}` (file ids as in the chronology; consent is needed with photos or a statement).
  The answer: `draft_id`, `lines` (each with `scene`, `text`, `refs`, `kind`, `cites` [exhibit, page, box, quote], `status`
  ok | check | unsupported | numbers | argument, and why), `tie` (every bill line, statement sums, printed totals, flags,
  `specials`), `raise` (gaps, prior conditions, date conflicts and bill problems, for the lawyer only), `usage`.
  With `Accept: text/event-stream` (or `"stream": true`): `ready`, `consent`, `bills`, `script`, a `line` per checked line,
  `report`, `budget`, `done`.
- `POST /settlement/check {"draft_id", "line": {...}}` re-checks an edited line; `POST /settlement/bills {"draft_id", "include": [...],
  "exclude": [...]}` redoes the tie-out and the bills scene in code.
- `POST /settlement/render {"draft_id", "edit": {"lines": [...], "approved": {"before": true, ...}}, "narration": {"mode": "house",
  "voice": "am_michael"}}` (modes: house, own with one recording per scene, consented, none). 409 until every scene in the
  script is approved and every line is `ok` or marked `argument`.
- `GET /settlement/jobs/{job_id}` until `status` is `done`; then `GET /settlement/jobs/{job_id}/video.mp4`, `/cite-sheet.pdf`
  and `/record.json` (the token as a header or `?token=`). Everything is deleted after two hours; `DELETE` the job sooner.
- `POST /record/verify` (no token) `{"record": {...}}` re-checks the signed record.

## Example (Python, `pip install httpx`)
```python
import httpx, os, time, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
d = httpx.post(f"{API}/settlement/draft", json={"sample": "whitlock"}, headers=H, timeout=300).json()
print(d["tie"]["specials"], [f["kind"] for f in d["tie"]["flags"]])
lines = [l for l in d["lines"] if l["status"] in ("ok", "argument")]      # the lawyer reviews and edits these first
edit = {"lines": lines, "approved": {l["scene"]: True for l in lines}}
job = httpx.post(f"{API}/settlement/render", json={"draft_id": d["draft_id"], "edit": edit, "narration": {"mode": "house"}}, headers=H).json()
while (j := httpx.get(f"{API}/settlement/jobs/{job['job_id']}", headers=H).json())["status"] not in ("done", "error"):
    time.sleep(5)
pathlib.Path("settlement-video.mp4").write_bytes(httpx.get(f"{API}{j['urls']['video']}", headers=H).content)
pathlib.Path("cite-sheet.pdf").write_bytes(httpx.get(f"{API}{j['urls']['cite_sheet']}", headers=H).content)
```

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa settlement video: run it yourself (containers)

You are setting up Decosa's settlement-video maker on this machine, so medical records, bills and client photos never
leave it. It drafts a 7-scene script from the matter's chronology where every line carries its exhibit, page and box,
checks each line against what it cites, ties the bills out in code, and renders an MP4 drawn from the exhibits (no
video-generation model) with content credentials, a cite sheet and a signed record, after the attorney approves every
scene. It is a drafting aid for lawyers: an attorney decides.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/settlement-video.zip (2 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py settlement-video` (the api image carries the same bundle under /app/rehearsal/settlement-video/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py settlement-video --bundle settlement-video.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the bills tie to $24,331.00 in code", "the statement whose printed total is off its own lines is flagged", "the duplicate therapy charge is flagged and counted once"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker and the NVIDIA container toolkit, as in Docker's and NVIDIA's official instructions; check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the `llm`
   service (Qwen3.8-27B on vLLM), the document reader (`services/docreader` with PaddleOCR-VL-1.6) and the `api`.
3. The api needs the Kokoro house voices for narration: build it with `docker/drama/Dockerfile` on top of the api image
   and set `DECOSA_STUDIO_KOKORO_PYTHON=/opt/kokoro/bin/python`. Set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`,
   `DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_DOCREADER_URL=http://docreader:8497`, `DECOSA_CHRONO_SYNTHETIC_ONLY=0` and
   `DECOSA_SETTLE_SYNTHETIC_ONLY=0`, and bind every port to 127.0.0.1. Never set the gateway route on a box with client files.
4. `docker compose up -d` and wait for the health checks. `GET /settlement/info` shows `synthetic_only: false`.
5. Smoke test: a token from `POST /demo/session {"vertical":"settlement-video"}`, then `POST /settlement/draft {"sample":"whitlock"}`.
   Expect `tie.specials` `$24,514.00` and the flags `mis_total`, `duplicate`, `pre_incident`, `no_record`. Render the lines with
   status `ok` (all scenes approved, `narration: {"mode": "house"}`), wait for the job, download the MP4 and the cite sheet,
   and send the record to `POST /record/verify`: `ok` must be true.
6. Real matters: record the client's consent with `POST /settlement/consent` (the client reads the sentence it shows), then
   draft from your chronology. Report back: the public key, the tie-out, the render time.

## provider network (optional)
Off by default, and never on a box that holds client material.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVFP4) needs a GPU.

  • GeForce RTX 4090Doesn't fit

    Needs about 25.4 GB of GPU memory at the smallest settings; 24 GB available.

  • GeForce RTX 5090Doesn't fit

    Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (63 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (63 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBCan't tell

    Memory not known for Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build.

  • Apple M5 Max, 64 GBCan't tell

    Memory not known for Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"settlement-video"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py settlement-video

Download the mock-data bundle (2 KB, 11 checks)expected.json

Dana Whitlock is fictional (the medical chronology's sample packet). The matter adds itemized bills from five providers and her signed statement, plus stand-in photos with no person in them. The run drafts the 7-scene script from the chronology (every line cited and checked), ties the bills out in code, then renders the approved lines with captions only (no narration) and signs a record. Planted in the bills: one statement's printed total is $18 off its own lines, one therapy charge is billed twice, a 2022 visit comes before the injury, and one therapy charge falls in the break in care with no record. It takes about a minute: one script call, one check per line, and about 3 minutes of frames drawn on CPU. The sample matter ships with the API (GET /settlement/samples/whitlock): the chronology report, Exhibits 1-7 and the stand-in photos, so the bundle needs no input files.

What the rehearsal checks
  • the bills tie to $24,331.00 in code
  • the statement whose printed total is off its own lines is flagged
  • the duplicate therapy charge is flagged and counted once
  • the 2022 charge before the injury is left out
  • the therapy charge with no matching record is flagged
  • most lines pass the checks and carry cites
  • the other side's points are listed for the lawyer (the 138-day gap among them)
  • the render finishes
  • the cite sheet lists the cited lines
  • the signed record verifies
  • a changed record fails

Licence: Synthetic: the client, providers, dates, bills and statement are invented (decosa_api/verticals/chronology/synth.py, decosa_api/verticals/settlement/synth.py). Stand-in photos: public domain or CC0, no person shown (decosa_api/verticals/settlement/data/photos/licences.json). Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa settlement video: run it yourself (containers)

You are setting up Decosa's settlement-video maker on this machine, so medical records, bills and client photos never
leave it. It drafts a 7-scene script from the matter's chronology where every line carries its exhibit, page and box,
checks each line against what it cites, ties the bills out in code, and renders an MP4 drawn from the exhibits (no
video-generation model) with content credentials, a cite sheet and a signed record, after the attorney approves every
scene. It is a drafting aid for lawyers: an attorney decides.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/settlement-video.zip (2 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py settlement-video` (the api image carries the same bundle under /app/rehearsal/settlement-video/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py settlement-video --bundle settlement-video.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the bills tie to $24,331.00 in code", "the statement whose printed total is off its own lines is flagged", "the duplicate therapy charge is flagged and counted once"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker and the NVIDIA container toolkit, as in Docker's and NVIDIA's official instructions; check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the `llm`
   service (Qwen3.8-27B on vLLM), the document reader (`services/docreader` with PaddleOCR-VL-1.6) and the `api`.
3. The api needs the Kokoro house voices for narration: build it with `docker/drama/Dockerfile` on top of the api image
   and set `DECOSA_STUDIO_KOKORO_PYTHON=/opt/kokoro/bin/python`. Set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`,
   `DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_DOCREADER_URL=http://docreader:8497`, `DECOSA_CHRONO_SYNTHETIC_ONLY=0` and
   `DECOSA_SETTLE_SYNTHETIC_ONLY=0`, and bind every port to 127.0.0.1. Never set the gateway route on a box with client files.
4. `docker compose up -d` and wait for the health checks. `GET /settlement/info` shows `synthetic_only: false`.
5. Smoke test: a token from `POST /demo/session {"vertical":"settlement-video"}`, then `POST /settlement/draft {"sample":"whitlock"}`.
   Expect `tie.specials` `$24,514.00` and the flags `mis_total`, `duplicate`, `pre_incident`, `no_record`. Render the lines with
   status `ok` (all scenes approved, `narration: {"mode": "house"}`), wait for the job, download the MP4 and the cite sheet,
   and send the record to `POST /record/verify`: `ok` must be true.
6. Real matters: record the client's consent with `POST /settlement/consent` (the client reads the sentence it shows), then
   draft from your chronology. Report back: the public key, the tie-out, the render time.

## provider network (optional)
Off by default, and never on a box that holds client material.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

Doesn't fitSettlement video from the case file on GeForce RTX 5090

Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available.

Lite · captions or your own recording, no voice models: what changesuses estimates

  • Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available.
Memory per component
  • The script checks: decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocks. CPU. Runs on CPU (vram_gb 0 in stack.json).
  • The 7-scene script: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
  • Finds the regions of each bill and statement...: Docling 2.130 with the Heron layout model (document reader block). ~1 GB (from stack.json). vram_gb 1 in stack.json.
  • Reads the bill tables as cells and the statem...: PaddleOCR-VL-1.6 (0.9B, document reader block). ~4.4 GB (from stack.json). vram_gb 4.4 in stack.json.

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Settlement video from the case file, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Settlement video from the case file on my hardware

Fetch https://decosa.ai/prompts/settlement-video-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=settlement-video)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Lite · captions or your own recording, no voice models (lite). Fit check: doesn't fit; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- The script checks: decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocks, CPU
- The 7-scene script: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB
- Finds the regions of each bill and statement...: Docling 2.130 with the Heron layout model (document reader block) (docling-project/docling-layout-heron), 1 GB
- Reads the bill tables as cells and the statem...: PaddleOCR-VL-1.6 (0.9B, document reader block) (PaddlePaddle/PaddleOCR-VL-1.6), 4.4 GB

Warning: the fit check says this tier does not fit: Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further.

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/settlement-video-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 30 Sep 2026 · measured 30 Sep 2026: · p50 104 s · p95 110 s (5 runs) · ~$0.010 per run · 23 receipts

Loading the nightly status…

Self-host: verified 29 Sep 2026 · fresh clone of the branch into a clean directory, run with the host's Python environment (no container build: the server's root disk was full at the time), direct route to the local Qwen3.8-27B, the running document reader, local signing, real-matters mode on

Measured cost to run: about $0.014 per video (hosted, 30 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The rehearsal bundle passed 11/11 in 38.5 s (draft, tie-out flags, render, record verified and failed when changed) and the smoke module passed in 19.9 s.

Known limits (6)
  • Hosted numbers are the whole sample task measured on production (draft + one line re-checked + render with a house voice, through the production API), run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
  • Measured on synthetic matters only; real records, bills and photos are not measured, and the lawyer's review time (10-20 minutes estimated) has not been timed with a lawyer.
  • When the chronology misses a visit, a real charge on that day is flagged as having no matching visit and left out of the default total until the lawyer puts it back (4 of 11 eval matters, $186-594).
  • A printed statement total on a degraded scan or fax is sometimes unreadable; the tool says so and the video uses the sum of the lines.
  • A child's photo is refused until a guardian consent path for this purpose is approved.
  • No causation, future-care, wage-loss or billed-versus-paid figures: only what the chronology and the itemized bills hold.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B and the document reader (as for the medical chronology); the voices, the checks and the renderer run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the settlement video from the case file API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.
The open stack

A narrated settlement video from the chronology, the bills and the client's own words, every fact cited to its exhibit page and the bills tied out in code.

For plaintiff's personal-injury lawyers taking a case to mediation. It starts from the matter's medical chronology, itemized bills, the client's signed statement and photos, drafts a 7-scene script where every line carries its exhibit, page and box, checks each line against what it cites, adds up the bills in code, and renders an MP4 drawn from the exhibits after the lawyer edits and approves every scene. No video-generation model is used.

Deployment
Self-host first
Regulatory
Written 29 Sep 2026. Settlement communications are generally inadmissible to prove or disprove a disputed claim's validity or amount under Federal Rule of Evidence 408 (https://www.law.cornell.edu/rules/fre/rule_408, read 29 Sep 2026); states have their own versions, and whether and how to label a video is the lawyer's call. Medical records are protected health information in a covered entity's hands (45 CFR Parts 160 and 164, https://www.hhs.gov/hipaa/for-professionals/privacy/index.html); a firm holds them under its own duties of confidentiality, so real matters run self-hosted or on the confidential API, never the hosted demo. The client's photos and statement are shown only with the client's recorded consent (consent ledger), and any generated narration voice is disclosed in the video and its content credential. It is a drafting aid, not legal advice, and it never writes a demand amount.
Architecture
Text description

The matter's chronology, bills, the client's statement and photos go to decosa-api. The document reader reads the bills and statement. Qwen3.8-27B drafts a seven-scene script naming the entries each line rests on; code attaches the cites. Every line is checked by the grounding judge and the numeric block, and the bills are tied out in code. The lawyer edits and approves every scene and picks the narration; the consent ledger gates the client's photos and statement and any voice. A renderer draws the frames from the exhibits with Pillow and ffmpeg, with no video model. Out come an MP4 with content credentials, a cite sheet and a signed record.

Architecture

At a glance

Data retention
Drafts, page images, photos and videos are kept for two hours on the server that made them, then deleted (Delete now removes a render at once). The signed record holds hashes and receipt ids, never record text, names or photos; logs carry counts, timings and ids only. The hosted demo takes synthetic matters only.
What leaves the box
Self-hosted on the direct route: nothing. The model, the document reader, the voices and the renderer run on the same machine. Hosted: record text and page images go through the Decosa API, and only synthetic matters are accepted.
What every line on screen carries
The exhibit, page and box it rests on, attached in code, and a check against the cited text; bill amounts are sums in code. Lines that fail are held back until the lawyer fixes or removes them; only the closing scenes may carry the lawyer's own words, labelled as such.
What it will not do
Generate a person, a face, a reenactment or 'what the crash looked like'; edit or generate from the client's photos; show a child's photo (refused for now); write a demand amount; render before every scene is approved.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    captions or your own recording, no voice models

    The same script, checks, tie-out, renderer, cite sheet and record, narrated by the attorney's own recording or captions only; no Kokoro or Chatterbox runtime to install.

    Models
    • decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocks
    • Qwen3.8-27B (NVFP4)
    • Docling 2.130 with the Heron layout model (document reader block)
    • PaddleOCR-VL-1.6 (0.9B, document reader block)
    Hardware
    1x GPU with about 26 GB free for Qwen3.8-27B and the reader, plus 8+ CPU cores for the frames
    Quality evidence
    • Script, checks and tie-outsame as standard (the voice does not change them)decosa-api docs/evals/settlement-video.md
    • Render time, captions onlynot measured yetdecosa-api docs/evals/settlement-video.md, render table
    Latency
    measured: frames and encode in under a minute for a short video on many CPU cores; narration with a house voice adds a little more
    Verification
    Proof: strongSelf-host onlyEvery model call is receipted; the record and the C2PA credential are signed by the instance.
  • In the hosted demo

    Standard

    reader, model, house voices and the consented clone (hosted demo)

    What the hosted demo runs on synthetic matters: the document reader for the bills and statement, Qwen3.8-27B for the script and the per-line checks, Kokoro-82M house voices and Chatterbox for the attorney's consented voice on CPU, and the renderer.

    Models
    • decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocks
    • Qwen3.8-27B (NVFP4)
    • Docling 2.130 with the Heron layout model (document reader block)
    • PaddleOCR-VL-1.6 (0.9B, document reader block)
    • Kokoro-82M (stock voicepacks am_michael, af_heart)
    • Chatterbox Multilingual
    Hardware
    1x RTX PRO 6000 96 GB (measured on shared cards) and 32 CPU cores
    Quality evidence
    • Cites on the right page, lines that render (6 held-out synthetic matters)376 / 376decosa-api docs/evals/settlement-video.md, held-out seeds, 29 Sep 2026
    • Cites on the right page, 2 matters written blind by another agent143 / 144 on review (139 / 144 by the writer's key)decosa-api docs/evals/settlement-video.md, blind split
    • Dates in rendered lines that match the answer key (all splits)100 / 100decosa-api docs/evals/settlement-video.md: 53 held out, 17 blind, 30 post-fix
    • Every itemized charge read and summed to the cent11 / 11 mattersdecosa-api docs/evals/settlement-video.md: the default total also leaves out charges with no matching visit, which equals the key in 7 / 11 (the chronology missed 1-3 visits in the others)
    • Planted bill problems flagged (before the injury, duplicate, no matching visit, statement total off)11/11, 11/11, 11/11, 10/11decosa-api docs/evals/settlement-video.md, all splits
    • Rendered lines a blind judge found supported by their cited text77 / 80 (3 partly, 0 not supported)decosa-api docs/evals/settlement-video.md, Claude Code Opus 5.5 blind, held-out sample
    Latency
    measured on the pre-release server, gateway route: the draft in under a minute, the render in about a minute for a short video with a house voice
    Verification
    Proof: strongEvery model call goes through our gateway with its own signed receipt; the record carries the receipt ids.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
The script checks (refs, cites attached in code, numbers, spinal levels and doses), the bills tie-out, the scene plan, the frames and the cite sheet (no model; CPU)decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocks
0 GBProof: partial
The 7-scene script (one call: which chronology entries and statement paragraphs each line rests on) and one grounding verdict per line against the cited record text; re-reads of bill and statement regions the parser was unsure ofQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Finds the regions of each bill and statement page (tables, text) with their boxesDocling 2.130 with the Heron layout model (document reader block)docling-project/docling-layout-heron on Hugging Face (opens in a new tab)
1 GBProof: partialIn the hosted demo
Reads the bill tables as cells and the statement's paragraphsPaddleOCR-VL-1.6 (0.9B, document reader block)PaddlePaddle/PaddleOCR-VL-1.6 on Hugging Face (opens in a new tab)
0.9B · 4.4 GBProof: partialIn the hosted demo
The stock house voice that reads the approved script, when the lawyer picks itKokoro-82M (stock voicepacks am_michael, af_heart)hexgrad/Kokoro-82M on Hugging Face (opens in a new tab)
82M · 0 GBProof: partialIn the hosted demo
The attorney's own voice, generated from the attorney's consent recording, only with an active consent-ledger entry for this matterChatterbox MultilingualResembleAI/chatterbox on Hugging Face (opens in a new tab)
0 GBProof: partialIn the hosted demo

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag> plus the Kokoro layer (docker/drama/Dockerfile)

    GET /settlement/info, /settlement/samples; POST /settlement/draft (SSE or JSON), /settlement/check, /settlement/bills, /settlement/render; GET /settlement/jobs/{id} and its files; POST /settlement/consent (self-host). Deletes drafts and videos after two hours.

  • decosa-docreader (+ PaddleOCR-VL on vLLM):8497
    built from services/docreader; parser on vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Bill and statement pages in, regions with boxes out.

  • vLLM (Qwen3.8-27B):8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    The script and the per-line grounding verdicts.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the model through the shared gateway on one card, the document reader on another shared card; the voices, checks and frames on CPU (32 cores).

  • 1x 48 GB card Fits

    Estimate: Qwen3.8-27B NVFP4 (about 20 GB with a small KV cache) plus about 6 GB for the reader. Not measured.

  • CPU only Does not fit

    The model and the page parser need a GPU. The renderer, voices and checks run on CPU: about a minute of drawing and encoding for a 3-minute video on our server.

Latency per lane

  • draft (script and per-line checks), held-out matters, gateway23.6 s

    Measuredmeasured on our server 2026-09-29 (pre-release server): 16.6-31.0 s over 6 held-out synthetic matters, 15-18 model calls each

  • render with a house voice, the Whitlock sample64.7 s

    Measuredmeasured on our server 2026-09-29: 61.3 s and 68.1 s for 3:20 and 3:17 videos (narration about 17 s on CPU, frames and encoding about 42 s on 32 cores)

Notes

  • On six held-out synthetic matters every cite of every line that would render pointed at a page where the answer key has that fact (376 of 376), every stated date matched (53 of 53), and every itemized charge was read and summed to the cent.
  • On two matters written blind by another agent, 143 of 144 cites were on the right page on review; a blind judge found 77 of 80 sampled lines supported by their cited text and none unsupported.
  • Planted bill problems were flagged 43 of 44 times; the flag for charges with no matching visit also fires when the chronology missed a visit (7 false flags), and those charges are left out of the default total until the lawyer puts them back.
  • A blind plaintiff's-attorney persona would use it on mid-value cases after changes (most now made); a blind adjuster persona said it does not move the number but makes the specials easy to trust and verify.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

settlement-video/assemble-prompt.md202 lines
# Assemble the Decosa settlement-video maker on this machine

You are setting up a self-hosted tool for a plaintiff's personal-injury firm on this Linux machine. From a matter's
medical records, itemized bills, the client's signed statement and the client's own photos, it makes:
- a medical chronology where every line cites the file, page and box (the "Build a medical chronology" tool);
- a 7-scene settlement video script where every line names the chronology entries or statement paragraphs it rests on,
  with the cites attached in code, each line checked against what it cites, every date, count and amount checked against
  the record, and the bills added up in code (a statement's wrong total, duplicate charges, charges before the injury and
  charges with no matching visit flagged);
- after the attorney edits and approves every scene: an MP4 drawn from the exhibits (no video-generation model), narrated
  by the attorney's own recording, a stock house voice, or the attorney's voice with a consent-ledger entry, with content
  credentials, a cite sheet PDF and a signed, hash-chained record.

Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- Medical records, bills and client photos are protected health information and work product. On this box nothing leaves
  the machine: the model, the document reader, the voices and the renderer run here. Never point it at a hosted gateway
  while it holds client material.
- The client's photos and statement are used only with the client's recorded consent (POST /settlement/consent), and the
  attorney approves every scene. It is a drafting aid, not legal advice; it never writes a demand figure.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/settlement-video.zip (2 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py settlement-video` (the api image carries the same bundle under /app/rehearsal/settlement-video/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py settlement-video --bundle settlement-video.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the bills tie to $24,331.00 in code", "the statement whose printed total is off its own lines is flagged", "the duplicate therapy charge is flagged and counted once"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `parser` | `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1` | `PaddlePaddle/PaddleOCR-VL-1.6` @ `c5630abae1d940eafe0697512a0325494b02ab42`, Apache-2.0 | internal 8498 |
| `docreader` | built from `decosa-api/services/docreader` | Docling layout `docling-project/docling-layout-heron` @ `8f39ad3c0b4c58e9c2d2c84a38465abf757272d8` (MIT + Apache-2.0) | internal 8497 |
| `api` | built: `decosa-api` + a Kokoro layer (`docker/drama/Dockerfile`), no GPU | Kokoro-82M `hexgrad/Kokoro-82M` @ `f3ff3571791e39611d31c381e3a41a3af07b4987`, Apache-2.0, CPU | `127.0.0.1:8445` |

The attorney's consented voice uses Chatterbox Multilingual (`ResembleAI/chatterbox` @ `5bb1f6ee58e50c3b8d408bc82a6d3740c2db6e18`,
MIT) in its own venv; it is optional (step 6). Without it, narrate with the attorney's own recording or a house voice.

## 1. Check the GPU, driver and Docker

1. `nvidia-smi`: one NVIDIA GPU with at least 48 GB (the model about 20 GB plus KV cache, the parser and layout model
   about 6 GB). Driver 580 or newer. Blackwell (RTX PRO 6000) runs the NVFP4 defaults below; on Hopper or a 48 GB card use
   `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main`, `LLM_GPU_UTIL=0.62`, `LLM_MAX_LEN=32768` (not measured).
2. `docker --version`, `docker compose version`, `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the NVIDIA
   Container Toolkit is missing, install them from the official repositories and run `sudo nvidia-ctk runtime configure --runtime=docker`.
3. About 80 GB of free disk; at least 8 CPU cores (frames are drawn on CPU, about a minute for a 3-minute video on 8 cores).

## 2. Get the source and build the api image

```bash
mkdir -p ~/decosa-settle && cd ~/decosa-settle
git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>     # publishing soon; if it fails, stop and tell me
cd decosa-api
docker build -f docker/api/Dockerfile -t decosa-api:local .
docker build -f docker/drama/Dockerfile --build-arg BASE=decosa-api:local -t decosa-api:settle .
cd ..
```
The second image adds Kokoro-82M and its stock voicepacks on CPU (`/opt/kokoro`), so the house voices run offline.

## 3. Write the compose file

Create `~/decosa-settle/.env`:
```bash
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.62
PARSER_GPU_UTIL=0.05
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Injury Law records team>"
```

Create `~/decosa-settle/docker-compose.yml`:
```yaml
name: decosa-settle
x-gpu: &gpu { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
x-health: &health { interval: 15s, timeout: 5s, retries: 5 }
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:0.1.0
    deploy: *gpu
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}", "--max-num-seqs", "16",
              "--kv-cache-dtype", "fp8_e4m3", "--seed", "0", "--enable-force-include-usage", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  parser:
    image: vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1
    deploy: *gpu
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["PaddlePaddle/PaddleOCR-VL-1.6", "--revision", "c5630abae1d940eafe0697512a0325494b02ab42", "--trust-remote-code",
              "--max-model-len", "8192", "--gpu-memory-utilization", "${PARSER_GPU_UTIL}", "--max-num-seqs", "16",
              "--no-enable-prefix-caching", "--mm-processor-cache-gb", "0", "--generation-config", "vllm", "--host", "0.0.0.0", "--port", "8498"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8498/health', timeout=4)"], start_period: 600s }
  docreader:
    build: { context: ./decosa-api/services/docreader }
    deploy: *gpu
    restart: unless-stopped
    depends_on: { parser: { condition: service_healthy } }
    environment: { DOCREADER_PARSER_URL: "http://parser:8498/v1", DOCREADER_PARSER_MODEL: PaddlePaddle/PaddleOCR-VL-1.6, DOCLING_DEVICE: cuda }
    volumes: [hf-cache:/models]
  api:
    image: decosa-api:settle
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy }, docreader: { condition: service_started } }
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_DOCREADER_URL: http://docreader:8497
      DECOSA_CHRONO_RETRIEVAL: bm25             # the chronology's cite index without the retrieval service
      DECOSA_CHRONO_SYNTHETIC_ONLY: "0"         # this box accepts real records
      DECOSA_SETTLE_SYNTHETIC_ONLY: "0"         # and real matters, with recorded consent
      DECOSA_CHRONO_MAX_PAGES: "600"
      DECOSA_STUDIO_KOKORO_PYTHON: /opt/kokoro/bin/python
      DECOSA_SETTLE_RENDER_WORKERS: "6"
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "600000"
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_LLM_TIMEOUT_S: "300"
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```
Use the named volume `decosa-data` exactly as written (a root-owned bind mount breaks `/data`). Run `docker compose up -d --build`
and poll `docker compose ps` until healthy (the LLM takes 5 to 10 minutes the first time).
`curl -s localhost:8445/settlement/info | jq '{synthetic_only, narration}'` should show `synthetic_only: false`.

## 4. Smoke test on the bundled synthetic matter (fictional client, stand-in photos)

```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"settlement-video"}' | jq -r .token)
curl -s $API/settlement/draft -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample":"whitlock"}' > /tmp/draft.json
jq '{counts, specials: .tie.specials, flags: [.tie.flags[].kind]}' /tmp/draft.json
jq -r '.lines[] | "\(.id)\t\(.status)\t\(.scene)\t\(.text[0:70])"' /tmp/draft.json
jq '{draft_id, edit: {lines: [.lines[] | select(.status=="ok" or .status=="argument")], approved: {before:true, incident:true, timeline:true, injuries:true, treatment:true, bills:true, after:true}}, narration: {mode: "house", voice: "am_michael"}}' /tmp/draft.json \
  | curl -s $API/settlement/render -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @- | tee /tmp/job.json
JOB=$(jq -r .job_id /tmp/job.json); until curl -s $API/settlement/jobs/$JOB -H "authorization: Bearer $TOKEN" | jq -e '.status=="done" or .status=="error"' >/dev/null; do sleep 5; done
curl -s $API/settlement/jobs/$JOB -H "authorization: Bearer $TOKEN" | jq '{status, error, outputs: .outputs | {seconds, cited_lines, credential, timings_s}}'
curl -s "$API/settlement/jobs/$JOB/video.mp4?token=$TOKEN" -o settlement-video.mp4
curl -s "$API/settlement/jobs/$JOB/cite-sheet.pdf?token=$TOKEN" -o cite-sheet.pdf
curl -s "$API/settlement/jobs/$JOB/record.json?token=$TOKEN" | jq '{record: .}' | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```
Pass if: the bills tie to `$24,514.00` with the flags `mis_total`, `duplicate`, `pre_incident` and `no_record`; most lines are
`ok`, and a line naming the MRI level "L4-15" (the scanned page's read slip) is held back as `numbers`; the job is `done`
with a video of about 3 minutes and `credential: "c2pa"`; and the record verifies.

## 5. A real matter

1. Build the chronology first: `POST /chronology/run` with the record files (see the medical chronology's own prompt).
2. Record the client's consent: show the client the sentence from `POST /settlement/consent` (it is in the 422 answer when the
   read-back fails), record it on a phone, then
   `POST /settlement/consent {"role":"client","name":"<client>","matter":"<matter id>","recording_b64":"<audio>"}`.
3. `POST /settlement/draft` with `{chronology, files: [{id, title, role: record|bills|statement, file_b64}], photos: [{file_b64, when}],
   matter: {client_name, incident, id, caption, prepared_by}, consent: {client: <identity_id>, matter: <matter>}}`. File ids must
   match the chronology's (F1, F2, ... in the order you sent them). Mark a photo of a child `minor: true`: it is refused for now.
4. Edit and approve in the web app, or send the final lines to `POST /settlement/render` as in step 4.

## 6. Optional: the attorney's consented voice

Create a venv with Chatterbox (`python -m venv /opt/chatterbox && /opt/chatterbox/bin/pip install chatterbox-tts==0.1.4`)
inside the api container or a derived image, set `DECOSA_DUB_TTS_PYTHON=/opt/chatterbox/bin/python`, enrol the attorney
with `POST /settlement/consent {"role":"attorney", ...}` and render with `{"mode":"consented","identity_id":"<id>","reference_b64":"<the same recording>"}`.
It runs on CPU here (`DECOSA_SETTLE_TTS_DEVICE=cpu`): several minutes for a 3-minute script. Generated voices are disclosed
on the end card and in the content credential.

## 7. Point the web app at it, and keep the network off

- Base URL `http://localhost:8445`; for the web app set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`.
- Drafts and renders are deleted after two hours (`DECOSA_SETTLE_TTL_S`); keep the MP4, the cite sheet and the record with the file.
- `DECOSA_LLM_ROUTE=gateway` would send record text to a the gateway: leave it off on a box with client material.

Finish with a summary: what is running, the health output, the smoke-test results and the reminders above.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A settlement video tool for personal-injury lawyers: it turns the matter's chronology, bills, the client's statement and photos into a narrated video where every fact on screen is cited to its exhibit page, and nothing renders until the lawyer approves each scene.
Who it's for
Plaintiff's personal-injury lawyers and paralegals preparing a case for mediation.
Where it runs
Self-host or the confidential API for real client files; the hosted demo takes synthetic matters only
Key numbers

On six held-out synthetic matters, 376 of 376 cites of rendered lines pointed at the right page and the bills tied out to the cent on all six; real records are not measured yet.

  • 143 / 144 Cites on the right page, matters written blind by another agent (held out, n = 144)
  • 376 / 376 Cites on the right page, rendered lines (test split, n = 376)
  • 53 / 53 Dates in rendered lines right (test split, n = 53)
  • 103.5 s Median end-to-end run, hosted (QA sweep 2026-09-30)
All results, datasets and caveats
Models
Qwen3.8-27B (script and per-line grounding) · document reader (bills and statement) · Kokoro-82M house voices or Chatterbox (the attorney's consented voice), CPU · Pillow + ffmpeg renderer (no generative video model)
Where
Self-host or the confidential API for real client files; the hosted demo takes synthetic matters only
Checks
Receipt per model call; every line's cites attached in code and checked; bills tied out in code; consent-ledger decisions for the voice, the photos and the statement; C2PA on the MP4; signed hash-chained record
Industry
Legal
Output
Media · Signed record or verdict
Data
Patient data (PHI) · Privileged or legal
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

Does the settlement video tool generate footage of the client or the accident?

No. There is no video-generation model. The frames are drawn from the matter's own exhibit pages, the client's own photos (shown as supplied) and graphics made from the record: a timeline, record pages zoomed to the cited box and a tally of the bills.

How are the medical bills added up?

In code, from the bill tables the document reader reads: each charge keeps its exhibit, page and row. A statement whose printed total is off its own lines, duplicate charges, charges before the injury and charges with no matching visit are flagged for the lawyer.

Can the video use my own voice?

Yes: record the narration yourself, one file per scene, or record a consent sentence once and use a voice generated from it for that matter only (disclosed in the video). A stock house voice that is no person's is the third option.

Can I use real client files on the hosted demo?

No. The hosted demo takes synthetic matters only and deletes drafts and videos after two hours. Real matters run on your own hardware or with confidential access.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Settlement video from the case file

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.