Intake checklist
waitingWhile the call runs: topics covered, what is still missing, names heard so far.
A cited intake memo, a conflicts list, a 0.1-hour time entry and follow-ups from a client call, with no notetaker bot and the audio deleted.
Built on: Live speech to text, Speaker diarization, Grounding, Signed record
Some states require every person on a call to agree before it is recorded. Say where you and the client are, and the tool applies the strictest rule. This picks a prompt from a dated table of state laws; it is not legal advice.
Choose where you and the client are to record.
Tell everyone on the call that it is recorded, and get every party's agreement where the law requires it. The hosted demo is for the synthetic sample calls or test calls with consenting people only: never a real client call. For real calls, self-host or use the Confidential tier.
Audio is streamed to the demo server, held in memory for this session only, and deleted when it ends.
Start recording or run a sample.
While the call runs: topics covered, what is still missing, names heard so far.
After hang-up: Attorney, Client and anyone else on the line, one line per turn.
The audio is dropped once the transcript is made; its hash stays in the record.
Facts, parties, dates, deadlines, issues, goals, documents, action items, open questions. Each line cites transcript lines.
Every memo line checked against the lines it cites; unsupported lines are marked and left out of the clean memo.
Every person and organisation named on the call, each looked up in the transcript.
Call length rounded up to 0.1 h, UTBMS A106, and a billing narrative.
Action items, documents to request and deadlines, with owners and dates.
To the client, in the client's language, marked privileged and confidential.
Consent, transcript lines, memo, time entry and audio deletion in one signed hash chain.
Each step is signed: which model ran, and a fingerprint of what went in and came out, so it can be checked later.
WAV, m4a or mp3, up to 5 minutes on the demo. Synthetic or consented test calls only: never a real client call on the hosted demo. The audio is held in memory and deleted once the transcript is made.
A record chains every entry of a run (each transcript line or finding, each result, each model receipt), then signs the result. This check recomputes every hash and signature itself, with no call to our servers. Change any word and it fails, naming the entry.
Run a session above, or load the sample, a recorded 59 s interview.
Recorded sessions from the live system, replayed event by event.
Loading recordings
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)# Decosa Privileged call notes: use the hosted API
You are adding Decosa's Privileged call notes to this project. A lawyer's client call (live audio, or a recording
uploaded afterwards) becomes a speaker transcript, an intake memo where every line cites transcript lines and is checked
against them, the names heard for a conflicts check, a time entry rounded up to 0.1 hour with a billing narrative,
follow-up tasks, a client email draft, and a signed session record. Decosa runs open models (Voxtral Mini 4B Realtime for
live captions, MOSS-Transcribe-Diarize for speakers, Qwen3.8-27B for text); every model output comes with a signed receipt.
Use only the endpoints below. If you need something that is not listed, stop and ask me; do not guess endpoints.
**The hosted API is for synthetic calls and development.** Real client calls are privileged: run them on a self-hosted
box or on the Confidential tier. Tell me this before writing any code.
- Base URL: `https://api.decosa.ai`
- WebSocket base: `wss://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz` returns `{"ok": true, "asr": bool, "llm": bool, ...}`.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in code.
Send `Authorization: Bearer $DECOSA_API_KEY`; WebSockets take `?token=$DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "privileged-call-notes"}` returns
`{"token", "expires_at", "budget": {"seconds_audio": 300, "llm_tokens": 20000}}`. A 3-minute call uses about 5,000-6,000 generated tokens.
3. Over a limit the API answers HTTP 429 with `Retry-After` (seconds): wait, then retry.
## Consent first (required)
`GET https://api.decosa.ai/callnotes/consent-rules?attorney_state=CA&client_state=TX` (no token) returns
`{rule: "all-party"|"one-party", prompt, why, sources, checked}`. Show `prompt` to the lawyer before recording. Every call
then carries a consent statement `{attorney_state, client_state, all_parties_consented: bool, method:
"stated_on_call"|"written"|"one_party_state"|"other", others?: [...]}`. States are two-letter codes, `DC`, `outside-us` or
`unknown`. When any place involved is all-party, mixed or unknown, `all_parties_consented` must be true, or the API refuses
(HTTP 400, or WebSocket close 4400) with the prompt. This picks a prompt from a dated table; it is not legal advice.
## Upload a recording
`POST https://api.decosa.ai/callnotes/audio?attorney_state=..&client_state=..&all_parties_consented=true&method=stated_on_call&call_date=YYYY-MM-DD&lang=en|es[&attorney=<name>]`
with the recording as the body (WAV, m4a, mp3, ogg or webm; up to 30 minutes and 48 MB) and `Authorization`. The answer is
Server-Sent Events: `receipt` events (one per model call), then `lane` events in this order: `final_transcript`,
`privacy` (the audio was dropped: `data.audio_deleted = {audio_sha256, bytes, audio_ms}`), `memo`, `verifier`,
`conflicts`, `time_entry`, `tasks`, `email`, `record`; then `result` (everything as one object), `budget`, `done`.
## A transcript you already have
`POST https://api.decosa.ai/callnotes/notes` with `{"transcript": "Attorney: ...\nClient: ...", "consent": {...}, "call_date", "lang",
"duration_s"}` (or `segments: [{start, end, speaker, text}]`). JSON by default, SSE with `"stream": true`.
## Live audio
`WS wss://api.decosa.ai/ws/live?vertical=privileged-call-notes&token=<token>&consent=<url-encoded JSON>&call_date=YYYY-MM-DD&lang=en`
- send 16 kHz mono PCM16 little-endian frames of about 100 ms, then the text frame `{"type":"stop"}`;
- you get `ready`, `transcript` (live captions), `receipt`, a live `lane` `live` (intake checklist: covered, still to
ask, names heard), then after stop the same lanes as the upload, and `done` with `summary` (memo, conflicts,
time_entry, tasks, followup_email, consent_heard, claim_check, audio_deleted, record). Expect `done` 30-90 s after the call ends.
## What the lanes hold
- `memo.data`: `{matter, client, practice_area, summary, facts, dates, deadlines, issues, client_goals, documents_to_request,
action_items, open_questions, consent_on_recording, third_party_present}`. Every item has `cites` (transcript line
numbers) and `label` (`SUPPORTED`, `PARTIAL`, `UNSUPPORTED`). Leave `UNSUPPORTED` items out of anything you file or send.
A deadline has `computed` (recomputed from `basis.from` + `basis.days`) and `check` (`recomputed`, `mismatch`, `computed`, `as_stated`, `no_date`).
- `conflicts.data.names`: `[{name, type, role, adverse, cites, aliases, heard}]`; `heard` is `exact`, `fuzzy`, `spelled`,
`partial` or `not_found` (a name nobody said: check it).
- `time_entry.data`: `{hours, minutes, date, matter, narrative, utbms: {code: "A106"}, rounding, basis}`.
- `tasks.data.tasks`: `[{owner, task, due, cites, from}]`; `email.data`: `{subject, body, lang}`.
- `record.data`: the signed record (`decosa.record.v1`); check it with `POST https://api.decosa.ai/record/verify` (no token).
## Testing without a microphone
`POST https://api.decosa.ai/demo/replay` with `{"vertical":"privileged-call-notes","script_id":"cn-family-custody-ortiz"}` (also
`cn-criminal-dui-reyes` in Spanish and `cn-pi-slip-fall-thibodeaux`) runs a synthetic call through the live pipeline.
Pass if the stream ends with lanes `memo`, `conflicts`, `time_entry`, `tasks`, `email`, `privacy` and `record`, and
`done.summary.record_check.ok` is true.
## Rules
- Never send a real client call to the hosted demo. Self-host or use the Confidential tier for real calls.
- The memo is a draft for a lawyer to check; it is not legal advice and not a certified transcript.
- Show the claim check next to the memo, and never claim more than the receipts prove.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa Privileged call notes: run it yourself (containers)
You are setting up Decosa Privileged call notes to run entirely on this machine's NVIDIA GPU(s), so client calls never
leave it. Nothing is sent to Decosa's hosted API and there are no Decosa charges. The local service speaks the same API as
the hosted one, so apps built against the hosted API only need a new base URL.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver. The speaker
service (MOSS-Transcribe-Diarize) is needed for uploaded recordings.
Ask me before any command that needs sudo, and show me the command first. Remind me first: get recording consent
(some states require every party's), and the memo is a draft for a lawyer to check, not legal advice.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/privileged-call-notes.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py privileged-call-notes` (the api image carries the same bundle under /app/rehearsal/privileged-call-notes/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py privileged-call-notes --bundle privileged-call-notes.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the memo has at least eight facts", "opposing counsel is on the conflicts list", "the 159-second call is a 0.1-hour entry"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
NVIDIA's official instructions, then run `sudo nvidia-ctk runtime configure --runtime=docker` and
`sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep `DECOSA_LLM_ROUTE=direct` so prompts stay on this box. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health: poll `curl -fsS http://localhost:<PORT>/healthz` every 10 s until `"ok": true` with `"asr": true`,
`"llm": true` and `"diarize": true`. Show me `docker compose logs --tail=50` if it is not up after 20 minutes.
7. Smoke test: `curl -fsS "http://localhost:<PORT>/callnotes/consent-rules?attorney_state=TX&client_state=CA"` should
print `"rule": "all-party"`; then take a demo token (`POST /demo/session {"vertical":"privileged-call-notes"}`) and
run `POST /demo/replay {"vertical":"privileged-call-notes","script_id":"cn-family-custody-ortiz"}`: the stream should
end with a `memo`, `conflicts`, `time_entry` and `record` lane.
8. Signing key: `curl -fsS http://localhost:<PORT>/attest/signing-key` shows the key this box generated. Tell me to back
up the data volume. Receipts from this box say `attested`: signed by our own key, an attestation, not a proof.
9. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.
## Shared network (leave it off)
Leave it off: this box handles privileged client material.
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Voxtral Mini 4B Realtime needs a GPU.
The standard tier does not fit: Needs about 40 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes.
Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.
The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions.
The standard tier does not fit: Needs about 53.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits.
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
The standard tier fits (85.6 of 96 GB).
The standard tier fits (85.6 of 192 GB).
The standard tier fits with changes: Replace Voxtral Mini 4B Realtime with Voxtral Mini 4B Realtime, MLX 4-bit. MLX build for Apple Silicon.
The standard tier does not fit: Needs about 52 GB of GPU memory at the smallest settings; 48 GB available to the GPU. The lite tier fits with changes.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
One file describes the API, the speech model and the language model as services.
mkdir -p ~/decosa && cd ~/decosa
curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yamlThe first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
Wait until the API reports ok with both models loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz
# {"ok": true, "asr": true, "llm": true, ...}
curl -fsS -X POST http://localhost:<PORT>/demo/session \
-H 'Content-Type: application/json' -d '{"vertical":"privileged-call-notes"}'expected.json. Every check must print PASS.docker compose exec api python scripts/rehearse.py privileged-call-notes
Download the mock-data bundle (3 KB, 10 checks)expected.json
A synthetic call between a family lawyer and a father who wants to change a custody schedule (fictional people; the transcript is the diarizer output of a recording voiced by Decosa house voices). The memo must cite the call, the conflicts list must carry opposing counsel, the time entry must be 0.1 hour under UTBMS A106, a run without a consent statement must be refused, and the signed record must verify and fail when changed.
Licence: Synthetic call written for Decosa: every person, firm and court is fictional. Part of decosa-api, AGPL-3.0-or-later.
# Decosa Privileged call notes: run it yourself (containers)
You are setting up Decosa Privileged call notes to run entirely on this machine's NVIDIA GPU(s), so client calls never
leave it. Nothing is sent to Decosa's hosted API and there are no Decosa charges. The local service speaks the same API as
the hosted one, so apps built against the hosted API only need a new base URL.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver. The speaker
service (MOSS-Transcribe-Diarize) is needed for uploaded recordings.
Ask me before any command that needs sudo, and show me the command first. Remind me first: get recording consent
(some states require every party's), and the memo is a draft for a lawyer to check, not legal advice.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/privileged-call-notes.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py privileged-call-notes` (the api image carries the same bundle under /app/rehearsal/privileged-call-notes/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py privileged-call-notes --bundle privileged-call-notes.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the memo has at least eight facts", "opposing counsel is on the conflicts list", "the 159-second call is a 0.1-hour entry"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
NVIDIA's official instructions, then run `sudo nvidia-ctk runtime configure --runtime=docker` and
`sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep `DECOSA_LLM_ROUTE=direct` so prompts stay on this box. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health: poll `curl -fsS http://localhost:<PORT>/healthz` every 10 s until `"ok": true` with `"asr": true`,
`"llm": true` and `"diarize": true`. Show me `docker compose logs --tail=50` if it is not up after 20 minutes.
7. Smoke test: `curl -fsS "http://localhost:<PORT>/callnotes/consent-rules?attorney_state=TX&client_state=CA"` should
print `"rule": "all-party"`; then take a demo token (`POST /demo/session {"vertical":"privileged-call-notes"}`) and
run `POST /demo/replay {"vertical":"privileged-call-notes","script_id":"cn-family-custody-ortiz"}`: the stream should
end with a `memo`, `conflicts`, `time_entry` and `record` lane.
8. Signing key: `curl -fsS http://localhost:<PORT>/attest/signing-key` shows the key this box generated. Tell me to back
up the data volume. Receipts from this box say `attested`: signed by our own key, an attestation, not a proof.
9. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.
## Shared network (leave it off)
Leave it off: this box handles privileged client material.
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Doesn't fitPrivileged call notes on GeForce RTX 5090
Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
The self-host prompt for Privileged call notes, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Privileged call notes on my hardware Fetch https://decosa.ai/prompts/privileged-call-notes-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=privileged-call-notes) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Lite · one 48 GB card, uploaded calls (lite). Fit check: doesn't fit; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - After hang-up: MOSS-Transcribe-Diarize 0.9B (OpenMOSS-Team/MOSS-Transcribe-Diarize), 4 GB - Lite tier: Qwen3.8-27B (official FP8) (Qwen/Qwen3.8-27B-FP8), 33.6 GB Warning: the fit check says this tier does not fit: Needs about 36 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further. During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/privileged-call-notes-assemble.md
Hosted: verified 28 Sep 2026 · measured 28 Sep 2026: · p50 148 s · p95 204 s (26 runs) · ~$0.017 per run · 44 receipts
Loading the nightly status…
Self-host: verified 28 Sep 2026 · fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume on the host network, direct route, local signing, against the already-running Voxtral, MOSS diarizer and Qwen3.8; torn down after
Measured cost to run: about $0.073 per call (hosted, 28 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
Assembly prompt §6: TX/CA gives all-party; 400 without consent; custody call upload 39 s, 11 names incl. opposing counsel, 0.1 h, audio dropped, record verified (135 entries), 41 receipts attested. Live replay at 2x: 155 s, 52 captions, record verified (238 entries).
For solo and small-firm lawyers (family, personal injury, estate, criminal, employment) who take client calls all day and lose the recap and the 0.1-hour entry afterwards. Record the call live or upload the recording. When it ends, an open diarization model writes who said what, and the language model writes an intake memo where every line cites the transcript: facts, parties, dates, deadlines (recomputed when the call gives a start date and a number of days), issues, goals, documents to request, action items and open questions. A claim check marks lines the call does not support and leaves them out of the clean memo. You also get every name heard for your conflicts check, a time entry rounded up to 0.1 hour with a billing narrative, follow-up tasks and a client email in the client's language. A state-aware recording-consent step comes first; the audio is dropped once the transcript is made, and the session ends in a signed record. For real client calls, self-host it or run it on the Confidential tier.
A consent step first: the lawyer says where each side is and the strictest recording rule applies. Call audio (live microphone or an uploaded recording) goes to decosa-api. Live, Voxtral Mini 4B Realtime (Apache-2.0) makes captions and Qwen3.8-27B keeps an intake checklist. When the call ends, MOSS-Transcribe-Diarize 0.9B (Apache-2.0) writes who said what with a receipt per line, then the audio is dropped and its hash recorded. Qwen3.8-27B (Apache-2.0) writes a cited intake memo, lists every name for conflicts, and checks each memo line against the lines it cites; code recomputes deadlines and rounds the call length up to 0.1 hour; the model writes the billing narrative and a client email. Consent, transcript lines, memo lines, model receipts, time entry and the audio deletion go into one signed hash chain. Self-hosted or in the Confidential tier, everything stays on that box or in the attested enclave.
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
one 48 GB card, uploaded calls
Diarized transcript and the full memo, conflicts, time entry and tasks on the FP8 checkpoint; no live captions.
the hosted demo
Live captions and checklist, then speakers, cited memo, claim check, conflicts, time entry, tasks and email, all in a signed record.
DeepSeek-V4-Flash writes and checks
Standard plus a larger memo writer and checker on two more 96 GB cards.
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
| Live, while it happens | ||||
Live captions during the call (streaming, no speakers)Voxtral Mini 4B Realtimemistralai/Voxtral-Mini-4B-Realtime-2602 on Hugging Face (opens in a new tab) 4.4B · 24 GBProof: partialIn the hosted demo | LiteStandardBest | 4.4B · 24 GB | Proof: partialIn the hosted demo | |
| ||||
| After the session | ||||
After hang-up (or on an upload): who said what, one line per turn, each with its own receiptMOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab) 0.9BProof: partialIn the hosted demo | LiteStandardBest | 0.9B | Proof: partialIn the hosted demo | |
| ||||
Speaker roles, live intake checklist, cited memo, conflicts names, claim check, time-entry narrative and client emailQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 57 GBProof: strongIn the hosted demo | StandardBest | 27.8B · 57 GB | Proof: strongIn the hosted demo | |
| ||||
Lite tier: the same text steps on the official FP8 checkpointQwen3.8-27B (official FP8)Qwen/Qwen3.8-27B-FP8 on Hugging Face (opens in a new tab) 27.8BProof: strongSelf-host only | Lite | 27.8B | Proof: strongSelf-host only | |
| ||||
Best tier: memo writer and claim checkerDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab) 284B (13B active) · 192 GBNo proof yetSelf-host only | Best | 284B (13B active) · 192 GB | No proof yetSelf-host only | |
| ||||
Picks the consent prompt: the strictest rule of the states involved. 28 of 51 entries re-read on the official page; the rest are marked not re-checked.
Recomputes a deadline said on the call from its start date and number of days, and flags a mismatch.
Activity code on the time entry: Communicate (with client).
${DECOSA_REGISTRY}/decosa-api:0.1.0Lane engine and HTTP/WS API (/ws/live, /demo/replay, /healthz). No GPU. Binds 127.0.0.1 by default.
${DECOSA_REGISTRY}/decosa-llm:0.1.0vLLM OpenAI endpoint for Qwen3.8-27B, served as qwen3.8-27b. Internal to the compose network.
${DECOSA_REGISTRY}/decosa-asr:0.1.0vLLM realtime endpoint for Voxtral Mini 4B Realtime (served as voxtral-realtime). Internal to the compose network.
MOSS-Transcribe-Diarize 0.9B pass-2 service (decosa-api services/diarize, GPU0, loopback only). decosa-api runs pass 2 on stop: diarize, role map, cited note, verifier. No published image yet; self-host builds it (see the assemble prompt).
Hosted demo layout on our server: Voxtral and MOSS-TD on GPU0, Qwen3.8-27B on GPU1 (two cards). The one-card compose split is not measured.
Not measured. FP8 language model plus the diarizer for uploaded calls; live captions need Voxtral as well.
Measuredp50 148 s, p95 204 s for calls with 2.6 min median audio, 3 calls in parallel on the busy shared gateway; an idle single call took 38-46 s; decosa-api docs/evals/privileged-call-notes.md (28 Sep 2026)
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble Decosa privileged call notes on this machine
You are setting up self-hosted call notes for a law firm on this Linux machine. A client call (recorded live in the browser, or uploaded afterwards) becomes: a speaker-attributed transcript, an intake memo where every line cites transcript lines and is checked against them, the names heard for a conflicts check, a time entry rounded up to 0.1 hour with a billing narrative, follow-up tasks, a client email draft (English or Spanish), and a session record signed with this machine's own key. The call audio is held in memory and dropped once the transcript is made. Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.
**Before anything else, remind me:**
- Client calls are privileged and confidential. Keep everything on this machine: the model route stays local (`direct`), and nothing goes to a hosted service.
- Get recording consent first. Some states (for example California, Penal Code 632) require every party's consent; the API refuses to record or process a call until a consent statement is given, and applies the strictest rule of the states involved. That is a prompt, not legal advice.
- The memo is a draft for a lawyer to check. Speech recognition mishears names and numbers; deadlines are only the ones said on the call; the time entry follows your billing rules.
Repeat these points in your final summary.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/privileged-call-notes.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py privileged-call-notes` (the api image carries the same bundle under /app/rehearsal/privileged-call-notes/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py privileged-call-notes --bundle privileged-call-notes.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the memo has at least eight facts", "opposing counsel is on the conflicts list", "the 159-second call is a 0.1-hour entry"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## What you are building
| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `asr` | `${DECOSA_REGISTRY}/decosa-asr:0.1.0` (vLLM 0.27.1 + `mistral-common[audio]`) | `mistralai/Voxtral-Mini-4B-Realtime-2602`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
| `diarize` | built from the decosa-api source (`services/diarize`) | `OpenMOSS-Team/MOSS-Transcribe-Diarize` @ `704aa4a9c304e8520be88901e0d1960158ef5b15`, Apache-2.0 | internal 8092 |
Any OpenAI-compatible endpoint serving an open model can replace `llm` (set `DECOSA_LLM_URL` and `DECOSA_LLM_MODEL`); the measured setup is the one above.
## 1. Check the GPU, driver and Docker
1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
- Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4).
- Hopper (H100/H200): set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main`, `LLM_GPU_UTIL=0.62`.
- 48 GB Ada/L40S: FP8 checkpoint as above, plus `LLM_MAX_LEN=16384`, `LLM_GPU_UTIL=0.70`, `ASR_GPU_UTIL=0.22`, `DECOSA_LIVE_CAP=2`.
- Under 48 GB: stop and tell me it will not fit.
Only the Blackwell defaults have been measured; the other rows are starting points.
2. Check `docker --version` and `docker compose version`. If Docker is missing, install Docker Engine from Docker's official apt/dnf repository for this distro.
3. Check `docker run --rm --gpus all ubuntu nvidia-smi`. If it fails, install the NVIDIA Container Toolkit (`nvidia-container-toolkit`) from NVIDIA's repository, run `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
4. Confirm about 80 GB of free disk for images and weights.
## 2. Get the images
The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,asr,api}:0.1.0`. If a pull fails (not published yet, or no access), build from source once the `decosa-api` source is published: clone it, then `docker compose build llm asr api` in the repo, which builds the same tags from `docker/`. If neither works, stop and tell me.
## 3. Write the compose file
Create `~/decosa-callnotes/.env`:
```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.60
ASR_GPU_UTIL=0.25
DECOSA_LLM_ROUTE=direct # local model; receipts are signed by this box's own key ("attested")
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Family Law PLLC>"
DECOSA_LIVE_CAP=4
```
Create `~/decosa-callnotes/docker-compose.yml` with exactly these services:
```yaml
name: decosa-callnotes
x-gpu: &gpu
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
x-health: &health
interval: 15s
timeout: 5s
retries: 5
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
<<: *gpu
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
"--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
"--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
"--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
asr:
image: ${DECOSA_REGISTRY}/decosa-asr:${DECOSA_TAG}
<<: *gpu
ipc: host
restart: unless-stopped
depends_on: { llm: { condition: service_healthy } } # start after llm so the memory split is stable
volumes: [hf-cache:/root/.cache/huggingface]
command: ["--model", "mistralai/Voxtral-Mini-4B-Realtime-2602", "--tokenizer-mode", "mistral", "--config-format", "mistral",
"--load-format", "mistral", "--compilation-config", '{"cudagraph_mode":"PIECEWISE"}', "--max-model-len", "45000",
"--max-num-batched-tokens", "8192", "--max-num-seqs", "16", "--gpu-memory-utilization", "${ASR_GPU_UTIL}",
"--served-model-name", "voxtral-realtime", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 600s }
api:
image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
restart: unless-stopped
depends_on: { llm: { condition: service_healthy }, asr: { condition: service_healthy } }
environment:
DECOSA_ASR_WS: ws://asr:8000/v1/realtime
DECOSA_LLM_ROUTE: ${DECOSA_LLM_ROUTE}
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LIVE_CAP: ${DECOSA_LIVE_CAP}
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_AUDIO_S: "3600" # per session; one call up to 30 min (DECOSA_CALLNOTES_MAX_AUDIO_S)
DECOSA_BUDGET_LLM_TOKENS: "200000"
DECOSA_SESSION_TTL_S: "28800"
DECOSA_LOCAL_SIGNING: "on" # Ed25519 key created at /data/attest/ed25519.pem on first start
DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
DECOSA_DIARIZE_URL: "" # step 4 sets this if you add speaker labels
DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
ports: ["127.0.0.1:8445:8445"]
volumes: [decosa-data:/data]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```
Run `docker compose up -d`, then poll `docker compose ps` until all three are healthy (the LLM takes 5–10 minutes the first time) and `curl -s localhost:8445/healthz` shows `"asr": true, "llm": true`. If `llm` runs out of memory, lower `LLM_GPU_UTIL` or `LLM_MAX_LEN`; the two GPU shares must add up to less than about 0.9.
## 4. Speaker labels (diarize service)
The `diarize` service (MOSS-Transcribe-Diarize 0.9B, Apache-2.0) has no published image yet. If I want speaker labels, run `services/diarize` from the decosa-api source on this box (install with the `uv` commands at the top of `services/diarize/requirements.txt`, then `DIARIZE_HOST=172.17.0.1 DIARIZE_PORT=8092 DIARIZE_DEVICE=cuda:0 .venv-diarize/bin/python services/diarize/server.py`; `172.17.0.1` is the docker bridge address from `ip -4 addr show docker0`, reachable from containers and not from the LAN). Add `extra_hosts: ["host.docker.internal:host-gateway"]` to `api` and set `DECOSA_DIARIZE_URL: http://host.docker.internal:8092`. `curl -s localhost:8445/healthz` then shows `"diarize": true`. The 0.9B model needs about 2 GB of weights; lower `LLM_GPU_UTIL` to 0.55 so it fits. That fit is an estimate, not measured. For call notes this step is **required for uploads** (`POST /callnotes/audio` answers 503 without it) and strongly recommended live: without it the memo is written from the live captions, with no Attorney/Client labels.
## 5. The signing key
1. `curl -s localhost:8445/attest/signing-key` shows the public key. Show me the `pubkey`, and tell me to back up the `decosa-data` volume (it holds `/data/attest/ed25519.pem`).
2. Receipts from this box say `status: "attested"`: signed by its own key. They show nothing was changed after signing and who signed; they do not prove the model heard correctly.
## 6. Smoke test: a synthetic call, uploaded
```bash
API=localhost:8445
curl -s "$API/callnotes/consent-rules?attorney_state=TX&client_state=CA" | jq '{rule, prompt}'
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"privileged-call-notes"}' | jq -r .token)
# the bundled synthetic call (from the decosa-api source): demo_scripts/audio/cn-family-custody-ortiz.wav
curl -sN "$API/callnotes/audio?attorney_state=OH&client_state=OH&all_parties_consented=true&method=stated_on_call&call_date=2026-09-15" \
-H "authorization: Bearer $TOKEN" -H 'content-type: audio/wav' --data-binary @cn-family-custody-ortiz.wav > /tmp/cn.sse
grep '"type": "result"' /tmp/cn.sse | sed 's/^data: //' > /tmp/cn.json
jq '{claims: .claim_check, hours: .time_entry.hours, names: [.conflicts[].name], audio_deleted: .audio_deleted.bytes, record_ok: .record_check.ok}' /tmp/cn.json
grep '"lane": "record"' /tmp/cn.sse | sed 's/^data: //' | jq '.data' > /tmp/cn-record.json
curl -s $API/record/verify -H 'content-type: application/json' --data-binary @/tmp/cn-record.json | jq '{ok, summary}'
```
Pass if: the first call prints `all-party`; the result has a `claim_check` with counts, `hours` 0.1, a names list that includes `Gregory Paskett`, a non-zero `audio_deleted`, and `record_ok: true`; and the verify prints `ok: true`. Without the consent parameters, `/callnotes/audio` answers 400.
Then the live path: `POST /demo/replay {"vertical":"privileged-call-notes","script_id":"cn-family-custody-ortiz"}` streams the same call through Voxtral in real time; the stream ends with lanes `memo`, `conflicts`, `time_entry`, `tasks`, `email`, `privacy` and `record`.
## 7. Point the app at the local API
- Base URL: `http://localhost:8445` (web app: `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`). Add other origins to `DECOSA_CORS_ORIGINS`.
- Live calls: `ws://localhost:8445/ws/live?vertical=privileged-call-notes&token=<t>&consent=<url-encoded JSON {attorney_state, client_state, all_parties_consented, method}>&call_date=YYYY-MM-DD&lang=en|es`, 16 kHz mono PCM16 frames of about 100 ms, then `{"type":"stop"}`. A missing or insufficient consent statement closes the socket with code 4400 and the prompt to show.
- Recordings: `POST /callnotes/audio` (WAV, m4a, mp3, ogg or webm; consent in the query string). Transcripts you already have: `POST /callnotes/notes` with `transcript` or `segments` and `consent`.
- `GET /callnotes/info` lists what it does and does not do; `GET /callnotes/consent-rules` the rule for a pair of states.
- The server stores nothing. Keep the memo and the signed record in your matter file; anyone can check the record with `POST /record/verify`.
- Keep the API on `127.0.0.1`; for the LAN, put a TLS reverse proxy with authentication in front and set `DECOSA_TRUSTED_PROXIES`.
## 8. Hosted gateway route (off, and leave it off)
`DECOSA_LLM_ROUTE=gateway` would send the prompts, which contain the client's words, to Decosa's hosted gateway. Never use it for client calls on a self-hosted box.
Finish with a summary: what is running, the health output, the signing key's pubkey, the smoke-test results, and the reminders above.Last reviewed
Nothing joins the call. You record it in the browser or upload the recording afterwards. An open diarization model writes who said what, and an open language model writes the intake memo, the conflicts names, the time entry and the follow-ups. Every memo line cites transcript lines and is checked against them. Self-hosted or on the Confidential tier, the audio never leaves your box or the attested enclave.
Yes, as a prompt, not legal advice. Before recording you say where you and the client are. If any state involved requires every party's consent (California and eight others), or the rule is mixed or unknown, recording stays off until you confirm everyone agreed. The table of state laws was checked on 28 Sep 2026 and links each statute.
It is held in memory only and dropped as soon as the transcript is made. Its hash, length and the time it was dropped go into the signed record, so you can match a copy you keep. Nothing is written to disk.
The call audio's length, rounded up to the next 0.1 hour by default (nearest is an option), with the ABA UTBMS activity code A106 and a one- or two-sentence billing narrative written from the checked memo. It is a draft: your firm's and the client's billing rules decide the entry.
On 20 synthetic test calls graded blind, the memo stated 89.3% of the key facts correctly, invented nothing (0 of 796 lines), and 1.5% of lines had a wrong detail. Conflicts-name recall was 89.6% (94.0% allowing for speech-recognition spellings). Every line cites the transcript so you can check it.
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…