Skip to content
decosa
LiveHostedSelf-host

Turn a client call into notes

A cited intake memo, a conflicts list, a 0.1-hour time entry and follow-ups from a client call, with no notetaker bot and the audio deleted.

Held-out test89.3%Facts from the call kept in the memo (blind grader, test set)
On production2.5 minmedian on production (2026-09-28); slower when the service is busy
List price~$0.073 per callmeasured, at list price

Built on: Live speech to text, Speaker diarization, Grounding, Signed record

Some states require every person on a call to agree before it is recorded. Say where you and the client are, and the tool applies the strictest rule. This picks a prompt from a dated table of state laws; it is not legal advice.

Choose where you and the client are to record.

Try it live

Live

Tell everyone on the call that it is recorded, and get every party's agreement where the law requires it. The hosted demo is for the synthetic sample calls or test calls with consenting people only: never a real client call. For real calls, self-host or use the Confidential tier.

Audio is streamed to the demo server, held in memory for this session only, and deleted when it ends.

Recording is off: Choose where you and the client are to record.

Demo audio left05:00
Demo usage left100%

Live captions · Voxtral Mini 4B Realtime

Start recording or run a sample.

Live lanes

Intake checklist

waiting

While the call runs: topics covered, what is still missing, names heard so far.

Transcript (speakers)

waiting

After hang-up: Attorney, Client and anyone else on the line, one line per turn.

Audio and privacy

waiting

The audio is dropped once the transcript is made; its hash stays in the record.

Intake memo (cited)

waiting

Facts, parties, dates, deadlines, issues, goals, documents, action items, open questions. Each line cites transcript lines.

Claim check

waiting

Every memo line checked against the lines it cites; unsupported lines are marked and left out of the clean memo.

Conflicts check names

waiting

Every person and organisation named on the call, each looked up in the transcript.

Time entry

waiting

Call length rounded up to 0.1 h, UTBMS A106, and a billing narrative.

Follow-up tasks

waiting

Action items, documents to request and deadlines, with owners and dates.

Follow-up email draft

waiting

To the client, in the client's language, marked privileged and confidential.

Signed record

waiting

Consent, transcript lines, memo, time entry and audio deletion in one signed hash chain.

Receipts

Proof · signed records appear as each step finishes

Each step is signed: which model ran, and a fingerprint of what went in and came out, so it can be checked later.

Or upload a call recording

WAV, m4a or mp3, up to 5 minutes on the demo. Synthetic or consented test calls only: never a real client call on the hosted demo. The audio is held in memory and deleted once the transcript is made.

Choose where you and the client are to record.

Verify a record

Checked in your browser

A record chains every entry of a run (each transcript line or finding, each result, each model receipt), then signs the result. This check recomputes every hash and signature itself, with no call to our servers. Change any word and it fails, naming the entry.

Run a session above, or load the sample, a recorded 59 s interview.

Paste a record
Watch a recorded run first

Watch a recorded session

Live

Recorded sessions from the live system, replayed event by event.

Loading recordings

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the privileged call notes API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for the language model; Voxtral and the diarizer need about 27 GB more (a second card on our server).
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
privileged-call-notes

Use the hosted API

# Decosa Privileged call notes: use the hosted API

You are adding Decosa's Privileged call notes to this project. A lawyer's client call (live audio, or a recording
uploaded afterwards) becomes a speaker transcript, an intake memo where every line cites transcript lines and is checked
against them, the names heard for a conflicts check, a time entry rounded up to 0.1 hour with a billing narrative,
follow-up tasks, a client email draft, and a signed session record. Decosa runs open models (Voxtral Mini 4B Realtime for
live captions, MOSS-Transcribe-Diarize for speakers, Qwen3.8-27B for text); every model output comes with a signed receipt.
Use only the endpoints below. If you need something that is not listed, stop and ask me; do not guess endpoints.

**The hosted API is for synthetic calls and development.** Real client calls are privileged: run them on a self-hosted
box or on the Confidential tier. Tell me this before writing any code.

- Base URL: `https://api.decosa.ai`
- WebSocket base: `wss://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz` returns `{"ok": true, "asr": bool, "llm": bool, ...}`.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in code.
   Send `Authorization: Bearer $DECOSA_API_KEY`; WebSockets take `?token=$DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "privileged-call-notes"}` returns
   `{"token", "expires_at", "budget": {"seconds_audio": 300, "llm_tokens": 20000}}`. A 3-minute call uses about 5,000-6,000 generated tokens.
3. Over a limit the API answers HTTP 429 with `Retry-After` (seconds): wait, then retry.

## Consent first (required)
`GET https://api.decosa.ai/callnotes/consent-rules?attorney_state=CA&client_state=TX` (no token) returns
`{rule: "all-party"|"one-party", prompt, why, sources, checked}`. Show `prompt` to the lawyer before recording. Every call
then carries a consent statement `{attorney_state, client_state, all_parties_consented: bool, method:
"stated_on_call"|"written"|"one_party_state"|"other", others?: [...]}`. States are two-letter codes, `DC`, `outside-us` or
`unknown`. When any place involved is all-party, mixed or unknown, `all_parties_consented` must be true, or the API refuses
(HTTP 400, or WebSocket close 4400) with the prompt. This picks a prompt from a dated table; it is not legal advice.

## Upload a recording
`POST https://api.decosa.ai/callnotes/audio?attorney_state=..&client_state=..&all_parties_consented=true&method=stated_on_call&call_date=YYYY-MM-DD&lang=en|es[&attorney=<name>]`
with the recording as the body (WAV, m4a, mp3, ogg or webm; up to 30 minutes and 48 MB) and `Authorization`. The answer is
Server-Sent Events: `receipt` events (one per model call), then `lane` events in this order: `final_transcript`,
`privacy` (the audio was dropped: `data.audio_deleted = {audio_sha256, bytes, audio_ms}`), `memo`, `verifier`,
`conflicts`, `time_entry`, `tasks`, `email`, `record`; then `result` (everything as one object), `budget`, `done`.

## A transcript you already have
`POST https://api.decosa.ai/callnotes/notes` with `{"transcript": "Attorney: ...\nClient: ...", "consent": {...}, "call_date", "lang",
"duration_s"}` (or `segments: [{start, end, speaker, text}]`). JSON by default, SSE with `"stream": true`.

## Live audio
`WS wss://api.decosa.ai/ws/live?vertical=privileged-call-notes&token=<token>&consent=<url-encoded JSON>&call_date=YYYY-MM-DD&lang=en`
- send 16 kHz mono PCM16 little-endian frames of about 100 ms, then the text frame `{"type":"stop"}`;
- you get `ready`, `transcript` (live captions), `receipt`, a live `lane` `live` (intake checklist: covered, still to
  ask, names heard), then after stop the same lanes as the upload, and `done` with `summary` (memo, conflicts,
  time_entry, tasks, followup_email, consent_heard, claim_check, audio_deleted, record). Expect `done` 30-90 s after the call ends.

## What the lanes hold
- `memo.data`: `{matter, client, practice_area, summary, facts, dates, deadlines, issues, client_goals, documents_to_request,
  action_items, open_questions, consent_on_recording, third_party_present}`. Every item has `cites` (transcript line
  numbers) and `label` (`SUPPORTED`, `PARTIAL`, `UNSUPPORTED`). Leave `UNSUPPORTED` items out of anything you file or send.
  A deadline has `computed` (recomputed from `basis.from` + `basis.days`) and `check` (`recomputed`, `mismatch`, `computed`, `as_stated`, `no_date`).
- `conflicts.data.names`: `[{name, type, role, adverse, cites, aliases, heard}]`; `heard` is `exact`, `fuzzy`, `spelled`,
  `partial` or `not_found` (a name nobody said: check it).
- `time_entry.data`: `{hours, minutes, date, matter, narrative, utbms: {code: "A106"}, rounding, basis}`.
- `tasks.data.tasks`: `[{owner, task, due, cites, from}]`; `email.data`: `{subject, body, lang}`.
- `record.data`: the signed record (`decosa.record.v1`); check it with `POST https://api.decosa.ai/record/verify` (no token).

## Testing without a microphone
`POST https://api.decosa.ai/demo/replay` with `{"vertical":"privileged-call-notes","script_id":"cn-family-custody-ortiz"}` (also
`cn-criminal-dui-reyes` in Spanish and `cn-pi-slip-fall-thibodeaux`) runs a synthetic call through the live pipeline.
Pass if the stream ends with lanes `memo`, `conflicts`, `time_entry`, `tasks`, `email`, `privacy` and `record`, and
`done.summary.record_check.ok` is true.

## Rules
- Never send a real client call to the hosted demo. Self-host or use the Confidential tier for real calls.
- The memo is a draft for a lawyer to check; it is not legal advice and not a certified transcript.
- Show the claim check next to the memo, and never claim more than the receipts prove.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Privileged call notes: run it yourself (containers)

You are setting up Decosa Privileged call notes to run entirely on this machine's NVIDIA GPU(s), so client calls never
leave it. Nothing is sent to Decosa's hosted API and there are no Decosa charges. The local service speaks the same API as
the hosted one, so apps built against the hosted API only need a new base URL.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver. The speaker
service (MOSS-Transcribe-Diarize) is needed for uploaded recordings.

Ask me before any command that needs sudo, and show me the command first. Remind me first: get recording consent
(some states require every party's), and the memo is a draft for a lawyer to check, not legal advice.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/privileged-call-notes.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py privileged-call-notes` (the api image carries the same bundle under /app/rehearsal/privileged-call-notes/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py privileged-call-notes --bundle privileged-call-notes.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the memo has at least eight facts", "opposing counsel is on the conflicts list", "the 159-second call is a 0.1-hour entry"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
   Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
   NVIDIA's official instructions, then run `sudo nvidia-ctk runtime configure --runtime=docker` and
   `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep `DECOSA_LLM_ROUTE=direct` so prompts stay on this box. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health: poll `curl -fsS http://localhost:<PORT>/healthz` every 10 s until `"ok": true` with `"asr": true`,
   `"llm": true` and `"diarize": true`. Show me `docker compose logs --tail=50` if it is not up after 20 minutes.
7. Smoke test: `curl -fsS "http://localhost:<PORT>/callnotes/consent-rules?attorney_state=TX&client_state=CA"` should
   print `"rule": "all-party"`; then take a demo token (`POST /demo/session {"vertical":"privileged-call-notes"}`) and
   run `POST /demo/replay {"vertical":"privileged-call-notes","script_id":"cn-family-custody-ortiz"}`: the stream should
   end with a `memo`, `conflicts`, `time_entry` and `record` lane.
8. Signing key: `curl -fsS http://localhost:<PORT>/attest/signing-key` shows the key this box generated. Tell me to back
   up the data volume. Receipts from this box say `attested`: signed by our own key, an attestation, not a proof.
9. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.

## Shared network (leave it off)
Leave it off: this box handles privileged client material.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Voxtral Mini 4B Realtime needs a GPU.

  • GeForce RTX 4090lite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 40 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes.

  • GeForce RTX 5090Doesn't fit

    Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions.

  • L40Slite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 53.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (85.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (85.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Replace Voxtral Mini 4B Realtime with Voxtral Mini 4B Realtime, MLX 4-bit. MLX build for Apple Silicon.

  • Apple M5 Max, 64 GBlite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 52 GB of GPU memory at the smallest settings; 48 GB available to the GPU. The lite tier fits with changes.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API, the speech model and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with both models loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "asr": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"privileged-call-notes"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py privileged-call-notes

Download the mock-data bundle (3 KB, 10 checks)expected.json

A synthetic call between a family lawyer and a father who wants to change a custody schedule (fictional people; the transcript is the diarizer output of a recording voiced by Decosa house voices). The memo must cite the call, the conflicts list must carry opposing counsel, the time entry must be 0.1 hour under UTBMS A106, a run without a consent statement must be refused, and the signed record must verify and fail when changed.

What the rehearsal checks
  • the memo has at least eight facts
  • opposing counsel is on the conflicts list
  • the 159-second call is a 0.1-hour entry
  • the entry uses the client-communication code
  • the recording notice was heard on the call
  • every memo line was checked against the call
  • a run without a consent statement is refused
  • the signed record verifies
  • a changed record fails
  • every model call has a signed receipt

Licence: Synthetic call written for Decosa: every person, firm and court is fictional. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa Privileged call notes: run it yourself (containers)

You are setting up Decosa Privileged call notes to run entirely on this machine's NVIDIA GPU(s), so client calls never
leave it. Nothing is sent to Decosa's hosted API and there are no Decosa charges. The local service speaks the same API as
the hosted one, so apps built against the hosted API only need a new base URL.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver. The speaker
service (MOSS-Transcribe-Diarize) is needed for uploaded recordings.

Ask me before any command that needs sudo, and show me the command first. Remind me first: get recording consent
(some states require every party's), and the memo is a draft for a lawyer to check, not legal advice.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/privileged-call-notes.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py privileged-call-notes` (the api image carries the same bundle under /app/rehearsal/privileged-call-notes/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py privileged-call-notes --bundle privileged-call-notes.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the memo has at least eight facts", "opposing counsel is on the conflicts list", "the 159-second call is a 0.1-hour entry"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
   Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
   NVIDIA's official instructions, then run `sudo nvidia-ctk runtime configure --runtime=docker` and
   `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep `DECOSA_LLM_ROUTE=direct` so prompts stay on this box. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health: poll `curl -fsS http://localhost:<PORT>/healthz` every 10 s until `"ok": true` with `"asr": true`,
   `"llm": true` and `"diarize": true`. Show me `docker compose logs --tail=50` if it is not up after 20 minutes.
7. Smoke test: `curl -fsS "http://localhost:<PORT>/callnotes/consent-rules?attorney_state=TX&client_state=CA"` should
   print `"rule": "all-party"`; then take a demo token (`POST /demo/session {"vertical":"privileged-call-notes"}`) and
   run `POST /demo/replay {"vertical":"privileged-call-notes","script_id":"cn-family-custody-ortiz"}`: the stream should
   end with a `memo`, `conflicts`, `time_entry` and `record` lane.
8. Signing key: `curl -fsS http://localhost:<PORT>/attest/signing-key` shows the key this box generated. Tell me to back
   up the data volume. Receipts from this box say `attested`: signed by our own key, an attestation, not a proof.
9. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.

## Shared network (leave it off)
Leave it off: this box handles privileged client material.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

Doesn't fitPrivileged call notes on GeForce RTX 5090

Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.

Lite · one 48 GB card, uploaded calls: what changesuses estimates

  • Needs about 36 GB of GPU memory at the smallest settings; 32 GB available.
Memory per component
  • After hang-up: MOSS-Transcribe-Diarize 0.9B. ~4 GB, weights 1.8 GB (estimate). MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate.
  • Lite tier: Qwen3.8-27B (official FP8). ~33.6 GB (at least ~32 GB), weights 29 GB (from stack.json). Qwen3.8-27B FP8: 33.6 GB is the sales lite tier's allotment (stack.json). Weights of about 29 GB are an estimate (27.8B parameters at one byte, plus higher-precision layers).

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Privileged call notes, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Privileged call notes on my hardware

Fetch https://decosa.ai/prompts/privileged-call-notes-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=privileged-call-notes)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Lite · one 48 GB card, uploaded calls (lite). Fit check: doesn't fit; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- After hang-up: MOSS-Transcribe-Diarize 0.9B (OpenMOSS-Team/MOSS-Transcribe-Diarize), 4 GB
- Lite tier: Qwen3.8-27B (official FP8) (Qwen/Qwen3.8-27B-FP8), 33.6 GB

Warning: the fit check says this tier does not fit: Needs about 36 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further.

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/privileged-call-notes-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 28 Sep 2026 · measured 28 Sep 2026: · p50 148 s · p95 204 s (26 runs) · ~$0.017 per run · 44 receipts

Loading the nightly status…

Self-host: verified 28 Sep 2026 · fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume on the host network, direct route, local signing, against the already-running Voxtral, MOSS diarizer and Qwen3.8; torn down after

Measured cost to run: about $0.073 per call (hosted, 28 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Assembly prompt §6: TX/CA gives all-party; 400 without consent; custody call upload 39 s, 11 names incl. opposing counsel, 0.1 h, audio dropped, record verified (135 entries), 41 receipts attested. Live replay at 2x: 155 s, 52 captions, record verified (238 entries).

Known limits (5)
  • Not a certified transcript: speech recognition mishears names and numbers. Check anything that matters against the call.
  • The conflicts list is the names heard on the call; it does not search your conflicts database.
  • Deadlines are the ones said on the call. It recomputes the arithmetic, but it does not know court rules or limitation periods.
  • The time entry is the call audio's length rounded up to 0.1 hour; your firm's and the client's billing rules decide the entry.
  • The hosted demo is for synthetic calls. On the hosted API the server sees the audio and text in memory while it works; real client calls belong on your own box or the Confidential tier.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the privileged call notes API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for the language model; Voxtral and the diarizer need about 27 GB more (a second card on our server).
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

A client call becomes a cited intake memo, a conflicts list, a 0.1-hour time entry and follow-ups, with no notetaker bot and the audio deleted.

For solo and small-firm lawyers (family, personal injury, estate, criminal, employment) who take client calls all day and lose the recap and the 0.1-hour entry afterwards. Record the call live or upload the recording. When it ends, an open diarization model writes who said what, and the language model writes an intake memo where every line cites the transcript: facts, parties, dates, deadlines (recomputed when the call gives a start date and a number of days), issues, goals, documents to request, action items and open questions. A claim check marks lines the call does not support and leaves them out of the clean memo. You also get every name heard for your conflicts check, a time entry rounded up to 0.1 hour with a billing narrative, follow-up tasks and a client email in the client's language. A state-aware recording-consent step comes first; the audio is dropped once the transcript is made, and the session ends in a signed record. For real client calls, self-host it or run it on the Confidential tier.

Deployment
Hosted or self-host
Regulatory
Not legal advice (checked 28 Sep 2026). Recording consent: federal law and most states allow a party to record a call (one-party consent, 18 U.S.C. 2511(2)(d): https://www.law.cornell.edu/uscode/text/18/2511); California and eight other states require every party's consent (Cal. Penal Code 632: https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632), and five states are mixed. The tool asks where each side is, applies the strictest rule, and will not record until the consent step is done; it picks a prompt from a dated table and does not decide whether a recording is lawful. Confidentiality: ABA Model Rule 1.6(c) asks lawyers to make reasonable efforts to prevent unauthorized disclosure of client information, and ABA Formal Opinion 512 (29 Jul 2024) applies the duties of competence and confidentiality to generative AI tools. Sending client calls to a third-party notetaker raises privilege-waiver questions; this tool runs on open models on your own box or in an attested enclave, but you decide whether that meets your duties. Not a certified transcript: check names, dates and amounts against the call.
Architecture
Text description

A consent step first: the lawyer says where each side is and the strictest recording rule applies. Call audio (live microphone or an uploaded recording) goes to decosa-api. Live, Voxtral Mini 4B Realtime (Apache-2.0) makes captions and Qwen3.8-27B keeps an intake checklist. When the call ends, MOSS-Transcribe-Diarize 0.9B (Apache-2.0) writes who said what with a receipt per line, then the audio is dropped and its hash recorded. Qwen3.8-27B (Apache-2.0) writes a cited intake memo, lists every name for conflicts, and checks each memo line against the lines it cites; code recomputes deadlines and rounds the call length up to 0.1 hour; the model writes the billing narrative and a client email. Consent, transcript lines, memo lines, model receipts, time entry and the audio deletion go into one signed hash chain. Self-hosted or in the Confidential tier, everything stays on that box or in the attested enclave.

Architecture

At a glance

Who it's for
Solo and small-firm lawyers who take client calls all day, and the assistants who write the intake recaps and time entries.
Data retention
The call audio is held in memory only and dropped as soon as the transcript is made; its hash and the time go into the signed record. Nothing is written to disk. The memo, transcript and signed record are returned to you; the server keeps receipts (hashes), not the text.
No notetaker bot
Nothing joins your call. Record in the browser or upload the recording afterwards; speech recognition, speaker labels and the memo run on open models: on Decosa's hosted service, on your own box, or in an attested enclave.
Recording consent
You say where you and the client are before recording. When any state involved requires every party's consent (or is mixed or unknown), recording stays off until you confirm everyone agreed.
Typical run cost
A few cents or less per call at list price (a few dozen receipted model calls for a short call), plus speech recognition on Decosa's hosted service. Measured on synthetic calls.
What leaves the box (self-host)
Nothing, with DECOSA_LLM_ROUTE=direct. Anyone can verify a record offline or with POST /record/verify.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 48 GB card, uploaded calls

    Diarized transcript and the full memo, conflicts, time entry and tasks on the FP8 checkpoint; no live captions.

    Models
    • Voxtral Mini 4B Realtime
    • MOSS-Transcribe-Diarize 0.9B
    • Qwen3.8-27B (official FP8)
    Hardware
    1x L40S or RTX 6000 Ada 48 GB (not measured)
    Quality evidence
    • call memo qualitynot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyDirect route: every model call and transcript line is attested by the box's key; no gateway receipts.
  • In the hosted demo

    Standard

    the hosted demo

    Live captions and checklist, then speakers, cited memo, claim check, conflicts, time entry, tasks and email, all in a signed record.

    Models
    • Voxtral Mini 4B Realtime
    • MOSS-Transcribe-Diarize 0.9B
    • Qwen3.8-27B (NVIDIA NVFP4)
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB (hosted demo uses two cards on our server)
    Quality evidence
    • memo fact recall, 20 synthetic test calls89.3%blind grader, 20 synthetic test calls; decosa-api docs/evals/privileged-call-notes.md (28 Sep 2026)
    • conflicts-name recall, 20 synthetic test calls89.6% strict / 94.0% phonetic20 synthetic test calls; decosa-api docs/evals/privileged-call-notes.md (28 Sep 2026)
    • invented memo items left after the claim check0 of 796blind grader on the full memo; 12 items (1.5%) had a wrong detail; decosa-api docs/evals/privileged-call-notes.md (28 Sep 2026)
    Latency
    Measured on the busy shared gateway: a couple of minutes after a short call (calls in parallel); under a minute when idle. The claim check (dozens of short model calls) is the slowest stage.
    Verification
    Proof: strongLanguage-model calls: gateway-signed receipts. Speech: attested receipts from decosa-api's key.
  • Best

    DeepSeek-V4-Flash writes and checks

    Standard plus a larger memo writer and checker on two more 96 GB cards.

    Models
    • Voxtral Mini 4B Realtime
    • MOSS-Transcribe-Diarize 0.9B
    • Qwen3.8-27B (NVIDIA NVFP4)
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    Hardware
    2x RTX PRO 6000 96 GB for the writer plus the standard card
    Quality evidence
    • call memo qualitynot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlySelf-host only: model calls are attested by the box's own key.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Live, while it happens
Live captions during the call (streaming, no speakers)Voxtral Mini 4B Realtimemistralai/Voxtral-Mini-4B-Realtime-2602 on Hugging Face (opens in a new tab)
4.4B · 24 GBProof: partialIn the hosted demo
After the session
After hang-up (or on an upload): who said what, one line per turn, each with its own receiptMOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab)
0.9BProof: partialIn the hosted demo
Speaker roles, live intake checklist, cited memo, conflicts names, claim check, time-entry narrative and client emailQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
Lite tier: the same text steps on the official FP8 checkpointQwen3.8-27B (official FP8)Qwen/Qwen3.8-27B-FP8 on Hugging Face (opens in a new tab)
27.8BProof: strongSelf-host only
Best tier: memo writer and claim checkerDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active) · 192 GBNo proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

  • State recording-consent table (51 jurisdictions, checked 28 Sep 2026)Decosa, links to official statute pages

    Picks the consent prompt: the strictest rule of the states involved. 28 of 51 entries re-read on the official page; the rest are marked not re-checked.

  • Deadline arithmetic (calendar and business days, US federal holidays)Decosa (decosa_api.verticals.claims.dates)

    Recomputes a deadline said on the call from its start date and number of days, and flags a mismatch.

  • Activity code on the time entry: Communicate (with client).

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    Lane engine and HTTP/WS API (/ws/live, /demo/replay, /healthz). No GPU. Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B, served as qwen3.8-27b. Internal to the compose network.

  • decosa-asr:8000
    ${DECOSA_REGISTRY}/decosa-asr:0.1.0

    vLLM realtime endpoint for Voxtral Mini 4B Realtime (served as voxtral-realtime). Internal to the compose network.

  • decosa-diarize:8092

    MOSS-Transcribe-Diarize 0.9B pass-2 service (decosa-api services/diarize, GPU0, loopback only). decosa-api runs pass 2 on stop: diarize, role map, cited note, verifier. No published image yet; self-host builds it (see the assemble prompt).

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Hosted demo layout on our server: Voxtral and MOSS-TD on GPU0, Qwen3.8-27B on GPU1 (two cards). The one-card compose split is not measured.

  • 1x L40S / RTX 6000 Ada 48 GB, lite tier

    Not measured. FP8 language model plus the diarizer for uploaded calls; live captions need Voxtral as well.

Latency per lane

  • memo, checks, time entry after a 2.5-3 min call ends (upload path)148.0 s

    Measuredp50 148 s, p95 204 s for calls with 2.6 min median audio, 3 calls in parallel on the busy shared gateway; an idle single call took 38-46 s; decosa-api docs/evals/privileged-call-notes.md (28 Sep 2026)

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

privileged-call-notes/assemble-prompt.md181 lines
# Assemble Decosa privileged call notes on this machine

You are setting up self-hosted call notes for a law firm on this Linux machine. A client call (recorded live in the browser, or uploaded afterwards) becomes: a speaker-attributed transcript, an intake memo where every line cites transcript lines and is checked against them, the names heard for a conflicts check, a time entry rounded up to 0.1 hour with a billing narrative, follow-up tasks, a client email draft (English or Spanish), and a session record signed with this machine's own key. The call audio is held in memory and dropped once the transcript is made. Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- Client calls are privileged and confidential. Keep everything on this machine: the model route stays local (`direct`), and nothing goes to a hosted service.
- Get recording consent first. Some states (for example California, Penal Code 632) require every party's consent; the API refuses to record or process a call until a consent statement is given, and applies the strictest rule of the states involved. That is a prompt, not legal advice.
- The memo is a draft for a lawyer to check. Speech recognition mishears names and numbers; deadlines are only the ones said on the call; the time entry follows your billing rules.

Repeat these points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/privileged-call-notes.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py privileged-call-notes` (the api image carries the same bundle under /app/rehearsal/privileged-call-notes/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py privileged-call-notes --bundle privileged-call-notes.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the memo has at least eight facts", "opposing counsel is on the conflicts list", "the 159-second call is a 0.1-hour entry"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `asr` | `${DECOSA_REGISTRY}/decosa-asr:0.1.0` (vLLM 0.27.1 + `mistral-common[audio]`) | `mistralai/Voxtral-Mini-4B-Realtime-2602`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
| `diarize` | built from the decosa-api source (`services/diarize`) | `OpenMOSS-Team/MOSS-Transcribe-Diarize` @ `704aa4a9c304e8520be88901e0d1960158ef5b15`, Apache-2.0 | internal 8092 |

Any OpenAI-compatible endpoint serving an open model can replace `llm` (set `DECOSA_LLM_URL` and `DECOSA_LLM_MODEL`); the measured setup is the one above.

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4).
   - Hopper (H100/H200): set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main`, `LLM_GPU_UTIL=0.62`.
   - 48 GB Ada/L40S: FP8 checkpoint as above, plus `LLM_MAX_LEN=16384`, `LLM_GPU_UTIL=0.70`, `ASR_GPU_UTIL=0.22`, `DECOSA_LIVE_CAP=2`.
   - Under 48 GB: stop and tell me it will not fit.
   Only the Blackwell defaults have been measured; the other rows are starting points.
2. Check `docker --version` and `docker compose version`. If Docker is missing, install Docker Engine from Docker's official apt/dnf repository for this distro.
3. Check `docker run --rm --gpus all ubuntu nvidia-smi`. If it fails, install the NVIDIA Container Toolkit (`nvidia-container-toolkit`) from NVIDIA's repository, run `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
4. Confirm about 80 GB of free disk for images and weights.

## 2. Get the images

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,asr,api}:0.1.0`. If a pull fails (not published yet, or no access), build from source once the `decosa-api` source is published: clone it, then `docker compose build llm asr api` in the repo, which builds the same tags from `docker/`. If neither works, stop and tell me.

## 3. Write the compose file

Create `~/decosa-callnotes/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.60
ASR_GPU_UTIL=0.25
DECOSA_LLM_ROUTE=direct   # local model; receipts are signed by this box's own key ("attested")
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Family Law PLLC>"
DECOSA_LIVE_CAP=4
```

Create `~/decosa-callnotes/docker-compose.yml` with exactly these services:

```yaml
name: decosa-callnotes
x-gpu: &gpu
  deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    <<: *gpu
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  asr:
    image: ${DECOSA_REGISTRY}/decosa-asr:${DECOSA_TAG}
    <<: *gpu
    ipc: host
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }   # start after llm so the memory split is stable
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["--model", "mistralai/Voxtral-Mini-4B-Realtime-2602", "--tokenizer-mode", "mistral", "--config-format", "mistral",
              "--load-format", "mistral", "--compilation-config", '{"cudagraph_mode":"PIECEWISE"}', "--max-model-len", "45000",
              "--max-num-batched-tokens", "8192", "--max-num-seqs", "16", "--gpu-memory-utilization", "${ASR_GPU_UTIL}",
              "--served-model-name", "voxtral-realtime", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 600s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy }, asr: { condition: service_healthy } }
    environment:
      DECOSA_ASR_WS: ws://asr:8000/v1/realtime
      DECOSA_LLM_ROUTE: ${DECOSA_LLM_ROUTE}
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LIVE_CAP: ${DECOSA_LIVE_CAP}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_AUDIO_S: "3600"                   # per session; one call up to 30 min (DECOSA_CALLNOTES_MAX_AUDIO_S)
      DECOSA_BUDGET_LLM_TOKENS: "200000"
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_LOCAL_SIGNING: "on"                      # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_DIARIZE_URL: ""                          # step 4 sets this if you add speaker labels
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

Run `docker compose up -d`, then poll `docker compose ps` until all three are healthy (the LLM takes 5–10 minutes the first time) and `curl -s localhost:8445/healthz` shows `"asr": true, "llm": true`. If `llm` runs out of memory, lower `LLM_GPU_UTIL` or `LLM_MAX_LEN`; the two GPU shares must add up to less than about 0.9.

## 4. Speaker labels (diarize service)

The `diarize` service (MOSS-Transcribe-Diarize 0.9B, Apache-2.0) has no published image yet. If I want speaker labels, run `services/diarize` from the decosa-api source on this box (install with the `uv` commands at the top of `services/diarize/requirements.txt`, then `DIARIZE_HOST=172.17.0.1 DIARIZE_PORT=8092 DIARIZE_DEVICE=cuda:0 .venv-diarize/bin/python services/diarize/server.py`; `172.17.0.1` is the docker bridge address from `ip -4 addr show docker0`, reachable from containers and not from the LAN). Add `extra_hosts: ["host.docker.internal:host-gateway"]` to `api` and set `DECOSA_DIARIZE_URL: http://host.docker.internal:8092`. `curl -s localhost:8445/healthz` then shows `"diarize": true`. The 0.9B model needs about 2 GB of weights; lower `LLM_GPU_UTIL` to 0.55 so it fits. That fit is an estimate, not measured. For call notes this step is **required for uploads** (`POST /callnotes/audio` answers 503 without it) and strongly recommended live: without it the memo is written from the live captions, with no Attorney/Client labels.

## 5. The signing key

1. `curl -s localhost:8445/attest/signing-key` shows the public key. Show me the `pubkey`, and tell me to back up the `decosa-data` volume (it holds `/data/attest/ed25519.pem`).
2. Receipts from this box say `status: "attested"`: signed by its own key. They show nothing was changed after signing and who signed; they do not prove the model heard correctly.

## 6. Smoke test: a synthetic call, uploaded

```bash
API=localhost:8445
curl -s "$API/callnotes/consent-rules?attorney_state=TX&client_state=CA" | jq '{rule, prompt}'
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"privileged-call-notes"}' | jq -r .token)
# the bundled synthetic call (from the decosa-api source): demo_scripts/audio/cn-family-custody-ortiz.wav
curl -sN "$API/callnotes/audio?attorney_state=OH&client_state=OH&all_parties_consented=true&method=stated_on_call&call_date=2026-09-15" \
  -H "authorization: Bearer $TOKEN" -H 'content-type: audio/wav' --data-binary @cn-family-custody-ortiz.wav > /tmp/cn.sse
grep '"type": "result"' /tmp/cn.sse | sed 's/^data: //' > /tmp/cn.json
jq '{claims: .claim_check, hours: .time_entry.hours, names: [.conflicts[].name], audio_deleted: .audio_deleted.bytes, record_ok: .record_check.ok}' /tmp/cn.json
grep '"lane": "record"' /tmp/cn.sse | sed 's/^data: //' | jq '.data' > /tmp/cn-record.json
curl -s $API/record/verify -H 'content-type: application/json' --data-binary @/tmp/cn-record.json | jq '{ok, summary}'
```

Pass if: the first call prints `all-party`; the result has a `claim_check` with counts, `hours` 0.1, a names list that includes `Gregory Paskett`, a non-zero `audio_deleted`, and `record_ok: true`; and the verify prints `ok: true`. Without the consent parameters, `/callnotes/audio` answers 400.

Then the live path: `POST /demo/replay {"vertical":"privileged-call-notes","script_id":"cn-family-custody-ortiz"}` streams the same call through Voxtral in real time; the stream ends with lanes `memo`, `conflicts`, `time_entry`, `tasks`, `email`, `privacy` and `record`.

## 7. Point the app at the local API

- Base URL: `http://localhost:8445` (web app: `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`). Add other origins to `DECOSA_CORS_ORIGINS`.
- Live calls: `ws://localhost:8445/ws/live?vertical=privileged-call-notes&token=<t>&consent=<url-encoded JSON {attorney_state, client_state, all_parties_consented, method}>&call_date=YYYY-MM-DD&lang=en|es`, 16 kHz mono PCM16 frames of about 100 ms, then `{"type":"stop"}`. A missing or insufficient consent statement closes the socket with code 4400 and the prompt to show.
- Recordings: `POST /callnotes/audio` (WAV, m4a, mp3, ogg or webm; consent in the query string). Transcripts you already have: `POST /callnotes/notes` with `transcript` or `segments` and `consent`.
- `GET /callnotes/info` lists what it does and does not do; `GET /callnotes/consent-rules` the rule for a pair of states.
- The server stores nothing. Keep the memo and the signed record in your matter file; anyone can check the record with `POST /record/verify`.
- Keep the API on `127.0.0.1`; for the LAN, put a TLS reverse proxy with authentication in front and set `DECOSA_TRUSTED_PROXIES`.

## 8. Hosted gateway route (off, and leave it off)

`DECOSA_LLM_ROUTE=gateway` would send the prompts, which contain the client's words, to Decosa's hosted gateway. Never use it for client calls on a self-hosted box.

Finish with a summary: what is running, the health output, the signing key's pubkey, the smoke-test results, and the reminders above.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Legal call notes without a notetaker bot: record or upload a client call and get an intake memo where every line cites the transcript, the names for your conflicts check, a 0.1-hour time entry and follow-up tasks.
Who it's for
Solo and small-firm lawyers, and their assistants, who turn client calls into intake memos and time entries.
Where it runs
Self-host or the Confidential tier for real client calls; the hosted demo is for synthetic calls
Key numbers
  • 89.3% Memo fact recall (blind grader) (test split, n = 20)
  • 0 / 12 of 796 Memo lines invented / with a wrong detail (test split, n = 796)
  • 89.6% / 94.0% Conflicts-name recall, strict / phonetic (test split, n = 20)
  • 148.0 s Median end-to-end run, hosted (QA sweep 2026-09-28)
All results, datasets and caveats
Models
Voxtral Mini 4B Realtime (live captions) · MOSS-Transcribe-Diarize 0.9B (who said what) · Qwen3.8-27B (memo, names, claim check, time entry)
Where
Self-host or the Confidential tier for real client calls; the hosted demo is for synthetic calls
Checks
Receipt per model call and per transcript line; signed, hash-chained session record with the consent statement and the audio deletion
Industry
Legal
Output
Notes, reports and drafts · Signed record or verdict
Data
Privileged or legal · Personal data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

How do legal call notes work without a notetaker bot?

Nothing joins the call. You record it in the browser or upload the recording afterwards. An open diarization model writes who said what, and an open language model writes the intake memo, the conflicts names, the time entry and the follow-ups. Every memo line cites transcript lines and is checked against them. Self-hosted or on the Confidential tier, the audio never leaves your box or the attested enclave.

Does it handle recording consent?

Yes, as a prompt, not legal advice. Before recording you say where you and the client are. If any state involved requires every party's consent (California and eight others), or the rule is mixed or unknown, recording stays off until you confirm everyone agreed. The table of state laws was checked on 28 Sep 2026 and links each statute.

What happens to the call audio?

It is held in memory only and dropped as soon as the transcript is made. Its hash, length and the time it was dropped go into the signed record, so you can match a copy you keep. Nothing is written to disk.

How is the time entry worked out?

The call audio's length, rounded up to the next 0.1 hour by default (nearest is an option), with the ABA UTBMS activity code A106 and a one- or two-sentence billing narrative written from the checked memo. It is a draft: your firm's and the client's billing rules decide the entry.

Is the memo accurate?

On 20 synthetic test calls graded blind, the memo stated 89.3% of the key facts correctly, invented nothing (0 of 796 lines), and 1.5% of lines had a wrong detail. Conflicts-name recall was 89.6% (94.0% allowing for speech-recognition spellings). Every line cites the transcript so you can check it.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Privileged call notes

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.