Skip to content
decosa
LabsHostedSelf-hostMac

Make a tamper-evident meeting record

Live captions, then a speaker-labelled transcript and a summary whose every sentence cites transcript lines, in a signed record that shows any edit.

Measured14/17 council, 9/9 interviewSummary sentences backed by the transcript (synthetic sessions)
On production17 smedian on production (2026-09-25); slower when the service is busy
List price~$0.024 per minute of audiomeasured, at list price

Built on: Live speech to text, Speaker diarization, Grounding, Signed record

Try it live

Live

Tell everyone that the session is recorded and transcribed, and get consent where the law requires it. Demo only: use public or synthetic content. Not a certified court record.

Tick the consent box to record.

Demo audio left05:00
Demo usage left100%

Live captions · Voxtral Mini 4B Realtime

Start recording or run a sample.

Live lanes

Motions, votes and actions

waiting

Formal items as they are stated.

Record chain

waiting

Entries, head hash and signed checkpoints.

Transcript (speakers)

waiting

Written when you stop.

Summary (cited)

waiting

Every sentence cites transcript lines.

Claim check

waiting

Each sentence checked against its lines.

Signed record

waiting

Sealed and signed when the session ends.

Receipts

Proof · signed records appear as each step finishes

Each step is signed: which model ran, and a fingerprint of what went in and came out, so it can be checked later.

Verify a record

Checked in your browser

A record chains every entry of a run (each transcript line or finding, each result, each model receipt), then signs the result. This check recomputes every hash and signature itself, with no call to our servers. Change any word and it fails, naming the entry.

Run a session above, or load the sample, a recorded 59 s interview.

Paste a record
Watch a recorded run first

Watch a recorded session

Live

Recorded sessions from the live system, replayed event by event.

Loading recordings

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the tamper-evident record API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB), or 2× RTX 5090.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
record

Use the hosted API

# Decosa Tamper-evident record: use the hosted API

You are adding Decosa's Tamper-evident record to this project. Decosa runs open models (Voxtral Mini 4B Realtime for speech,
MOSS-Transcribe-Diarize for speakers, Qwen3.8-27B for text) through the Decosa API. Every model output comes with a signed receipt.
Use only the endpoints below. If you need something that is not listed, stop and ask me; do not guess endpoints.

- Base URL: `https://api.decosa.ai`
- WebSocket base: `wss://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz` returns `{"ok": true, "asr": bool, "llm": bool, "live_sessions": n, "queue": n}`.

## Auth: API key (or a demo session)
1. Preferred: an API key. Create one on the tool page with "Get an API key"; it looks like `dk_…` and is shown
   only once. Keep it in an environment variable, never in code: `DECOSA_API_KEY=dk_…`. Send
   `Authorization: Bearer $DECOSA_API_KEY` on calls that need auth; WebSockets take `?token=$DECOSA_API_KEY` in the URL.
2. Without a key, use a short demo session: `POST https://api.decosa.ai/demo/session` with JSON `{"vertical": "record"}` returns
   `{"token": "<opaque>", "expires_at": <unix seconds>, "budget": {"seconds_audio": 300, "llm_tokens": 20000}}`.
3. Send `Authorization: Bearer <token>` on calls that need it (session-bound calls such as live audio, replay, chat and
   render jobs). WebSockets take `?token=<token>` in the URL instead. These need no token: `GET /healthz`,
   `GET /demo/scripts`, `GET /demo/recordings`, `GET /demo/recordings/{id}/events`, `GET /studio/gallery`,
   `GET /studio/jobs/{id}`, `GET /receipts/{id}`, `GET /attest/signing-key`, `POST /record/verify`.
4. Demo-session limits: a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz), and a global cap on concurrent live audio sessions. Over a limit the API answers
   HTTP 429 with a `Retry-After` header (seconds): wait that long, then retry. Reuse a token until `expires_at`.
5. The API keeps no PII; session transcripts live in memory and are deleted when the session ends.

## Live audio
`WS wss://api.decosa.ai/ws/live?vertical=record&token=<token>`

Client to server:
- binary frames: 16 kHz mono PCM16 little-endian, about 100 ms each (1600 samples, 3200 bytes);
- a text frame `{"type":"stop"}` when the session is over. The server then writes the transcript, summary, claim check
  and the signed record, sends `done` and closes. Expect `done` 10-40 s after audio ends for a 1-2 minute session.

Server to client, JSON text frames:
- `{"type":"ready","vertical":"record","models":{...},"receipts":"signed","asr_receipts":"attested","signer":{"pubkey":"<hex>","key_id":"...","name":"..."},"record":{...}}`
- `{"type":"transcript","t":12.4,"text":"...","final":true,"i":3}`: these are the live captions. Partials
  (`final:false`) are replaced by the next update.
- `{"type":"receipt", ...}`: one per model call (`status:"signed"`, gateway receipt) and one per final caption
  (`kind:"asr"`, `status:"attested"`, `signer`, `sig`). `GET https://api.decosa.ai/receipts/{id}` returns either kind.
- `{"type":"lane","lane":"<lane id>","title":"...","body":"<markdown>","data":{...}}`. Each lane event replaces that lane.
- `{"type":"budget",...}`, `{"type":"error","message":"..."}`, `{"type":"done","summary":{...}}`

Lanes for `record`:
- `actions` (live): `data.items` = `[{type: motion|vote|decision|action|allegation|response|exhibit, text, who, status}]`
- `chain` (live): `data` = `{session, count, head, kinds, checkpoints, last_checkpoint, signer}`
- `final_transcript` (on stop): `data.lines` = `[{n, speaker, role, start, end, text}]`, `data.roles` = `{S01: "Chair", ...}`
- `summary` (on stop): sections SUMMARY, DECISIONS AND VOTES, ACTION ITEMS, OPEN QUESTIONS; every sentence ends in `[[n]]` or `[[n,m]]` (line numbers)
- `verifier` (on stop): `data` = `{counts, flagged, claims: [{i, section, claim, cites, label, reason}]}`
- `record` (last): `data` = the signed record

Final artifact: `done.summary.record`, format `decosa.record.v1`:
`{format, statement: {session, vertical, title, count, head, root, models, signer, signer_name, ...}, sig, entries, checkpoints, note}`.
Store it as-is (JSON). Do not re-serialise it with changed key order or number formatting before verifying; verification
recomputes hashes from the parsed values, so pretty-printing is fine but editing values is not.

## Verifying a record
- Server: `POST https://api.decosa.ai/record/verify` with the record as the body (no token, up to 8 MiB) returns
  `{ok, checks:[{name,ok,detail}], bad:[{seq,what,problems}], first_bad, summary, issued_here}`.
- Offline: canonical JSON = sorted keys, no spaces, UTF-8, no floats. Entry hash =
  sha256("decosa.chain-entry.v1\n" + canonical(entry without "hash" and "text")); `text_sha256` = sha256(text);
  each `prev` = previous entry's hash (first = 64 zeros). Merkle root: leaf = sha256(0x00 || hash), node =
  sha256(0x01 || left || right), an odd node carried up. Record signature: Ed25519 over
  "decosa.record.v1\n" + canonical(statement) with `statement.signer` as the public key.
- Compare `statement.signer` with `GET https://api.decosa.ai/attest/signing-key` to know the record came from this server.
- What it proves: the record was not changed after signing, and which key signed it. It does not prove that speech was
  recognised correctly. Not a certified court record.

## Testing without a microphone
`POST https://api.decosa.ai/demo/replay` with `{"vertical":"record","script_id":"record-council-meeting"}` (or
`record-investigation-interview`) and `Authorization: Bearer <token>` streams the same events as Server-Sent Events,
running canned synthetic audio through the real pipeline. Pass if you get `ready`, captions, `receipt` events with
`kind:"asr"`, lanes `actions`, `final_transcript`, `summary`, `verifier`, `record`, and `done.summary.record_check.ok == true`;
then POST the record to `/record/verify`, change one word in an entry's `text`, and check that `first_bad` names it.

## Rules
- Tell people they are being recorded; get consent where the law requires it. Use public or synthetic content on the
  hosted demo. Investigations, newsroom sources and research interviews belong on a self-hosted box.
- Show the receipts and the verification result next to the record; never claim more than they prove.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Tamper-evident record: run it yourself (containers)

You are setting up Decosa Tamper-evident record to run entirely on this machine's NVIDIA GPU(s). Nothing is sent to Decosa's
hosted API and there are no Decosa charges. The local service speaks the same API as the hosted one, so apps built
against the hosted API only need a new base URL.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/record.zip (630 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py record` (the api image carries the same bundle under /app/rehearsal/record/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py record --bundle record.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "final captions arrive"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
   Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
   NVIDIA's official instructions (docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html),
   then run `sudo nvidia-ctk runtime configure --runtime=docker` and `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. If it references an `.env` file or variables, ask me for any values. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health. Find the host port that compose.yaml publishes for the API (`docker compose ps`), then poll
   `curl -fsS http://localhost:<PORT>/healthz` every 10 s until it returns `"ok": true` with `"asr": true` and
   `"llm": true`. The first start can take a while as weights download. Show me `docker compose logs --tail=50` if it
   has not come up after 20 minutes.
7. Smoke test: `curl -fsS -X POST http://localhost:<PORT>/demo/session -H 'Content-Type: application/json' -d '{"vertical":"record"}'`
   should return a token.
8. Signing key: `curl -fsS http://localhost:<PORT>/attest/signing-key` shows the Ed25519 key this box generated on first
   start (kept in the data volume). Tell me to back up the volume and to publish that public key where people will
   check our records. Receipts from this box say `attested`: signed by our own key, an attestation, not a proof.
9. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/record-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Voxtral Mini 4B Realtime needs a GPU.

  • GeForce RTX 4090Doesn't fit

    Needs about 40 GB of GPU memory at the smallest settings; 24 GB available.

  • GeForce RTX 5090Doesn't fit

    Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions.

  • L40Slite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 53.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (85.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (85.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (48 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (48 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API, the speech model and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with both models loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "asr": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"record"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py record

Download the mock-data bundle (630 KB, 9 checks)expected.json

The first 30 seconds of a synthetic town council meeting, streamed from a WAV file over the live WebSocket as if it came from a microphone. The session must produce captions and end in a signed, hash-chained record that verifies, and one edited caption must be caught.

What the rehearsal checks
  • the session ends with a done event
  • the session reports no errors
  • final captions arrive
  • the captions carry the roll call of the town council
  • the server's own check of the sealed record passes
  • the signed record verifies
  • a record with one caption edited no longer verifies
  • the verifier names the first bad entry
  • every speech and model call has a signed receipt

Licence: Synthetic: a script written for Decosa (fictional town and people) read by Decosa house voices (Kokoro-82M stock voicepacks, Apache-2.0), each allowed by the consent ledger for project decosa-record-demo. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa Tamper-evident record: run it yourself (containers)

You are setting up Decosa Tamper-evident record to run entirely on this machine's NVIDIA GPU(s). Nothing is sent to Decosa's
hosted API and there are no Decosa charges. The local service speaks the same API as the hosted one, so apps built
against the hosted API only need a new base URL.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/record.zip (630 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py record` (the api image carries the same bundle under /app/rehearsal/record/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py record --bundle record.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "final captions arrive"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
   Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
   NVIDIA's official instructions (docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html),
   then run `sudo nvidia-ctk runtime configure --runtime=docker` and `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. If it references an `.env` file or variables, ask me for any values. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health. Find the host port that compose.yaml publishes for the API (`docker compose ps`), then poll
   `curl -fsS http://localhost:<PORT>/healthz` every 10 s until it returns `"ok": true` with `"asr": true` and
   `"llm": true`. The first start can take a while as weights download. Show me `docker compose logs --tail=50` if it
   has not come up after 20 minutes.
7. Smoke test: `curl -fsS -X POST http://localhost:<PORT>/demo/session -H 'Content-Type: application/json' -d '{"vertical":"record"}'`
   should return a token.
8. Signing key: `curl -fsS http://localhost:<PORT>/attest/signing-key` shows the Ed25519 key this box generated on first
   start (kept in the data volume). Tell me to back up the volume and to publish that public key where people will
   check our records. Receipts from this box say `attested`: signed by our own key, an attestation, not a proof.
9. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/record-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

Doesn't fitTamper-evident record on GeForce RTX 5090

Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.

Lite · one 48 GB card, captions only: what changesuses estimates

  • Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.
Memory per component
  • Live captions: Voxtral Mini 4B Realtime. ~24 GB (at least ~16 GB), weights 8.3 GB (from stack.json). Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more.
  • Lite tier: Qwen3.8-27B (official FP8). ~33.6 GB (at least ~32 GB), weights 29 GB (from stack.json). Qwen3.8-27B FP8: 33.6 GB is the sales lite tier's allotment (stack.json). Weights of about 29 GB are an estimate (27.8B parameters at one byte, plus higher-precision layers).

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Tamper-evident record, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Tamper-evident record on my hardware

Fetch https://decosa.ai/prompts/record-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=record)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Lite · one 48 GB card, captions only (lite). Fit check: doesn't fit; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Live captions: Voxtral Mini 4B Realtime (mistralai/Voxtral-Mini-4B-Realtime-2602), 24 GB
- Lite tier: Qwen3.8-27B (official FP8) (Qwen/Qwen3.8-27B-FP8), 33.6 GB

Warning: the fit check says this tier does not fit: Needs about 48 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further.

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/record-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 48 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh --profile live

Mac prompt for your coding agent

# Decosa Tamper-evident record: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Tamper-evident record on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 48 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/record.zip (630 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py record` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "final captions arrive"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Live captions (streaming, no speakers) | vLLM realtime WebSocket | MLX 4-bit (mlx-community/Voxtral-Mini-4B-Realtime-2602-4bit) on mlx-audio 0.5.6, behind scripts/mac/asr_server.py | Runs, measured |
| After the session: speaker-attributed transcript, one line per turn | transformers on CUDA | MLX 8-bit (vanch007/mlx-MOSS-Transcribe-Diarize-8bit) on mlx-audio, scripts/mac/diarize_server.py | Runs, measured |
| Actions lane, speaker roles, cited summary, claim verifier | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 48 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh --profile live`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model, plus about 5 GB for speech recognition and diarization), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`, `"asr": true` and `"diarize": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py record`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --profile live --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 17 s · ~$0.009 per run · 54 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers

Measured cost to run: about $0.024 per minute of audio (hosted, 25 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The step 6 smoke passed as written: the record verified (132 entries), and after one edited word verification failed at the first edited transcript line. With the diarize service the record also carried speaker labels. Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified.

Known limits (2)
  • Proves the record was not changed after signing and which key signed it; it does not prove the speech was recognised correctly. Not a certified court record.
  • Speaker labels need the diarize service; without it transcript lines have no speaker names.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the tamper-evident record API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB), or 2× RTX 5090.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

Live captions now; afterwards a speaker transcript and a cited summary, sealed in a signed record anyone can re-check.

For public meetings, internal investigations, arbitration and newsroom interviews. Captions stream while people speak. When the session ends, an open diarization model writes who said what, the language model writes a summary where every sentence cites transcript lines, and a verifier checks each sentence. Every caption, line, sentence, model receipt and audio-segment hash goes into one hash chain that the server signs. Change one word and verification fails, naming the entry. Public meetings can use the hosted API; investigations, newsrooms and research should self-host, and the box signs records with its own key.

Deployment
Hosted or self-host
Regulatory
Not a certified court record: official transcripts need a certified reporter. Get consent before recording where the law requires it. Receipts attest that the record was not changed after signing; they do not prove that speech was recognised correctly. Captions for ADA Title II web content need accuracy checks; WER on meeting audio is not measured yet.
Architecture
Text description

Audio from a microphone streams 16 kHz PCM to decosa-api. Voxtral Mini 4B Realtime (Apache-2.0) turns it into live captions; each caption gets an ASR receipt signed by decosa-api's own key, and every 5 seconds of audio is hashed. Qwen3.8-27B (Apache-2.0) keeps a list of motions, votes and actions. When the session stops, MOSS-Transcribe-Diarize 0.9B (Apache-2.0) writes speaker-attributed lines with their own receipts, Qwen3.8-27B names the roles, writes a summary where each sentence cites lines, and checks each sentence. Every caption, line, sentence, model receipt and audio hash is appended to a hash chain with signed checkpoints; the chain is sealed with a Merkle root and signed. Language-model calls on the hosted route carry gateway-signed receipts. Anyone can verify the record in a browser; editing one word breaks it. Self-hosted, everything runs on one box that signs with its own key.

Architecture

At a glance

Measured latency
Hosted p50 below is the time from the end of the audio to the final `done`, over 3 replays of the sample (2026-09-25, gateway route, one session at a time).
Typical run cost
About a cent for the council sample: a couple of dozen model calls at the gateway list price, plus caption receipts signed by the server's key at no model cost.
Data retention
The signed record is returned to you in done; the server keeps receipts (hashes), not the transcript.
What leaves the box (self-host)
Nothing, with DECOSA_LLM_ROUTE=direct. Anyone can verify a record offline or with POST /record/verify.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 48 GB card, captions only

    Captions, actions lane and a cited summary from the live captions, all chained and signed. No speaker labels.

    Models
    • Voxtral Mini 4B Realtime
    • Qwen3.8-27B (official FP8)
    Hardware
    1x L40S or RTX 6000 Ada 48 GB (not measured)
    Quality evidence
    • WER on meeting audionot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyDirect route: every model call and caption is attested by the box's key; no gateway receipts.
  • In the hosted demo

    Standard

    the hosted demo, two passes

    Live captions, then speakers, roles, cited summary and claim check; the record embeds gateway-signed receipts for every model call.

    Models
    • Voxtral Mini 4B Realtime
    • MOSS-Transcribe-Diarize 0.9B
    • Qwen3.8-27B (NVIDIA NVFP4)
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB (hosted demo uses two cards on our server)
    Quality evidence
    • MOSS-TD WER / DER on PriMock57 (clinical proxy)10.3 / 11.4scribe-bench RESULTS.md
    • claim check on the synthetic sessions: sentences supported19/19 and 9/9measured on our server 2026-09-23: a decosa-api pre-release test instance, scripts/replay_client.py, one session at a time, gateway route
    • WER on meeting audionot measured yet
    Latency
    measured on our server: the signed record under a minute after a meeting ends, seconds after a short interview
    Verification
    Proof: strongLanguage-model calls: gateway-signed receipts. Speech: attested receipts from decosa-api's key.
  • Best

    DeepSeek-V4-Flash writes and checks

    Standard plus a larger writer and verifier on two more 96 GB cards.

    Models
    • Voxtral Mini 4B Realtime
    • MOSS-Transcribe-Diarize 0.9B
    • Qwen3.8-27B (NVIDIA NVFP4)
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    Hardware
    2x RTX PRO 6000 96 GB for the writer plus the standard card
    Quality evidence
    • summary quality on meetingsnot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyNot a hosted model: model calls are attested by the box's key only.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Live, while it happens
Live captions (streaming, no speakers)Voxtral Mini 4B Realtimemistralai/Voxtral-Mini-4B-Realtime-2602 on Hugging Face (opens in a new tab)
4.4B · 24 GBProof: partialIn the hosted demo
After the session
After the session: speaker-attributed transcript, one line per turnMOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab)
0.9BProof: partialIn the hosted demo
Actions lane, speaker roles, cited summary, claim verifierQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
Lite tier: actions lane and a cited summary written from the live captionsQwen3.8-27B (official FP8)Qwen/Qwen3.8-27B-FP8 on Hugging Face (opens in a new tab)
27.8BProof: strongSelf-host only
Best tier: summary writer and verifierDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active) · 192 GBNo proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    Lane engine and HTTP/WS API (/ws/live, /demo/replay, /healthz). No GPU. Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B, served as qwen3.8-27b. Internal to the compose network.

  • decosa-asr:8000
    ${DECOSA_REGISTRY}/decosa-asr:0.1.0

    vLLM realtime endpoint for Voxtral Mini 4B Realtime (served as voxtral-realtime). Internal to the compose network.

  • decosa-diarize:8092

    MOSS-Transcribe-Diarize 0.9B pass-2 service (decosa-api services/diarize, GPU0, loopback only). decosa-api runs pass 2 on stop: diarize, role map, cited note, verifier. No published image yet; self-host builds it (see the assemble prompt).

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Hosted demo layout on our server: Voxtral and MOSS-TD on GPU0, Qwen3.8-27B on GPU1 (two cards). The one-card compose split is not measured.

  • 1x L40S / RTX 6000 Ada 48 GB, lite tier

    Not measured. FP8 LLM plus Voxtral; no diarizer, so the record uses the live captions as transcript lines.

  • 2x RTX PRO 6000 96 GB + a card for the live pass, best tier

    DeepSeek-V4-Flash ran on our server across two cards (TP2); the full record stack on it is not measured.

Latency per lane

  • first caption1.4 s

    Measuredmeasured on our server 2026-09-23: a decosa-api pre-release test instance, scripts/replay_client.py, one session at a time, gateway route

  • actions (median per update)2.0 s

    Measuredmeasured on our server 2026-09-23: a decosa-api pre-release test instance, scripts/replay_client.py, one session at a time, gateway route (1.2-2.0 s)

  • signed record after audio ends, 91.7 s meeting35.0 s

    Measuredmeasured on our server 2026-09-23: a decosa-api pre-release test instance, scripts/replay_client.py, one session at a time, gateway route (32.4-38.9 s over two runs: diarize 6-10, roles 1, summary 4, verifier 14-25)

  • signed record after audio ends, 58.8 s interview11.2 s

    Measuredmeasured on our server 2026-09-23: a decosa-api pre-release test instance, scripts/replay_client.py, one session at a time, gateway route

  • verify in the browser (131-entry record)n/a

    Measurednot measured yet

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

record/assemble-prompt.md173 lines
# Assemble the Decosa tamper-evident record on this machine

You are setting up a self-hosted tamper-evident record on this Linux machine: live captions, a list of motions, votes and actions, then (when a session stops) a speaker-attributed transcript, a summary where every sentence cites transcript lines, a claim check, and a session record signed with this machine's own key. Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:** tell everyone that a session is being recorded and get consent where the law requires it. This is not a certified court record. Repeat both points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/record.zip (630 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py record` (the api image carries the same bundle under /app/rehearsal/record/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py record --bundle record.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "final captions arrive"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `asr` | `${DECOSA_REGISTRY}/decosa-asr:0.1.0` (vLLM 0.27.1 + `mistral-common[audio]`) | `mistralai/Voxtral-Mini-4B-Realtime-2602`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
| `diarize` (optional) | built from the decosa-api source (`services/diarize`) | `OpenMOSS-Team/MOSS-Transcribe-Diarize` @ `704aa4a9c304e8520be88901e0d1960158ef5b15`, Apache-2.0 | internal 8092 |

Audio, captions, transcripts and records stay on this machine. Without `diarize`, records still work: the live captions become the transcript lines, without speaker labels.

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4).
   - Hopper (H100/H200): set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main`, `LLM_GPU_UTIL=0.62`.
   - 48 GB Ada/L40S: FP8 checkpoint as above, plus `LLM_MAX_LEN=16384`, `LLM_GPU_UTIL=0.70`, `ASR_GPU_UTIL=0.22`, `DECOSA_LIVE_CAP=2`.
   - Under 48 GB: stop and tell me it will not fit.
   Only the Blackwell defaults have been measured; the other rows are starting points.
2. Check `docker --version` and `docker compose version`. If Docker is missing, install Docker Engine from Docker's official apt/dnf repository for this distro.
3. Check `docker run --rm --gpus all ubuntu nvidia-smi`. If it fails, install the NVIDIA Container Toolkit (`nvidia-container-toolkit`) from NVIDIA's repository, run `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
4. Confirm about 80 GB of free disk for images and weights.

## 2. Get the images

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,asr,api}:0.1.0`. If a pull fails (not published yet, or no access), build from source once the `decosa-api` source is published: clone it, then `docker compose build llm asr api` in the repo, which builds the same tags from `docker/`. If neither works, stop and tell me.

## 3. Write the compose file

Create `~/decosa-record/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.60
ASR_GPU_UTIL=0.25
DECOSA_LLM_ROUTE=direct   # local model; receipts are signed by this box's own key ("attested")
DECOSA_SIGNER_NAME="<who signs these records, e.g. Riverbend town clerk>"
DECOSA_LIVE_CAP=4
```

Create `~/decosa-record/docker-compose.yml` with exactly these services:

```yaml
name: decosa-record
x-gpu: &gpu
  deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    <<: *gpu
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  asr:
    image: ${DECOSA_REGISTRY}/decosa-asr:${DECOSA_TAG}
    <<: *gpu
    ipc: host
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }   # start after llm so the memory split is stable
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["--model", "mistralai/Voxtral-Mini-4B-Realtime-2602", "--tokenizer-mode", "mistral", "--config-format", "mistral",
              "--load-format", "mistral", "--compilation-config", '{"cudagraph_mode":"PIECEWISE"}', "--max-model-len", "45000",
              "--max-num-batched-tokens", "8192", "--max-num-seqs", "16", "--gpu-memory-utilization", "${ASR_GPU_UTIL}",
              "--served-model-name", "voxtral-realtime", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 600s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy }, asr: { condition: service_healthy } }
    environment:
      DECOSA_ASR_WS: ws://asr:8000/v1/realtime
      DECOSA_LLM_ROUTE: ${DECOSA_LLM_ROUTE}
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LIVE_CAP: ${DECOSA_LIVE_CAP}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_AUDIO_S: "3600"
      DECOSA_BUDGET_LLM_TOKENS: "200000"
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_LOCAL_SIGNING: "on"                      # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_DIARIZE_URL: ""                          # step 4 sets this if you add speaker labels
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

Run `docker compose up -d`, then poll `docker compose ps` until all three are healthy (the LLM takes 5–10 minutes the first time) and `curl -s localhost:8445/healthz` shows `"asr": true, "llm": true`. If `llm` runs out of memory, lower `LLM_GPU_UTIL` or `LLM_MAX_LEN`; the two GPU shares must add up to less than about 0.9.

## 4. Optional: speaker labels (diarize service)

The `diarize` service (MOSS-Transcribe-Diarize 0.9B, Apache-2.0) has no published image yet. If I want speaker labels, run `services/diarize` from the decosa-api source on this box (install with the `uv` commands at the top of `services/diarize/requirements.txt`, then `DIARIZE_HOST=172.17.0.1 DIARIZE_PORT=8092 DIARIZE_DEVICE=cuda:0 .venv-diarize/bin/python services/diarize/server.py`; `172.17.0.1` is the docker bridge address from `ip -4 addr show docker0`, reachable from containers and not from the LAN). Add `extra_hosts: ["host.docker.internal:host-gateway"]` to `api` and set `DECOSA_DIARIZE_URL: http://host.docker.internal:8092`. `curl -s localhost:8445/healthz` then shows `"diarize": true`. The 0.9B model needs about 2 GB of weights; lower `LLM_GPU_UTIL` to 0.55 so it fits. That fit is an estimate, not measured. Without it the record still works: `done.summary.pass2` reads `unavailable` and the transcript lines have no speaker names. If the source isn't available, skip this step.

## 5. The signing key

1. `curl -s localhost:8445/attest/signing-key` returns `{"scheme":"ed25519","pubkey":"<64 hex>","key_id":...,"name":...}`. Show me the `pubkey`.
2. Tell me to **back up the `decosa-data` volume**: it holds the private key (`/data/attest/ed25519.pem`, mode 0600). With a new key, new records still verify, but old records no longer show as issued by this box.
3. Tell me to publish the public key where people will check my records (website, agenda, case file).
4. Receipts from this box say `status: "attested"`: signed by our own key. They prove that nothing was changed after signing and who signed. They do not prove that the model ran or heard correctly.

## 6. Smoke test: record, verify, tamper

```bash
TOKEN=$(curl -s localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"record"}' | jq -r .token)
curl -sN localhost:8445/demo/replay -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
  -d '{"vertical":"record","script_id":"record-council-meeting"}' > /tmp/record.sse
grep -oE '"(type|lane|kind)": "[a-z_]+"' /tmp/record.sse | sort | uniq -c
grep '"type": "done"' /tmp/record.sse | sed 's/^data: //' | jq '.summary.record' > /tmp/record.json
curl -s localhost:8445/record/verify -H 'content-type: application/json' --data-binary @/tmp/record.json | jq '{ok, summary}'
jq '(.entries[] | select(.kind=="line") | .text) |= sub("the"; "not the")' /tmp/record.json > /tmp/edited.json
curl -s localhost:8445/record/verify -H 'content-type: application/json' --data-binary @/tmp/edited.json | jq '{ok, first_bad, summary}'
```

The meeting is 91.7 s of synthetic audio, paced in real time. Pass if:
- the stream has `ready` with `asr_receipts: "attested"`, `transcript` events, `receipt` events with `kind: "asr"` and with `status: "attested"`, and lanes `actions`, `chain`, `final_transcript`, `summary`, `verifier`, `record`;
- the first verify prints `ok: true`;
- the second prints `ok: false`, with `first_bad` at the first edited transcript line.

The other script is `record-investigation-interview` (58.8 s).

## 7. Point the app at the local API

- Base URL: `http://localhost:8445` (for the web app, `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`). Add the app's origin to `DECOSA_CORS_ORIGINS` if it is not localhost.
- `POST /demo/session {"vertical":"record"}` returns a token. Open `ws://localhost:8445/ws/live?vertical=record&token=<t>`, send 16 kHz mono PCM16 little-endian frames of about 100 ms, then `{"type":"stop"}`. Save `done.summary.record` as JSON: that file is the record. Close code `4403` means the origin is not allowed.
- Anyone can check a record: `POST /record/verify`, or the "Verify a record" panel on the Decosa site, which checks it in the browser.
- Keep the API on `127.0.0.1`. To reach it from the LAN, put a TLS reverse proxy in front and set `DECOSA_TRUSTED_PROXIES`.

Finish with a summary: what is running, the health output, the signing key's pubkey, both verify results, and the consent and not-a-court-record reminders.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

1 law, rule and guidance page cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Live captions now; afterwards a speaker transcript and a cited summary, sealed in a signed record anyone can re-check.
Who it's for
Teams in public sector and compliance and trust.
Where it runs
Hosted for public meetings, self-host for investigations
Key numbers
  • 14/17 council, 9/9 interview Claim check on the synthetic sessions: sentences supported (synthetic, n = 26)
  • 5 speakers found; one short turn given to the wrong member Speaker attribution on the synthetic council meeting (5 TTS voices) (synthetic, n = 5)
  • 16.6 s Median end-to-end run, hosted (QA sweep 2026-09-25)
All results, datasets and caveats
Models
Voxtral 4B · MOSS-TD 0.9B · Qwen3.8-27B
Where
Hosted for public meetings, self-host for investigations
Checks
Signed hash chain over every line
Output
Signed record or verdict · Notes, reports and drafts
Data
Personal data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

What is a tamper-evident meeting transcript in Decosa?

Decosa's tamper-evident record puts every live caption, transcript line, summary sentence, model receipt and audio-segment hash into one hash chain that the server signs. Change one word and verification fails, naming the entry. Anyone can verify a record offline or with POST /record/verify. In a self-host test, a 132-entry record verified, and after one edited word verification failed at the first edited transcript line.

Does a signed record prove the transcript is accurate?

No. Decosa's tamper-evident record proves the record was not changed after signing and which key signed it; it does not prove the speech was recognised correctly. Word error rate on meeting or interview audio has not been measured yet; published ASR numbers are on clinical consultations. Speaker labels need the diarize service, and in one synthetic council test two male TTS voices were merged.

Is it an official court transcript or approved minutes?

No. Decosa's tamper-evident record is not a certified court record; official transcripts need a certified reporter. It produces a speaker transcript and a summary where every sentence cites transcript lines, which staff can review. Get consent before recording where the law requires it. Public meetings can use the hosted API; investigations, newsrooms and research should self-host so the box signs with its own key.

How well does the cited meeting summary hold up?

In Decosa's tamper-evident record, a verifier checks each summary sentence against the transcript lines it cites. On two synthetic TTS sessions it found 14 of 17 council sentences and 9/9 interview sentences supported (2 of the 3 flags come from one summary sentence split at "Mr."). That is two sessions only: summary quality on meetings and verifier recall on meeting minutes have not been measured yet.

Can it provide live captions for ADA Title II public meetings?

Decosa's tamper-evident record streams live captions while people speak, but captions for ADA Title II web content need accuracy checks, and word error rate on meeting audio is not measured yet. The server keeps receipts (hashes), not the transcript; the signed record is returned to you. The 92 s council sample cost about $0.0085 at the gateway list price.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Tamper-evident record

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.