Translation
waitingStreamed per sentence.
Live translated captions about a second behind the speaker, with names and terms kept consistent, and a short summary at the end.
Built on: Live speech to text
Samples use their own language pair.
Start recording or run a sample.
Streamed per sentence.
Terms and how they were rendered.
Written when you stop.
Each step is signed: which model ran, and a fingerprint of what went in and came out, so it can be checked later.
Recorded sessions from the live system, replayed event by event.
Loading recordings
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)# Decosa Live translation: use the hosted API
You are adding Decosa's Live translation to this project. Decosa runs open models (Voxtral Mini 4B Realtime for speech,
Qwen3.8-27B for text) through the Decosa API. Every model output comes with a signed receipt.
Use only the endpoints below. If you need something that is not listed, stop and ask me; do not guess endpoints.
- Base URL: `https://api.decosa.ai`
- WebSocket base: `wss://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz` returns `{"ok": true, "asr": bool, "llm": bool, "live_sessions": n, "queue": n}`.
## Auth: API key (or a demo session)
1. Preferred: an API key. Create one on the tool page with "Get an API key"; it looks like `dk_…` and is shown
only once. Keep it in an environment variable, never in code: `DECOSA_API_KEY=dk_…`. Send
`Authorization: Bearer $DECOSA_API_KEY` on calls that need auth; WebSockets take `?token=$DECOSA_API_KEY` in the URL.
2. Without a key, use a short demo session: `POST https://api.decosa.ai/demo/session` with JSON `{"vertical": "translate"}` returns
`{"token": "<opaque>", "expires_at": <unix seconds>, "budget": {"seconds_audio": 300, "llm_tokens": 20000}}`.
3. Send `Authorization: Bearer <token>` on calls that need it (session-bound calls such as live audio, replay, chat and
render jobs). WebSockets take `?token=<token>` in the URL instead. These need no token: `GET /healthz`,
`GET /demo/scripts`, `GET /demo/recordings`, `GET /demo/recordings/{id}/events`, `GET /studio/gallery`,
`GET /studio/jobs/{id}`, `GET /receipts/{id}`.
4. Demo-session limits: a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz), and a global cap on concurrent live audio sessions. Over a limit the API answers
HTTP 429 with a `Retry-After` header (seconds): wait that long, then retry. Reuse a token until `expires_at`.
5. The API keeps no PII; session transcripts live in memory and are deleted when the session ends.
## Live audio
`WS wss://api.decosa.ai/ws/live?vertical=translate&token=<token>&lang=<src>&target=<dst>`
`lang` is the spoken language and `target` the output language, as short codes such as `es` and `en`. They apply to `translate` only.
Client to server:
- binary frames: 16 kHz mono PCM16 little-endian, about 100 ms each (1600 samples, 3200 bytes);
- a text frame `{"type":"stop"}` when the speaker is done. The server then writes final lanes, sends `done` and closes.
Server to client, JSON text frames:
- `{"type":"ready","vertical":"translate","models":{"asr":"...","llm":"..."}}`
- `{"type":"transcript","t":12.4,"text":"...","final":true}` (partials have `final:false` and are replaced by the next update)
- `{"type":"lane","lane":"<lane id>","title":"<Title>","body":"<markdown or text>","data":{...optional...},"latency_ms":850}`
Each lane event replaces the previous content of that lane.
- `{"type":"receipt","id":"<completion id>","model":"qwen3.8-27b","gateway_sig":"<hex>","provider":"<provider id>"}`
- `{"type":"budget","seconds_audio_left":240,"llm_tokens_left":15000}`
- `{"type":"error","message":"..."}`
- `{"type":"done","summary":{...final artifact...}}`
Lanes for `translate`:
- `translation`: the translation in the target language, streamed per sentence
- `glossary`: terms and how they were rendered
- `summary`: a summary, written when you send stop
Final artifact in `done.summary`: the translation and summary.
Close codes the server uses:
- `4401` bad or expired token: get a new session, then reconnect.
- `4409` token already has a live socket: close the other one, or get a new session.
- `4429` busy (live-session cap): wait, then retry. Back off at least 15 s.
- `4400` bad parameters, `4402` budget exhausted, `4403` vertical or origin mismatch: do not retry; fix the cause.
On any other drop before `done`, reconnect with backoff (0.5 s, 1 s, 2 s, ... up to 8 s) using the same token while it
is valid. Stop sending audio when `seconds_audio_left` reaches 0.
## Testing without a microphone
`POST https://api.decosa.ai/demo/replay` with `Authorization: Bearer <token>` and JSON
`{"vertical":"translate","script_id":"<script id>"}` streams the same event types over SSE (`text/event-stream`, one JSON
event per `data:` line). It runs a canned audio script through the real pipeline, so the output is live.
Script ids for `translate`: `translate-es-en-repair` (Spanish to English), `translate-en-es-tour` (English to Spanish). `GET https://api.decosa.ai/demo/scripts` lists all scripts (no token).
## Receipts (public, read-only)
`GET https://api.decosa.ai/receipts/{id}` returns
`{"id", "model", "weights_root", "request_hash", "output_hash", "provider": {"miner_id", "pubkey", "sig"}, "gateway": {"pubkey", "sig"}, "proof": {"format", "verified"}, "checks": [{"name", "ok", "detail"}]}`.
## Audio format
Convert a recording to the wire format with ffmpeg:
`ffmpeg -i input.wav -ar 16000 -ac 1 -f s16le input.pcm`
## Example: TypeScript (Node 22+, global fetch and WebSocket)
```ts
import { readFileSync } from "node:fs";
const BASE = "https://api.decosa.ai";
// Prefer your API key (dk_…); fall back to a short demo session.
let token = process.env.DECOSA_API_KEY;
if (!token) {
const session = await fetch(`${BASE}/demo/session`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ vertical: "translate" }),
});
if (session.status === 429) throw new Error(`busy, retry after ${session.headers.get("Retry-After")} s`);
({ token } = await session.json());
}
const ws = new WebSocket(
`${BASE.replace(/^http/, "ws")}/ws/live?vertical=translate&token=${encodeURIComponent(token)}&lang=es&target=en`,
);
ws.binaryType = "arraybuffer";
ws.onmessage = (m) => {
const ev = JSON.parse(String(m.data));
if (ev.type === "transcript" && ev.final) console.log("heard:", ev.text);
if (ev.type === "lane") console.log(`[${ev.lane}]`, ev.body);
if (ev.type === "receipt") console.log("receipt:", `${BASE}/receipts/${ev.id}`);
if (ev.type === "done") { console.log("final:", ev.summary); ws.close(); }
};
ws.onopen = async () => {
const pcm = readFileSync("input.pcm"); // 16 kHz mono PCM16 LE
for (let i = 0; i < pcm.length; i += 3200) {
ws.send(pcm.subarray(i, i + 3200));
await new Promise((r) => setTimeout(r, 100)); // real time
}
ws.send(JSON.stringify({ type: "stop" }));
};
```
## Example: Python (`pip install requests websockets`)
```python
import asyncio, json, os, requests, websockets
BASE = "https://api.decosa.ai"
token = os.environ.get("DECOSA_API_KEY") # dk_… API key preferred
if not token:
r = requests.post(f"{BASE}/demo/session", json={"vertical": "translate"})
if r.status_code == 429:
raise SystemExit(f"busy, retry after {r.headers.get('Retry-After')} s")
token = r.json()["token"]
async def main():
url = BASE.replace("http", "ws", 1) + f"/ws/live?vertical=translate&token={token}&lang=es&target=en"
async with websockets.connect(url) as ws:
async def send_audio():
with open("input.pcm", "rb") as f: # 16 kHz mono PCM16 LE
while chunk := f.read(3200):
await ws.send(chunk)
await asyncio.sleep(0.1)
await ws.send(json.dumps({"type": "stop"}))
sender = asyncio.create_task(send_audio())
async for msg in ws:
ev = json.loads(msg)
if ev["type"] == "lane":
print(f"[{ev['lane']}]", ev["body"])
elif ev["type"] == "receipt":
print("receipt:", f"{BASE}/receipts/{ev['id']}")
elif ev["type"] == "done":
print("final:", ev["summary"])
break
await sender
asyncio.run(main())
```
## What to build
1. A small client module for the calls above (session, socket, reconnect, 429 handling).
2. UI or CLI output that shows the transcript, each lane, and a link to every receipt.
3. A config value for the base URL, so it can point at a self-hosted box later with no code change.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa Live translation: run it yourself (containers)
You are setting up Decosa Live translation to run entirely on this machine's NVIDIA GPU(s). Nothing is sent to Decosa's
hosted API and there are no Decosa charges. The local service speaks the same API as the hosted one, so apps built
against the hosted API only need a new base URL.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/translate.zip (634 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py translate` (the api image carries the same bundle under /app/rehearsal/translate/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py translate --bundle translate.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the Spanish captions carry the sink and Thursday"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
NVIDIA's official instructions (docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html),
then run `sudo nvidia-ctk runtime configure --runtime=docker` and `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. If it references an `.env` file or variables, ask me for any values. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health. Find the host port that compose.yaml publishes for the API (`docker compose ps`), then poll
`curl -fsS http://localhost:<PORT>/healthz` every 10 s until it returns `"ok": true` with `"asr": true` and
`"llm": true`. The first start can take a while as weights download. Show me `docker compose logs --tail=50` if it
has not come up after 20 minutes.
7. Smoke test: `curl -fsS -X POST http://localhost:<PORT>/demo/session -H 'Content-Type: application/json' -d '{"vertical":"translate"}'`
should return a token.
8. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/translate-mac.md instead.
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Voxtral Mini 4B Realtime needs a GPU.
Needs about 36 GB of GPU memory at the smallest settings; 24 GB available.
Needs about 44 GB of GPU memory at the smallest settings; 32 GB available.
The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions.
The standard tier does not fit: Needs about 49.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes.
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
The standard tier fits (81.6 of 96 GB).
The standard tier fits (81.6 of 192 GB).
The standard tier fits (48 of 96 GB).
The standard tier fits (48 of 64 GB).
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
One file describes the API, the speech model and the language model as services.
mkdir -p ~/decosa && cd ~/decosa
curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yamlThe first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
Wait until the API reports ok with both models loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz
# {"ok": true, "asr": true, "llm": true, ...}
curl -fsS -X POST http://localhost:<PORT>/demo/session \
-H 'Content-Type: application/json' -d '{"vertical":"translate"}'expected.json. Every check must print PASS.docker compose exec api python scripts/rehearse.py translate
Download the mock-data bundle (634 KB, 8 checks)expected.json
Two cuts from a synthetic Spanish phone call booking a plumber (a leak under the kitchen sink; a technician on Thursday between 8 and 10; the visit costs 45 dollars; the address), streamed over the live WebSocket with lang=es and target=en. Every Spanish sentence must come back translated into English as it is said, and the call must end with an English summary that keeps the day, the price and the address.
Licence: Synthetic: a script written for Decosa (no real people or companies) read by Decosa house voices (Kokoro-82M stock voicepacks, Apache-2.0), each allowed by the consent ledger for project decosa-translate-demo. Part of decosa-api, AGPL-3.0-or-later.
# Decosa Live translation: run it yourself (containers)
You are setting up Decosa Live translation to run entirely on this machine's NVIDIA GPU(s). Nothing is sent to Decosa's
hosted API and there are no Decosa charges. The local service speaks the same API as the hosted one, so apps built
against the hosted API only need a new base URL.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/translate.zip (634 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py translate` (the api image carries the same bundle under /app/rehearsal/translate/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py translate --bundle translate.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the Spanish captions carry the sink and Thursday"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
NVIDIA's official instructions (docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html),
then run `sudo nvidia-ctk runtime configure --runtime=docker` and `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. If it references an `.env` file or variables, ask me for any values. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health. Find the host port that compose.yaml publishes for the API (`docker compose ps`), then poll
`curl -fsS http://localhost:<PORT>/healthz` every 10 s until it returns `"ok": true` with `"asr": true` and
`"llm": true`. The first start can take a while as weights download. Show me `docker compose logs --tail=50` if it
has not come up after 20 minutes.
7. Smoke test: `curl -fsS -X POST http://localhost:<PORT>/demo/session -H 'Content-Type: application/json' -d '{"vertical":"translate"}'`
should return a token.
8. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/translate-mac.md instead.
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Doesn't fitLive translation on GeForce RTX 5090
Needs about 44 GB of GPU memory at the smallest settings; 32 GB available.
Nothing: it runs as listed in the stack.
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
The self-host prompt for Live translation, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Live translation on my hardware Fetch https://decosa.ai/prompts/translate-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=translate) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Specialist translation models (alternate-specialist-mt). Fit check: runs, about 9.4 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Specialist translation model for the translat...: Hy-MT2-7B / Hy-MT2-30B-A3B-FP8 (tencent/Hy-MT2-7B), 9.4 GB GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Hy-MT2-7B / Hy-MT2-30B-A3B-FP8 ~9.4 GB (29%); about 22.6 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/translate-assemble.md
No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 48 GB of unified memory or more. Measured speeds and what runs where
From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh --profile live
# Decosa Live translation: run it on this Mac (Apple Silicon, no NVIDIA GPU) You are setting up the Decosa Live translation on this Mac, natively on Apple Silicon. The models run on the Mac's GPU through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API. Every model this tool needs runs on the Mac. It needs 48 GB of unified memory or more. Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop or kill a process this setup did not start; if a port is taken, pick another one. ## Step 0: set up with a coding agent, rehearse on mock data, then go private This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works. Work in this order: 1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to "test with something realistic". 2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool, https://decosa.ai/samples/translate.zip (634 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json), show me what is in it, and run the rehearsal against the local API: `.venv/bin/python scripts/rehearse.py translate` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key). It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the Spanish captions carry the sink and Thursday"). Show me the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json` to make a check pass. 3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this machine. For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent can read. Switch to your own data only after the rehearsal has passed and the agent's work is done. ## What runs where | Part | On an NVIDIA GPU | On this Mac | Status | |---|---|---|---| | Speech recognition (streaming) | vLLM realtime WebSocket | MLX 4-bit (mlx-community/Voxtral-Mini-4B-Realtime-2602-4bit) on mlx-audio 0.5.6, behind scripts/mac/asr_server.py | Runs, measured | | Translation lane model (translation, glossary, summary) | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured | ## Steps 1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and `sysctl -n hw.memsize` should be at least 48 GB for this tool. Check about 30 GB of free disk with `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory. 2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`. 3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`. Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me. 4. Start everything with one command: `scripts/mac/setup.sh --profile live`. It creates `.venv` (decosa-api) and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the language model, plus about 5 GB for speech recognition and diarization), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them. If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`. 5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must report `"llm": true`, `"asr": true` and `"diarize": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key: show it to me, because it is what others pin to check the receipts and records this Mac signs. 6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py translate`. It runs the tool's own sample end to end against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts. `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found. 7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`, the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`. 8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of `scripts/mac/setup.sh status`. ## Good to know - Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a self-hosted Mac. - The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published evals use. Expect small differences in wording and scores. - Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --profile live --engine omlx` serves the model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel calls; typed judgments then use sampling because oMLX returns no log-probabilities). - Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details: `docs/self-host-mac.md` in the checkout.
Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 2.3 s · ~$0.003 per run · 18 receipts
Loading the nightly status…
Self-host: verified 25 Sep 2026 · fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers
Measured cost to run: about $0.27 per 100 minutes of audio (hosted, 25 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The step 4 replay passed as written (14 translation lanes, glossary, summary, 18 attested receipts; a translation 157 ms median after its sentence), and the step 5 captions overlay translated live fake-microphone audio in headless Chromium once its createScriptProcessor buffer size was fixed. Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified.
Streams speech through open-weight speech recognition and translates each sentence as soon as it ends, so captions trail the speaker by about a second. Every 20 seconds or so it rebuilds a glossary of names and domain terms and feeds it back into later translations to keep them consistent. When you stop, it writes a short summary in the target language. It is for tours, meetings, customer calls and events where a caption track helps. It is not an interpreter for court, medical or other settings where a mistranslation carries legal or clinical risk.
Architecture of the live translation stack. A microphone streams 16 kHz PCM16 audio over a WebSocket to decosa-api, which feeds Voxtral Mini 4B Realtime (open weights, Apache-2.0) for streaming speech recognition in 13 source languages. Each finished sentence goes to the translation lane on Qwen3.8-27B NVFP4 (open weights, Apache-2.0), which emits three lanes: translation (one event per sentence, ordered by data.i), glossary (every 20 seconds, fed back into later translations) and summary (on stop). A captions overlay consumes the WebSocket and shows the translated sentences. Everything inside the dashed box stays on your machine when self-hosted. On the gateway route, each LLM call produces a signed receipt with request, output and weights hashes; the gateway countersigns it with Ed25519; and anyone can re-check it at GET /receipts/{id}. On the direct route receipts are signed by the box's own key (attested). The speech recognition step emits no receipt of its own.
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
one 48 GB card
A smaller mixture-of-experts model on cheaper hardware. Translations are likely less polished and no quality has been measured; no signed receipts.
one 96 GB card (hosted demo)
The stack the hosted demo runs: Qwen3.8-27B NVFP4 with MTP. Every lane call can carry a gateway-signed receipt.
two 96 GB cards
DeepSeek-V4-Flash (284B MoE) on two cards: the strongest model measured on the owner's own benchmarks, at twice the hardware and without signed receipts.
the largest open flash models
GLM-5.3-Flash and DeepSeek-V4.1-Flash translate, served by network providers; speech recognition stays on your machine. Not served yet.
We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Speech recognition (streaming)Voxtral Mini 4B Realtimemistralai/Voxtral-Mini-4B-Realtime-2602 on Hugging Face (opens in a new tab) 4.4B · 24 GBProof: partial | LiteStandardBestWanted | 4.4B · 24 GB | Proof: partial | |
| ||||
Translation lane model (translation, glossary, summary)Qwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 57.6 GBProof: strong | Standard | 27.8B · 57.6 GB | Proof: strong | |
| ||||
Translation lane model, lite tierGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab) 25.2B (3.8B active)No proof yet | Lite | 25.2B (3.8B active) | No proof yet | |
| ||||
Translation lane model, best tierDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab) 284B (13B active) · 192 GBNo proof yet | Best | 284B (13B active) · 192 GB | No proof yet | |
| ||||
Translation lane model, network-hosted (wanted)GLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab) about 170 GB (estimate)No proof yet | Wanted | about 170 GB (estimate) | No proof yet | |
| ||||
Translation lane model, network-hosted (wanted)DeepSeek-V4.1-Flashdeepseek-ai/DeepSeek-V4.1-Flash on Hugging Face (opens in a new tab) about 476 GB (estimate)No proof yet | Wanted | about 476 GB (estimate) | No proof yet | |
| ||||
Specialist translation model for the translation laneHy-MT2-7B / Hy-MT2-30B-A3B-FP8tencent/Hy-MT2-7B on Hugging Face (opens in a new tab) 7B / 30B (3B active) · about 31 GB (estimate)No proof yetSelf-host only | Alternate | 7B / 30B (3B active) · about 31 GB (estimate) | No proof yetSelf-host only | |
| ||||
${DECOSA_REGISTRY}/decosa-api:0.1.0Sessions, WS /ws/live?vertical=translate&lang=<src>&target=<dst>, SSE /demo/replay, the sentence splitter and translate lanes, receipt lookup. No GPU. Binds 127.0.0.1 by default.
${DECOSA_REGISTRY}/decosa-llm:0.1.0Qwen3.8-27B on vLLM 0.29.0, served as qwen3.8-27b. Internal to the compose network.
${DECOSA_REGISTRY}/decosa-asr:0.1.0Voxtral Mini 4B Realtime on vLLM 0.27.1, served as voxtral-realtime at /v1/realtime. Internal to the compose network.
Default compose split: LLM 0.60, ASR 0.25 of the card, about 4 concurrent sessions. The latencies below were measured on our server with the two models on separate RTX PRO 6000 cards; the one-card split has not been measured.
No NVFP4 on Hopper: use Qwen/Qwen3.8-27B-FP8 (Apache-2.0), LLM_GPU_UTIL=0.62, ASR_GPU_UTIL=0.25. Starting point, not measured.
FP8 checkpoint, LLM_MAX_LEN=16384, LLM_GPU_UTIL=0.70, ASR_GPU_UTIL=0.22, DECOSA_LIVE_CAP=2. Starting point, not measured.
Both models do not fit with useful KV cache on one card.
Measuredmeasured on our server 2026-09-23 (replay of translate-en-es-tour, one session, gateway route; 1,600 ms on translate-es-en-repair; ops/record-all.json)
Measuredmeasured on our server 2026-09-23 (median per sentence after it ends, translate-en-es-tour, max 1,272 ms; median 1,459 ms / max 2,235 ms on translate-es-en-repair; ops/record-all.json)
Measuredmeasured on our server 2026-09-23 (median, translate-en-es-tour; 2,712 ms on translate-es-en-repair; ops/record-all.json)
Measuredmeasured on our server 2026-09-23 (after stop, translate-en-es-tour; 2,961 ms on translate-es-en-repair; ops/record-all.json)
Measuredmeasured on our server 2026-09-23 (median while one session per live vertical ran in parallel, ops/smoke-parallel-1.json; a second parallel run gave 851 ms, ops/smoke-parallel-2.json). Shows the worst case seen, not the typical one.
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa live-translation stack on this machine
You are setting up a self-hosted live-translation service: speech in, translated captions out, one sentence at a time, plus a running glossary and an end summary. Work step by step, show me each command before running anything that needs sudo, and stop and ask if a check fails. Not for court or medical interpreting.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/translate.zip (634 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py translate` (the api image carries the same bundle under /app/rehearsal/translate/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py translate --bundle translate.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the Spanish captions carry the sink and Thursday"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. What you are building
- `llm`: Qwen3.8-27B, `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80` (Apache-2.0), on vLLM 0.29.0 (`vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`), NVFP4 + MTP (3 draft tokens) + FP8 KV cache. Needs a Blackwell GPU; on Hopper/Ada use `Qwen/Qwen3.8-27B-FP8` with `LLM_REVISION=main`.
- `asr`: Voxtral Mini 4B Realtime, `mistralai/Voxtral-Mini-4B-Realtime-2602` @ `2769294da9567371363522aac9bbcfdd19447add` (Apache-2.0, BF16), on `vllm/vllm-openai:v0.27.1` + `mistral-common[audio]`, realtime WebSocket `/v1/realtime`.
- `api`: decosa-api on `127.0.0.1:8445`: sessions, `WS /ws/live`, replay, the `translate` lanes (`translation`, `glossary`, `summary`). No GPU.
- Source (speech) languages the ASR supports: ar, de, en, es, fr, hi, it, ja, ko, nl, pt, ru, zh. Target languages tested end to end: en, es. Other targets the API accepts (fr, de, it, pt, nl, pl, ru, uk, tr, ar, hi, zh, ja, ko, vi, id, sv) are untested.
## 1. Check the machine
1. `nvidia-smi`: need one NVIDIA GPU with ≥48 GB (96 GB recommended) and driver 580+. Record the GPU name and memory.
2. `docker --version` and `docker compose version`. If Docker is missing, install Docker Engine from docs.docker.com for this distro.
3. `docker run --rm --gpus all ubuntu nvidia-smi`. If it fails, install the NVIDIA Container Toolkit (`nvidia-container-toolkit`), run `sudo nvidia-ctk runtime configure --runtime=docker` and `sudo systemctl restart docker`, then retry.
4. `df -h`: need about 80 GB free for images and weights (~25 GB of weights download on first start).
5. Choose settings for `.env` from the GPU:
- RTX PRO 6000 Blackwell 96 GB / B200: defaults (`LLM_GPU_UTIL=0.60`, `ASR_GPU_UTIL=0.25`).
- H100/H200: `LLM_MODEL=Qwen/Qwen3.8-27B-FP8 LLM_REVISION=main LLM_GPU_UTIL=0.62`.
- 48 GB (L40S, RTX 6000 Ada): FP8 as above plus `LLM_MAX_LEN=16384 LLM_GPU_UTIL=0.70 ASR_GPU_UTIL=0.22 DECOSA_LIVE_CAP=2`.
Only the first row has been measured; the others are starting points. The two GPU shares must add up to less than 0.9.
## 2. Images
The images `${DECOSA_REGISTRY}/decosa-{api,asr,llm}:0.1.0` are **publishing soon**. Try `docker pull ${DECOSA_REGISTRY}/decosa-api:0.1.0`. If it fails (403/404), build from source instead: from a checkout of the decosa-api repository (source release also publishing soon) run `docker compose build` in its root; it builds the same three images from `docker/`. Don't substitute other images or model ids.
## 3. Write `~/decosa/docker-compose.yml`
```yaml
name: decosa
x-gpu: &gpu
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:0.1.0
<<: *gpu
ipc: host
restart: unless-stopped
environment: { HF_TOKEN: "${HF_TOKEN:-}" }
volumes: [ "hf-cache:/root/.cache/huggingface" ]
command: ["${LLM_MODEL:-nvidia/Qwen3.8-27B-NVFP4}", "--revision", "${LLM_REVISION:-482ca0f3832238542f8f5295dde86b5f22711d80}",
"--served-model-name", "qwen3.8-27b", "--language-model-only", "--max-model-len", "${LLM_MAX_LEN:-65536}",
"--gpu-memory-utilization", "${LLM_GPU_UTIL:-0.60}", "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3",
"--speculative-config", '{"method":"mtp","num_speculative_tokens":3}', "--seed", "0", "--enable-force-include-usage",
"--host", "0.0.0.0", "--port", "8000"]
healthcheck: { test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], interval: 15s, timeout: 5s, retries: 5, start_period: 900s }
asr:
image: ${DECOSA_REGISTRY}/decosa-asr:0.1.0
<<: *gpu
ipc: host
restart: unless-stopped
depends_on: { llm: { condition: service_healthy } }
environment: { HF_TOKEN: "${HF_TOKEN:-}" }
volumes: [ "hf-cache:/root/.cache/huggingface" ]
command: ["--model", "mistralai/Voxtral-Mini-4B-Realtime-2602", "--revision", "2769294da9567371363522aac9bbcfdd19447add",
"--tokenizer-mode", "mistral", "--config-format", "mistral", "--load-format", "mistral",
"--compilation-config", '{"cudagraph_mode":"PIECEWISE"}', "--max-model-len", "45000", "--max-num-batched-tokens", "8192",
"--max-num-seqs", "16", "--gpu-memory-utilization", "${ASR_GPU_UTIL:-0.25}", "--served-model-name", "voxtral-realtime",
"--host", "0.0.0.0", "--port", "8000"]
healthcheck: { test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], interval: 15s, timeout: 5s, retries: 5, start_period: 600s }
api:
image: ${DECOSA_REGISTRY}/decosa-api:0.1.0
restart: unless-stopped
depends_on: { llm: { condition: service_healthy }, asr: { condition: service_healthy } }
environment:
DECOSA_ASR_WS: ws://asr:8000/v1/realtime
DECOSA_LLM_ROUTE: direct # local model; receipts are signed by this box's key ("attested")
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LIVE_CAP: "${DECOSA_LIVE_CAP:-4}"
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_AUDIO_S: "3600"
DECOSA_BUDGET_LLM_TOKENS: "200000"
DECOSA_SESSION_TTL_S: "28800"
DECOSA_STUDIO_WORKER: none
ports: [ "127.0.0.1:8445:8445" ]
volumes: [ "decosa-data:/data" ]
volumes: { hf-cache: {}, decosa-data: {} }
```
Then `cd ~/decosa && docker compose up -d`. The LLM takes 5–10 minutes to become healthy on first start; the ASR starts after it (1–3 minutes). Wait until `docker compose ps` shows all three healthy and `curl -s localhost:8445/healthz` returns `"ok": true, "asr": true, "llm": true`.
## 4. Smoke test (no microphone needed)
```bash
curl -s localhost:8445/demo/scripts | jq '.[] | select(.vertical=="translate")' # expect translate-es-en-repair and translate-en-es-tour
TOKEN=$(curl -s localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"translate"}' | jq -r .token)
curl -sN localhost:8445/demo/replay -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
-d '{"vertical":"translate","script_id":"translate-es-en-repair"}' | tee /tmp/translate-smoke.sse
```
Pass if the SSE stream contains, in order: a `ready` event with `"lang":"es","target":"en"`; `transcript` events (`final:false` partials, then `final:true` sentences with `i` and `t`); about 14 `lane` events with `"lane":"translation"` whose `data` is `{i, t, source, translation, target:"en"}` and whose text is English; at least one `glossary` lane; one `summary` lane after the audio ends; `receipt` events with `"status": "attested"` (signed by the box's own key on the direct route); and a final `done` whose `summary` has `translation`, `glossary`, `summary`, `source_language` and `target_language`. On our server a translated sentence arrived about 1.1–1.5 s after it ended (median, one session); the one-GPU split may differ. Report what you measure.
## 5. Point an app at the local API
- Base URL `http://localhost:8445`. Get a token with `POST /demo/session {"vertical":"translate"}`, then open `ws://localhost:8445/ws/live?vertical=translate&token=<t>&lang=<src|auto>&target=<dst>`.
- Send binary frames of 16 kHz mono PCM16 little-endian, ~100 ms each, then the text frame `{"type":"stop"}`.
- `translation` events are one per sentence and may arrive out of order: keep them in a map keyed by `data.i`. For `glossary` and `summary`, the latest event is the full state. Close codes: `4400` bad `lang`/`target`, `4429` busy, `4402` budget used up.
- Browsers are allowed from `localhost`/`127.0.0.1` on any port by default; add other origins with `DECOSA_CORS_ORIGINS`.
Captions overlay: save as `captions.html`, serve with `python3 -m http.server 5173` (not `file://`), open `http://localhost:5173/captions.html?lang=es&target=en`:
```html
<!doctype html><meta charset="utf-8"><title>Captions</title>
<style>body{margin:0;background:transparent;font:28px/1.3 system-ui}#cap{position:fixed;left:5%;right:5%;bottom:4%;color:#fff;background:#000b;padding:12px 18px;border-radius:8px}#src{font-size:18px;opacity:.7}</style>
<div id="cap"><div id="lines"></div><div id="src"></div></div><button id="go">Start</button>
<script>
const API = "http://localhost:8445", q = new URLSearchParams(location.search);
const lang = q.get("lang") || "", target = q.get("target") || "en", byI = new Map();
go.onclick = async () => {
go.remove();
const { token } = await (await fetch(API + "/demo/session", { method: "POST", headers: { "content-type": "application/json" }, body: JSON.stringify({ vertical: "translate" }) })).json();
const ws = new WebSocket(`${API.replace("http", "ws")}/ws/live?vertical=translate&token=${token}${lang ? `&lang=${lang}` : ""}&target=${target}`);
ws.binaryType = "arraybuffer";
ws.onmessage = (m) => {
const ev = JSON.parse(m.data);
if (ev.type === "transcript") src.textContent = ev.final ? "" : ev.text; // live source text while a sentence is open
if (ev.type === "lane" && ev.lane === "translation") { // one per sentence; order by data.i
byI.set(ev.data.i, ev.data.translation);
lines.innerHTML = [...byI.keys()].sort((a, b) => a - b).slice(-2).map((i) => `<div>${byI.get(i).replace(/</g, "<")}</div>`).join("");
}
if (ev.type === "error") console.warn(ev.message);
};
const stream = await navigator.mediaDevices.getUserMedia({ audio: { channelCount: 1, echoCancellation: true } });
const ctx = new AudioContext({ sampleRate: 16000 }), node = ctx.createScriptProcessor(2048, 1, 1); // 2048 samples = 128 ms (must be a power of two)
node.onaudioprocess = (e) => {
if (ws.readyState !== 1) return;
const f = e.inputBuffer.getChannelData(0), pcm = new Int16Array(f.length);
for (let k = 0; k < f.length; k++) pcm[k] = Math.max(-1, Math.min(1, f[k])) * 0x7fff;
ws.send(pcm.buffer);
};
ctx.createMediaStreamSource(stream).connect(node); node.connect(ctx.destination);
addEventListener("keydown", (e) => { if (e.key === "Escape") ws.send(JSON.stringify({ type: "stop" })); }); // Esc = stop, summary follows
};
</script>
```
The page works as an OBS browser source over video (transparent background).
Finish by printing: GPU and settings chosen, `docker compose ps`, `/healthz`, and the smoke-test result with the measured translation latency.Last reviewed
Decosa live translation streams speech through open-weight speech recognition and translates each sentence as soon as it ends, so captions trail the speaker. On the hosted route each translation arrived about 0.6 s (median) after its sentence. About every 20 seconds it rebuilds a glossary of names and domain terms to keep later translations consistent, and on stop it writes a short summary in the target language. Output is text captions, not synthesized speech.
On the public FLORES-200 devtest set (1,012 sentences per direction, held out), Decosa live translation scored chrF++ / BLEU / COMET-22 of 55.3 / 29.9 / 87.2 for English to Spanish and 59.7 / 31.8 / 87.7 for Spanish to English. These are automatic metrics on clean text, not speech-recogniser output, so live-speech quality will be lower. Only the standard tier was measured, and there is no human rating.
No. Decosa live translation is machine translation, not a certified interpreter, and is out of scope for court, legal and medical interpreting or any setting where a mistranslation carries legal or clinical risk. It is meant for tours, meetings, customer calls and events where a caption track helps. Get consent before you record or transcribe anyone.
Yes. Decosa live translation runs self-hosted, and with DECOSA_LLM_ROUTE=direct nothing leaves your machine; the hosted demo keeps transcripts in memory only and deletes them when the session ends. The standard tier uses one 96 GB card. A lite tier on a 48 GB card uses a smaller model, has no measured quality and no signed receipts.
Decosa live translation accepts speech in ar, de, en, es, fr, hi, it, ja, ko, nl, pt, ru and zh, or auto-detect. Only English and Spanish targets are tested end to end; other targets are accepted but untested. FLORES-200 scores are also published for English to French, German and Chinese. A 56 s Spanish repair call cost about $0.0025 at the gateway list price.
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…