{"schema_version":"1","site":"https://decosa.ai","id":"translate","num":"06","name":"Live translation","tool_name":"Translate a talk live","short":"Translate","blurb":"Speak in one language and read it in another, sentence by sentence, with a running glossary and a summary at the end.","status":"live","labels":{"industry":["general"],"job":["translate"],"input":["voice"],"deploy":["hosted","selfhost"],"status":"live","output":["text"],"data":["pii"],"hardware":"gpu-96","licence":"permissive"},"industries":["general"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["live-asr"],"models":"Voxtral 4B · Qwen3.8-27B","where":"Hosted or self-host","hardware":"1× RTX PRO 6000 (96 GB), or 2× RTX 5090","final_artifact":"The translation and a summary.","self_host_first":false,"verification":{"receipt_coverage":"full","summary":"Strong on text lanes","manual_qa":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":2300,"p95_ms":null,"runs":null,"receipts_per_run":18,"cost_per_run_usd":0.0025},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers","notes":"The step 4 replay passed as written (14 translation lanes, glossary, summary, 18 attested receipts; a translation 157 ms median after its sentence), and the step 5 captions overlay translated live fake-microphone audio in headless Chromium once its createScriptProcessor buffer size was fixed. Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified."},"known_limits":["Only English and Spanish targets are tested; other targets are accepted but untested.","Before the merge, the hosted API refused lang=auto (the console's default), so Start recording on the translate page failed; samples worked."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"FLORES-200 devtest en→es: chrF++ / BLEU / COMET-22","value":"55.3 / 29.9 / 87.2","unit":null,"n":1012,"split":"heldout","note":null},{"name":"FLORES-200 devtest es→en: chrF++ / BLEU / COMET-22","value":"59.7 / 31.8 / 87.7","unit":null,"n":1012,"split":"heldout","note":null},{"name":"FLORES-200 devtest en→fr: chrF++ / BLEU / COMET-22","value":"68.8 / 49.1 / 88.9","unit":null,"n":1012,"split":"heldout","note":null},{"name":"FLORES-200 devtest en→de: chrF++ / BLEU / COMET-22","value":"62.8 / 38.2 / 88.6","unit":null,"n":1012,"split":"heldout","note":null},{"name":"FLORES-200 devtest en→zh: chrF++ / BLEU / COMET-22","value":"29.7 / 45.7 / 89.5","unit":null,"n":1012,"split":"heldout","note":"zh BLEU with the zh tokenizer; chrF++ understates unsegmented Chinese"}],"dataset":"FLORES-200 devtest (public, 1,012 sentences per direction) through the live NVFP4 endpoint with decosa-api's translate prompt.","held_out":true,"caveats":["Clean text rather than speech-recogniser output, so live-speech quality will be lower.","Standard tier only; lite, best and wanted tiers not measured yet.","Automatic metrics only; no human rating."],"date":"2026-09-23","doc_url":null},"stack":{"summary":"Streams speech through open-weight speech recognition and translates each sentence as soon as it ends, so captions trail the speaker by about a second. Every 20 seconds or so it rebuilds a glossary of names and domain terms and feeds it back into later translations to keep them consistent. When you stop, it writes a short summary in the target language. It is for tours, meetings, customer calls and events where a caption track helps. It is not an interpreter for court, medical or other settings where a mistranslation carries legal or clinical risk.","tagline":"Speech in one language, captions in another, one sentence at a time, with a running glossary and a summary at the end.","deployment":"hosted-or-self-host","regulatory_note":"Machine translation, not a certified interpreter: out of scope for court, legal and medical interpreting. Get consent before you record or transcribe anyone. The hosted demo keeps transcripts in memory only and deletes them when the session ends; for confidential conversations, self-host with DECOSA_LLM_ROUTE=direct so audio and text stay on your machine.","components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","hf_repo":"mistralai/Voxtral-Mini-4B-Realtime-2602","license":"Apache-2.0","params":"4.4B","quant":"BF16","vram_gb":24,"memory_gb_estimate":null,"engine":"vLLM realtime WebSocket (/v1/realtime), 480 ms transcription delay","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard","best"],"alternative_to":null},{"id":"llm","role":"Translation lane model (translation, glossary, summary)","name":"Qwen3.8-27B (NVIDIA NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP) + FP8 (attention/GDN), FP8 KV cache, MTP speculative decoding (3 draft tokens)","vram_gb":57.6,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 (OpenAI-compatible)","receipt_coverage":"strong","in_hosted_demo":null,"tiers":["standard"],"alternative_to":null},{"id":"llm-lite","role":"Translation lane model, lite tier","name":"Gemma 4 26B A4B (instruction-tuned)","hf_repo":"google/gemma-4-26B-A4B-it","license":"Apache-2.0 (model card also links the Gemma 4 licence page)","params":"25.2B","quant":"BF16 weights; FP8 at load time (vLLM --quantization fp8) to fit a 48 GB card","vram_gb":null,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 (Gemma4ForConditionalGeneration is in its model registry)","receipt_coverage":"none","in_hosted_demo":null,"tiers":["lite"],"alternative_to":null},{"id":"llm-best","role":"Translation lane model, best tier","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","hf_repo":"nvidia/DeepSeek-V4-Flash-NVFP4","license":"MIT","params":"284B","quant":"NVFP4 experts (MXFP4 MTP draft), MTP speculative decoding","vram_gb":192,"memory_gb_estimate":null,"engine":"vLLM B12X native SM120 build, TP2 across two cards, with the owner's two MTP patches (dsv4-flash-nvfp4-sm120 kit)","receipt_coverage":"none","in_hosted_demo":null,"tiers":["best"],"alternative_to":null},{"id":"llm-network-glm","role":"Translation lane model, network-hosted (wanted)","name":"GLM-5.3-Flash","hf_repo":"zai-org/GLM-5.3-Flash","license":"MIT","params":null,"quant":null,"vram_gb":null,"memory_gb_estimate":170,"engine":null,"receipt_coverage":"none","in_hosted_demo":null,"tiers":["wanted"],"alternative_to":null},{"id":"llm-network-dsv41","role":"Translation lane model, network-hosted (wanted)","name":"DeepSeek-V4.1-Flash","hf_repo":"deepseek-ai/DeepSeek-V4.1-Flash","license":"MIT","params":null,"quant":"FP8 (block scale)","vram_gb":null,"memory_gb_estimate":476,"engine":null,"receipt_coverage":"none","in_hosted_demo":null,"tiers":["wanted"],"alternative_to":null},{"id":"mt-hy","role":"Specialist translation model for the translation lane","name":"Hy-MT2-7B / Hy-MT2-30B-A3B-FP8","hf_repo":"tencent/Hy-MT2-7B","license":"Apache-2.0 (no territory clause)","params":"7B / 30B (3B active)","quant":"BF16 (16.1 GB) / FP8 (30.6 GB)","vram_gb":null,"memory_gb_estimate":31,"engine":"vLLM","receipt_coverage":"none","in_hosted_demo":false,"tiers":["alternates"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · one 48 GB card","summary":"A smaller mixture-of-experts model on cheaper hardware. Translations are likely less polished and no quality has been measured; no signed receipts.","components":["asr","llm-lite"],"hardware":"1× L40S / RTX 6000 Ada 48 GB (or any 48 GB+ card)","quality_evidence":[{"metric":"Translation quality","value":"not measured yet","source":null}],"latency_note":"estimate: not run in this stack yet. 3.8B active parameters should decode faster than the 27B dense model. Fitting Voxtral (8.3 GB BF16 weights) and the LLM on one 32 GB card is not verified, so this tier starts at 48 GB.","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Not a hosted model yet; no receipts today (the speech recognition step never emits one). Becomes strong once the LLM is registered as a hosted model with gateway receipts.","hosting":null},{"id":"standard","label":"Standard · one 96 GB card (hosted demo)","summary":"The stack the hosted demo runs: Qwen3.8-27B NVFP4 with MTP. Every lane call can carry a gateway-signed receipt.","components":["asr","llm"],"hardware":"1× RTX PRO 6000 Blackwell 96 GB","quality_evidence":[{"metric":"FLORES-200 devtest, en→es / es→en (n = 1,012 sentences each): chrF++ / BLEU / COMET-22","value":"55.3 / 29.9 / 87.2 and 59.7 / 31.8 / 87.7","source":"; eval results file translation-flores200-qwen38-27b.json (live NVFP4 endpoint, decosa-api's translate prompt, clean text rather than speech-recogniser output, 2026-09-23)"},{"metric":"FLORES-200 devtest, en→fr / en→de / en→zh (n = 1,012 each): chrF++ / BLEU / COMET-22","value":"68.8 / 49.1 / 88.9; 62.8 / 38.2 / 88.6; 29.7 / 45.7 / 89.5 (zh BLEU with the zh tokenizer; chrF++ understates unsegmented Chinese)","source":"; eval results file translation-flores200-qwen38-27b.json (live NVFP4 endpoint, decosa-api's translate prompt, clean text rather than speech-recogniser output, 2026-09-23)"},{"metric":"Replay completeness, es→en and en→es","value":"28/28 sentences translated, 36/36 receipts signed","source":"decosa-api ops/record-all.json, our server 2026-09-23 (functional check, not a quality score)"},{"metric":"Engine quality smoke (math / code)","value":"38/40 and 12/12","source":"Decosa model benchmarks (Sep 2026) (pinned NVFP4+MTP3+FP8 KV stack)"}],"latency_note":"measured on our server: each translated sentence and the first transcript arrive about a second after speech (one session, models on two separate cards).","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":null,"hosting":null},{"id":"best","label":"Best · two 96 GB cards","summary":"DeepSeek-V4-Flash (284B MoE) on two cards: the strongest model measured on the owner's own benchmarks, at twice the hardware and without signed receipts.","components":["asr","llm-best"],"hardware":"2× RTX PRO 6000 Blackwell 96 GB (TP2)","quality_evidence":[{"metric":"Translation quality","value":"not measured yet","source":null},{"metric":"Clinical note ROUGE-L on ACI-Bench (not translation; for relative ranking only)","value":"35.8 vs 34.2 for Qwen3.8-27B","source":"scribe-bench wiki/models.md (V4-Flash measured on a Mac Studio build, not this NVFP4 build)"}],"latency_note":"measured decode, not lane latency: 150.6 tok/s single-stream with MTP, about 328 tok/s aggregate at 4 streams (dsv4-flash-nvfp4-sm120 README, 2× RTX PRO 6000). Per-sentence translation latency in this stack not measured.","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Not a hosted model yet; no receipts today (the speech recognition step never emits one). Becomes strong once the LLM is registered as a hosted model with gateway receipts.","hosting":null},{"id":"wanted","label":"Wanted · the largest open flash models","summary":"GLM-5.3-Flash and DeepSeek-V4.1-Flash translate, served by network providers; speech recognition stays on your machine. Not served yet.","components":["asr","llm-network-glm","llm-network-dsv41"],"hardware":"Network providers: an 8x H200-class node for DeepSeek-V4.1-Flash (476 GB of weights); 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash (about 170 GB). Speech recognition stays local. Estimate.","quality_evidence":[{"metric":"Translation quality","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Not hosted yet, so no receipts today.","hosting":"network"}],"alternates":[{"id":"specialist-mt","label":"Specialist translation models","components":["mt-hy"],"hardware":"1x 24 GB card (7B) or 1x 48 GB card (30B-A3B FP8)","use":"Hy-MT2-7B runs today in the language pack (served on Decosa's hosted service; this tool's hosted demo doesn't use it yet); Hy-MT2-30B-A3B would be a one-card best tier (page 32 #8) and is not served. Re-run our FLORES COMET-22 harness before any quality claim.","status":"runs today"}],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:0.1.0","purpose":"Sessions, WS /ws/live?vertical=translate&lang=<src>&target=<dst>, SSE /demo/replay, the sentence splitter and translate lanes, receipt lookup. No GPU. Binds 127.0.0.1 by default."},{"name":"decosa-llm","port":8000,"image":"${DECOSA_REGISTRY}/decosa-llm:0.1.0","purpose":"Qwen3.8-27B on vLLM 0.29.0, served as qwen3.8-27b. Internal to the compose network."},{"name":"decosa-asr","port":8000,"image":"${DECOSA_REGISTRY}/decosa-asr:0.1.0","purpose":"Voxtral Mini 4B Realtime on vLLM 0.27.1, served as voxtral-realtime at /v1/realtime. Internal to the compose network."}],"tools":[],"hardware":[{"tier":"1× RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Default compose split: LLM 0.60, ASR 0.25 of the card, about 4 concurrent sessions. The latencies below were measured on our server with the two models on separate RTX PRO 6000 cards; the one-card split has not been measured."},{"tier":"1× H100 / H200 (80–141 GB, Hopper)","fits":true,"notes":"No NVFP4 on Hopper: use Qwen/Qwen3.8-27B-FP8 (Apache-2.0), LLM_GPU_UTIL=0.62, ASR_GPU_UTIL=0.25. Starting point, not measured."},{"tier":"1× L40S / RTX 6000 Ada 48 GB","fits":true,"notes":"FP8 checkpoint, LLM_MAX_LEN=16384, LLM_GPU_UTIL=0.70, ASR_GPU_UTIL=0.22, DECOSA_LIVE_CAP=2. Starting point, not measured."},{"tier":"Cards under 48 GB","fits":false,"notes":"Both models do not fit with useful KV cache on one card."}],"latency":[{"lane":"transcript (first text)","typical_ms":1200,"source":"measured on our server 2026-09-23 (replay of translate-en-es-tour, one session, gateway route; 1,600 ms on translate-es-en-repair; ops/record-all.json)"},{"lane":"translation","typical_ms":1119,"source":"measured on our server 2026-09-23 (median per sentence after it ends, translate-en-es-tour, max 1,272 ms; median 1,459 ms / max 2,235 ms on translate-es-en-repair; ops/record-all.json)"},{"lane":"glossary","typical_ms":2152,"source":"measured on our server 2026-09-23 (median, translate-en-es-tour; 2,712 ms on translate-es-en-repair; ops/record-all.json)"},{"lane":"summary","typical_ms":3103,"source":"measured on our server 2026-09-23 (after stop, translate-en-es-tour; 2,961 ms on translate-es-en-repair; ops/record-all.json)"},{"lane":"translation (under load)","typical_ms":15122,"source":"measured on our server 2026-09-23 (median while one session per live vertical ran in parallel, ops/smoke-parallel-1.json; a second parallel run gave 851 ms, ops/smoke-parallel-2.json). Shows the worst case seen, not the typical one."}],"benchmark":null,"notes":[]},"buyer_facts":[{"label":"Measured latency","value":"The hosted median below is the time from the end of the audio to the final `done`, over replays of the sample (gateway route, one session at a time). Each sentence's translation arrives about a second after the sentence ends."},{"label":"Typical run cost","value":"A fraction of a cent for the Spanish repair call (under two dozen model calls) at the gateway list price."},{"label":"Data retention","value":"Nothing stored: transcript and translations live in memory for the session."},{"label":"What leaves the box (self-host)","value":"Nothing, with DECOSA_LLM_ROUTE=direct."},{"label":"Languages","value":"Speech: ar, de, en, es, fr, hi, it, ja, ko, nl, pt, ru, zh or auto. Targets tested end to end: en, es."}],"data_handling":{"page":"/data#translate","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Nothing stored: transcript and translations live in memory for the session.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/tools/operations/translate","input":"mic","lanes":[{"id":"translation","title":"Translation","kind":"list"},{"id":"glossary","title":"Glossary","kind":"list"},{"id":"summary","title":"Summary","kind":"markdown"}],"samples":[{"n":1,"id":"translate-es-en-repair","title":"Repair visit, Spanish to English","deep_link":"/tools/operations/translate?sample=1&autorun=0"},{"n":2,"id":"translate-en-es-tour","title":"Tour, English to Spanish","deep_link":"/tools/operations/translate?sample=2&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/translate-hosted.md","selfhost":"/prompts/translate-selfhost.md","assemble":"/prompts/translate-assemble.md","mac":"/prompts/translate-mac.md"},"rehearsal":{"bundle":"/samples/translate.zip","bundle_url":"https://decosa.ai/samples/translate.zip","folder":"/samples/translate/","expected":"/samples/translate/expected.json","files":["/samples/translate/expected.json","/samples/translate/inputs/repair-call-es-28s.wav","/samples/translate/inputs/repair-call-es-script.json"],"bytes":649621,"checks":["the session ends with a done event","the session reports no errors","the Spanish captions carry the sink and Thursday","translations arrive in English (target en)","at least 4 sentences are translated","the translations carry the leak and Thursday","the English summary keeps the day, the 45-dollar price and the street (Olmo)","every speech and model call has a signed receipt"],"licence":"Synthetic: a script written for Decosa (no real people or companies) read by Decosa house voices (Kokoro-82M stock voicepacks, Apache-2.0), each allowed by the consent ledger for project decosa-translate-demo. Part of decosa-api, AGPL-3.0-or-later.","about":"Two cuts from a synthetic Spanish phone call booking a plumber (a leak under the kitchen sink; a technician on Thursday between 8 and 10; the visit costs 45 dollars; the address), streamed over the live WebSocket with lang=es and target=en. Every Spanish sentence must come back translated into English as it is said, and the call must end with an English summary that keeps the day, the price and the address.","run":{"containers":"docker compose exec api python scripts/rehearse.py translate","checkout":"python scripts/rehearse.py translate --bundle translate.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py translate"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=translate","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":64,"basis":"estimate","unknown":[]},{"id":"standard","gpu_gb":81.6,"basis":"stack","unknown":[]},{"id":"best","gpu_gb":216,"basis":"stack","unknown":[]},{"id":"wanted","gpu_gb":1344,"basis":"estimate","unknown":[]},{"id":"alternate-specialist-mt","gpu_gb":9.4,"basis":"estimate","unknown":[]}],"mac":{"fit":"full","memory_gb":48}},"links":{"page":"/tools/operations/translate","json":"/use-cases/translate.json","metrics":"/metrics/translate","console":"/tools/operations/translate","console_sample":"/tools/operations/translate?sample=1&autorun=0","stack":"/tools/operations/translate#stack","try_live":"/tools/operations/translate","watch":"/tools/operations/translate","build":"/tools/operations/translate#build","self_host":"/tools/operations/translate#self-host","prompts":{"hosted":"/prompts/translate-hosted.md","selfhost":"/prompts/translate-selfhost.md","assemble":"/prompts/translate-assemble.md","mac":"/prompts/translate-mac.md"}}}