06 · Any industry · live
Live translation
Eval results
Scored on a held-out or test splitRun 23 Sep 2026
- FLORES-200 devtest en→es: chrF++ / BLEU / COMET-2255.3 / 29.9 / 87.2held outn = 1,012
- FLORES-200 devtest es→en: chrF++ / BLEU / COMET-2259.7 / 31.8 / 87.7held outn = 1,012
- FLORES-200 devtest en→fr: chrF++ / BLEU / COMET-2268.8 / 49.1 / 88.9held outn = 1,012
- FLORES-200 devtest en→de: chrF++ / BLEU / COMET-2262.8 / 38.2 / 88.6held outn = 1,012
- FLORES-200 devtest en→zh: chrF++ / BLEU / COMET-2229.7 / 45.7 / 89.5held outn = 1,012zh BLEU with the zh tokenizer; chrF++ understates unsegmented Chinese
Dataset
FLORES-200 devtest (public, 1,012 sentences per direction) through the live NVFP4 endpoint with decosa-api's translate prompt.
Caveats
- Clean text rather than speech-recogniser output, so live-speech quality will be lower.
- Standard tier only; lite, best and wanted tiers not measured yet.
- Automatic metrics only; no human rating.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 2.3 s
- Receipts
- 18
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.003
Self-host verification
Verified on 25 Sep 2026: fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers
The step 4 replay passed as written (14 translation lanes, glossary, summary, 18 attested receipts; a translation 157 ms median after its sentence), and the step 5 captions overlay translated live fake-microphone audio in headless Chromium once its createScriptProcessor buffer size was fixed. Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified.
Rehearsal bundle: translate.zip (634 KB, 8 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Only English and Spanish targets are tested; other targets are accepted but untested.
- Before the merge, the hosted API refused lang=auto (the console's default), so Start recording on the translate page failed; samples worked.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Speech recognition (streaming)Voxtral Mini 4B RealtimeApache-2.0
- Translation lane model (translation, glossary, summary)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 48 GB card (1)
- Translation quality: not measured yet
Standard · one 96 GB card (hosted demo) (4)
- FLORES-200 devtest, en→es / es→en (n = 1,012 sentences each): chrF++ / BLEU / COMET-22: 55.3 / 29.9 / 87.2 and 59.7 / 31.8 / 87.7; eval results file translation-flores200-qwen38-27b.json (live NVFP4 endpoint, decosa-api's translate prompt, clean text rather than speech-recogniser output, 2026-09-23)
- FLORES-200 devtest, en→fr / en→de / en→zh (n = 1,012 each): chrF++ / BLEU / COMET-22: 68.8 / 49.1 / 88.9; 62.8 / 38.2 / 88.6; 29.7 / 45.7 / 89.5 (zh BLEU with the zh tokenizer; chrF++ understates unsegmented Chinese); eval results file translation-flores200-qwen38-27b.json (live NVFP4 endpoint, decosa-api's translate prompt, clean text rather than speech-recogniser output, 2026-09-23)
- Replay completeness, es→en and en→es: 28/28 sentences translated, 36/36 receipts signeddecosa-api ops/record-all.json, our server 2026-09-23 (functional check, not a quality score)
- Engine quality smoke (math / code): 38/40 and 12/12Decosa model benchmarks (Sep 2026) (pinned NVFP4+MTP3+FP8 KV stack)
Best · two 96 GB cards (2)
- Translation quality: not measured yet
- Clinical note ROUGE-L on ACI-Bench (not translation; for relative ranking only): 35.8 vs 34.2 for Qwen3.8-27Bscribe-bench wiki/models.md (V4-Flash measured on a Mac Studio build, not this NVFP4 build)
Wanted · the largest open flash models (1)
- Translation quality: not measured yet