Skip to content
decosa

06 · Any industry · live

Live translation

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 23 Sep 2026

  • FLORES-200 devtest en→es: chrF++ / BLEU / COMET-2255.3 / 29.9 / 87.2held outn = 1,012
  • FLORES-200 devtest es→en: chrF++ / BLEU / COMET-2259.7 / 31.8 / 87.7held outn = 1,012
  • FLORES-200 devtest en→fr: chrF++ / BLEU / COMET-2268.8 / 49.1 / 88.9held outn = 1,012
  • FLORES-200 devtest en→de: chrF++ / BLEU / COMET-2262.8 / 38.2 / 88.6held outn = 1,012
  • FLORES-200 devtest en→zh: chrF++ / BLEU / COMET-2229.7 / 45.7 / 89.5held outn = 1,012zh BLEU with the zh tokenizer; chrF++ understates unsegmented Chinese

Dataset

FLORES-200 devtest (public, 1,012 sentences per direction) through the live NVFP4 endpoint with decosa-api's translate prompt.

Caveats

  • Clean text rather than speech-recogniser output, so live-speech quality will be lower.
  • Standard tier only; lite, best and wanted tiers not measured yet.
  • Automatic metrics only; no human rating.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
2.3 s
Receipts
18
Model calls
n/a
Tokens
n/a
Cost per run
$0.003

Self-host verification

Verified on 25 Sep 2026: fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers

The step 4 replay passed as written (14 translation lanes, glossary, summary, 18 attested receipts; a translation 157 ms median after its sentence), and the step 5 captions overlay translated live fake-microphone audio in headless Chromium once its createScriptProcessor buffer size was fixed. Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified.

Rehearsal bundle: translate.zip (634 KB, 8 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Only English and Spanish targets are tested; other targets are accepted but untested.
  • Before the merge, the hosted API refused lang=auto (the console's default), so Start recording on the translate page failed; samples worked.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 48 GB card (1)
  • Translation quality: not measured yet
Standard · one 96 GB card (hosted demo) (4)
  • FLORES-200 devtest, en→es / es→en (n = 1,012 sentences each): chrF++ / BLEU / COMET-22: 55.3 / 29.9 / 87.2 and 59.7 / 31.8 / 87.7; eval results file translation-flores200-qwen38-27b.json (live NVFP4 endpoint, decosa-api's translate prompt, clean text rather than speech-recogniser output, 2026-09-23)
  • FLORES-200 devtest, en→fr / en→de / en→zh (n = 1,012 each): chrF++ / BLEU / COMET-22: 68.8 / 49.1 / 88.9; 62.8 / 38.2 / 88.6; 29.7 / 45.7 / 89.5 (zh BLEU with the zh tokenizer; chrF++ understates unsegmented Chinese); eval results file translation-flores200-qwen38-27b.json (live NVFP4 endpoint, decosa-api's translate prompt, clean text rather than speech-recogniser output, 2026-09-23)
  • Replay completeness, es→en and en→es: 28/28 sentences translated, 36/36 receipts signeddecosa-api ops/record-all.json, our server 2026-09-23 (functional check, not a quality score)
  • Engine quality smoke (math / code): 38/40 and 12/12Decosa model benchmarks (Sep 2026) (pinned NVFP4+MTP3+FP8 KV stack)
Best · two 96 GB cards (2)
  • Translation quality: not measured yet
  • Clinical note ROUGE-L on ACI-Bench (not translation; for relative ranking only): 35.8 vs 34.2 for Qwen3.8-27Bscribe-bench wiki/models.md (V4-Flash measured on a Mac Studio build, not this NVFP4 build)
Wanted · the largest open flash models (1)
  • Translation quality: not measured yet

How we measure · All tools