48 · Film, TV and games · Creative and media · live
Consented creator dubbing
Eval results
Not held outRun 26 Sep 2026Eval write-up (decosa-api, access required)
- Dub tracks exactly the video's length (sample count)11 of 11syntheticn = 11Plus 6 of 6 synthetic frame-rate edge cases exact
- Consent gate decisions as expected9 of 9syntheticn = 9Deterministic code; all 9 receipted
- ASR word error rate against the narration script3.9% (own-voice), 1.0% (dubber-voice)syntheticSynthetic, clean narration; real creators will score worse
- Glossary terms rendered as required98 of 98syntheticn = 98
- Back-translation chrF against the sourcemean 73.4 (range 70.0-76.0)syntheticA proxy for meaning kept, not a quality score
- Lines cut short to fitGPU runs 12 of 160; CPU run 3 of 18syntheticn = 178
- Planted subtitle errors caught793 of 793syntheticn = 79316 kinds x 25 tries; shows each check fires, not real-world prevalence. 0 false alarms on the clean hand-made file
Dataset
Two self-made narration videos (57 s and 43.2 s) of fictional creators with synthetic stock voices, 11 pipeline runs; a hand-written Spanish SDH file with planted errors; 6 synthetic length edge cases.
Caveats
- No human rating of the Spanish or of the voice: the numbers are measurable proxies.
- Synthetic stock voices and clean narration; real voices depend on the consent clip's quality.
- The subtitle checker and its planted errors have the same author.
- The voice keeps some English accent; a native Spanish speaker has not rated it. No lip-sync, no stem separation, one speaker only.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 26 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 74 s
- Receipts
- 43
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.003
Self-host verification
Verified on 26 Sep 2026: Fresh clone of the branch into a clean directory, docker build of the api image plus the documented voice-runtime layer, compose with named volumes, pointed at the already-running local Voxtral and Qwen3.8-27B (direct route), voice on CPU; then torn down.
The revoked sample was refused; the own-voice dub reached review in 337 s (including the first download of the voice weights) with an exact 2,736,000-sample track, glossary 10 of 10 and a render receipt; approval gave a C2PA-signed release (state Valid) and a record that verifies; no transcript or approver text in the container logs. Found on the way: chatterbox-tts pins numpy<1.26, which has no wheels for the image's Python 3.12, so the documented layer now builds the voice venv with Python 3.11 via uv.
Rehearsal bundle: consented-dubbing.zip (756 KB, 17 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- One speaker, English to Spanish; no lip-sync; music under the voice is not carried into the dub track.
- The voice keeps some English accent (cross-language cloning), and no native speaker has rated it yet: that is why approval is required.
- About 1 line in 13 is cut short to fit its slot (12 of 160 on GPU runs).
- Approval proves the job's key or session pressed Approve on the exact draft, not which person did.
- C2PA credentials use a development certificate, so public validators show the issuer as untrusted.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Transcribes each speech span of the source, and listens back to every dubbed lineVoxtral Mini 4B RealtimeApache-2.0
- Translation with the glossary and a length budget per line, repair of lines that break the glossary or budget, back-translation for the reviewerQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
- Speaks each line in the consented voice (the enrolled consent clip is the reference), with the Perth watermarkChatterbox Multilingual (t3_23lang)MIT
- Consent ledger's speaker check: the reference clip must match the enrolled voiceprint (tool 47)ECAPA-TDNN speaker embeddings (ONNX export)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · voice on CPU (2)
- Voice real-time factor on CPU: 3.28measured on our server 2026-09-25, 1 run
- Exact video length: 1 of 1 run (3 of 18 lines cut short)decosa-api docs/evals/consented-dubbing.md, 2026-09-25
Standard · the hosted demo, voice on a shared GPU (6)
- Track length equals the video's to the sample: 11 of 11 runs; 6 of 6 synthetic edge cases (29.97 fps, audio longer or shorter than video, audio-only)decosa-api docs/evals/consented-dubbing.md, 2026-09-25
- Consent gate: decisions as expected: 9 of 9 (allowed, revoked, strike, territory, purpose, project, swapped voice, no entry), all receipteddecosa-api docs/evals/consented-dubbing.md, 2026-09-25
- Subtitle QA: planted errors caught: 793 of 793 planted by script (16 kinds), 15 of 15 in the hand-made demo file, 0 false alarms on the clean hand-made filedecosa-api docs/evals/consented-dubbing.md, 2026-09-25; the planter and the checker have the same author
- ASR WER against the script: 1.0-3.9%decosa-api docs/evals/consented-dubbing.md, 2026-09-25
- Lines cut short to fit: 12 of 160 lines across 10 GPU runsdecosa-api docs/evals/consented-dubbing.md, 2026-09-25
- Dub quality rated by a native speaker: not measured yet