{"schema_version":"1","site":"https://decosa.ai","id":"consented-dubbing","num":"48","name":"Consented creator dubbing","tool_name":"Dub a video in your own voice","short":"Consented dubbing","blurb":"A Spanish track for a creator's video in the creator's own consented voice, or in a consented dubber's. The consent ledger is checked when the job starts, before the voice is rendered and again at approval. The track is cut to the video's exact length, the original track is never replaced, and subtitles get a QA pass. Nothing is published until a person approves. Approval seals a signed record and a C2PA label: \"AI-dubbed, voice consented by\" the named person.","status":"live","labels":{"industry":["entertainment","creative-media"],"job":["translate","generate"],"input":["files"],"deploy":["hosted","selfhost"],"status":"live","output":["media","record"],"data":["pii"],"hardware":"gpu-96","licence":"permissive"},"industries":["entertainment","creative-media"],"runs_in":["hosted","selfhost"],"part_of":["studio-voice"],"built_from":["live-asr","consent-gate","content-credentials","signed-record"],"models":"Voxtral Mini 4B Realtime · Qwen3.8-27B · Chatterbox Multilingual","where":"Hosted with the demo's fictional voices; self-host with your own consented voices","hardware":"1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and Voxtral; the voice model needs about 4 GB of GPU or runs on CPU at about 4× real time","final_artifact":"A C2PA-labelled WAV cut to the exact video length, SRT and WebVTT subtitles with a QA report, a hand-off script for a human dubber, and a signed record.","self_host_first":false,"verification":{"receipt_coverage":"partial","summary":"Receipts for every ASR segment and model call; signed consent decisions; signed record and C2PA manifest at approval","manual_qa":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":74000,"p95_ms":null,"runs":null,"receipts_per_run":43,"cost_per_run_usd":0.0026},"selfhost":{"date":"2026-09-26","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of the api image plus the documented voice-runtime layer, compose with named volumes, pointed at the already-running local Voxtral and Qwen3.8-27B (direct route), voice on CPU; then torn down.","notes":"The revoked sample was refused; the own-voice dub reached review in 337 s (including the first download of the voice weights) with an exact 2,736,000-sample track, glossary 10 of 10 and a render receipt; approval gave a C2PA-signed release (state Valid) and a record that verifies; no transcript or approver text in the container logs. Found on the way: chatterbox-tts pins numpy<1.26, which has no wheels for the image's Python 3.12, so the documented layer now builds the voice venv with Python 3.11 via uv."},"known_limits":["One speaker, English to Spanish; no lip-sync; music under the voice is not carried into the dub track.","The voice keeps some English accent (cross-language cloning), and no native speaker has rated it yet: that is why approval is required.","About 1 line in 13 is cut short to fit its slot (12 of 160 on GPU runs).","Approval proves the job's key or session pressed Approve on the exact draft, not which person did.","C2PA credentials use a development certificate, so public validators show the issuer as untrusted."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Dub tracks exactly the video's length (sample count)","value":"11 of 11","unit":null,"n":11,"split":"synthetic","note":"Plus 6 of 6 synthetic frame-rate edge cases exact"},{"name":"Consent gate decisions as expected","value":"9 of 9","unit":null,"n":9,"split":"synthetic","note":"Deterministic code; all 9 receipted"},{"name":"ASR word error rate against the narration script","value":"3.9% (own-voice), 1.0% (dubber-voice)","unit":null,"n":null,"split":"synthetic","note":"Synthetic, clean narration; real creators will score worse"},{"name":"Glossary terms rendered as required","value":"98 of 98","unit":null,"n":98,"split":"synthetic","note":null},{"name":"Back-translation chrF against the source","value":"mean 73.4 (range 70.0-76.0)","unit":null,"n":null,"split":"synthetic","note":"A proxy for meaning kept, not a quality score"},{"name":"Lines cut short to fit","value":"GPU runs 12 of 160; CPU run 3 of 18","unit":null,"n":178,"split":"synthetic","note":null},{"name":"Planted subtitle errors caught","value":"793 of 793","unit":null,"n":793,"split":"synthetic","note":"16 kinds x 25 tries; shows each check fires, not real-world prevalence. 0 false alarms on the clean hand-made file"}],"dataset":"Two self-made narration videos (57 s and 43.2 s) of fictional creators with synthetic stock voices, 11 pipeline runs; a hand-written Spanish SDH file with planted errors; 6 synthetic length edge cases.","held_out":false,"caveats":["No human rating of the Spanish or of the voice: the numbers are measurable proxies.","Synthetic stock voices and clean narration; real voices depend on the consent clip's quality.","The subtitle checker and its planted errors have the same author.","The voice keeps some English accent; a native Spanish speaker has not rated it. No lip-sync, no stem separation, one speaker only."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/consented-dubbing"},"stack":{"summary":"For YouTubers, educators and small localisation vendors. The service transcribes a one-speaker video and translates it with the creator's glossary. It then speaks the Spanish in the creator's own consented voice, or in a consented dubber's, and cuts the track to exactly the video's length so platforms accept it. The consent ledger is checked when the job starts, before the voice render and at approval. Subtitles get a QA pass that also works on human-made files, and a hand-off script supports paid human dubbing. Nothing is released until a person approves; the release carries a signed record and a C2PA label.","tagline":"Your voice, your approval: a Spanish track cut to the exact video length, released only after you approve it.","deployment":"hosted-or-self-host","regulatory_note":"Not legal advice. Laws read on 25 Sep 2026 (the consent ledger's sources, linked there): Tennessee's ELVIS Act (in force 1 Jul 2024) makes whoever distributes a tool whose primary purpose is producing someone's voice without authorization liable. That is why every render here goes through the consent ledger. California AB 2602 (signed 17 Sep 2024) and New York S7676B (contracts from 1 Jan 2025) make a replica clause that replaces in-person work unenforceable without a reasonably specific use description and counsel or union representation; the ledger stores both. EU AI Act Art. 50(2) (applies from 2 Aug 2026) requires machine-readable marking of synthetic audio; the C2PA manifest and the Perth watermark cover it. Credentials use a development certificate, so public validators show the issuer as untrusted. YouTube's disclosure rules (Help 14328491, read 25 Sep 2026) exempt cloning your own voice for dubs; a dubber's clone of someone else is not in that exemption, so label it. Guild terms (SAG-AFTRA 2023 and 2026) were not verified against the contract text.","components":[{"id":"asr","role":"Transcribes each speech span of the source, and listens back to every dubbed line","name":"Voxtral Mini 4B Realtime","hf_repo":"mistralai/Voxtral-Mini-4B-Realtime-2602","license":"Apache-2.0","params":"4.4B","quant":"BF16","vram_gb":24,"memory_gb_estimate":null,"engine":"vLLM realtime WebSocket (/v1/realtime)","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["all"],"alternative_to":null},{"id":"llm","role":"Translation with the glossary and a length budget per line, repair of lines that break the glossary or budget, back-translation for the reviewer","name":"Qwen3.8-27B (NVIDIA NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP) + FP8 (attention/GDN), FP8 KV cache","vram_gb":57,"memory_gb_estimate":null,"engine":"vLLM 0.29.0","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["all"],"alternative_to":null},{"id":"tts","role":"Speaks each line in the consented voice (the enrolled consent clip is the reference), with the Perth watermark","name":"Chatterbox Multilingual (t3_23lang)","hf_repo":"ResembleAI/chatterbox","license":"MIT","params":"0.5B","quant":"FP32","vram_gb":4,"memory_gb_estimate":null,"engine":"chatterbox-tts 0.1.4 in its own venv, one worker process per job (CPU torch 2.6.0 or CUDA torch 2.7.1)","receipt_coverage":"none","in_hosted_demo":true,"tiers":["lite","standard"],"alternative_to":null},{"id":"verifier","role":"Consent ledger's speaker check: the reference clip must match the enrolled voiceprint (tool 47)","name":"ECAPA-TDNN speaker embeddings (ONNX export)","hf_repo":"speechbrain/spkrec-ecapa-voxceleb","license":"Apache-2.0","params":null,"quant":"FP32","vram_gb":0,"memory_gb_estimate":null,"engine":"ONNX Runtime on CPU (the consent extra)","receipt_coverage":"none","in_hosted_demo":true,"tiers":["all"],"alternative_to":null},{"id":"tts-v3","role":"Alternate: the newer multilingual weights and the Latin American Spanish single-language pack from the same repository","name":"Chatterbox Multilingual V3 / Single Language Pack (LatAm Spanish)","hf_repo":"ResembleAI/chatterbox","license":"MIT","params":"0.5B","quant":null,"vram_gb":null,"memory_gb_estimate":null,"engine":"needs a newer chatterbox-tts than 0.1.4 (not tested)","receipt_coverage":"none","in_hosted_demo":false,"tiers":["alternates"],"alternative_to":null},{"id":"lipsync","role":"Alternate: lip-sync for talking-head video","name":"InfiniteTalk","hf_repo":"MeiGen-AI/InfiniteTalk","license":"Apache-2.0","params":null,"quant":null,"vram_gb":null,"memory_gb_estimate":null,"engine":"not integrated","receipt_coverage":"none","in_hosted_demo":false,"tiers":["alternates"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · voice on CPU","summary":"The same pipeline and gates with the voice model on CPU: no GPU needed for the voice, about five times slower end to end.","components":["asr","llm","tts","verifier"],"hardware":"A 96 GB card (or remote endpoints) for Voxtral and Qwen3.8-27B; the voice on a 16-thread CPU","quality_evidence":[{"metric":"Voice real-time factor on CPU","value":"3.28","source":"measured on our server 2026-09-25, 1 run"},{"metric":"Exact video length","value":"1 of 1 run (3 of 18 lines cut short)","source":"decosa-api docs/evals/consented-dubbing.md, 2026-09-25"}],"latency_note":"measured: several minutes for a short video (one run)","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"Translation calls get Decosa API receipts on the gateway route; ASR and the voice render are attested by the server's key.","hosting":null},{"id":"standard","label":"Standard · the hosted demo, voice on a shared GPU","summary":"Voxtral, Qwen3.8-27B through our gateway, and Chatterbox Multilingual on GPU when 8 GB are free (CPU otherwise).","components":["asr","llm","tts","verifier"],"hardware":"1x RTX PRO 6000 96 GB for the text models plus about 4 GB of GPU for the voice while a job runs","quality_evidence":[{"metric":"Track length equals the video's to the sample","value":"11 of 11 runs; 6 of 6 synthetic edge cases (29.97 fps, audio longer or shorter than video, audio-only)","source":"decosa-api docs/evals/consented-dubbing.md, 2026-09-25"},{"metric":"Consent gate: decisions as expected","value":"9 of 9 (allowed, revoked, strike, territory, purpose, project, swapped voice, no entry), all receipted","source":"decosa-api docs/evals/consented-dubbing.md, 2026-09-25"},{"metric":"Subtitle QA: planted errors caught","value":"793 of 793 planted by script (16 kinds), 15 of 15 in the hand-made demo file, 0 false alarms on the clean hand-made file","source":"decosa-api docs/evals/consented-dubbing.md, 2026-09-25; the planter and the checker have the same author"},{"metric":"ASR WER against the script","value":"1.0-3.9%","source":"decosa-api docs/evals/consented-dubbing.md, 2026-09-25"},{"metric":"Lines cut short to fit","value":"12 of 160 lines across 10 GPU runs","source":"decosa-api docs/evals/consented-dubbing.md, 2026-09-25"},{"metric":"Dub quality rated by a native speaker","value":"not measured yet","source":null}],"latency_note":"measured on our server: about a minute from submit to review for short videos","in_hosted_demo":true,"receipt_coverage":"partial","receipt_note":"Translation calls get Decosa API receipts; ASR segments, the voice render and the consent decisions are signed by the server and the ledger.","hosting":null}],"alternates":[{"id":"voice-lipsync","label":"Better voice and lip-sync","components":["tts-v3","lipsync"],"hardware":"One 96 GB card (not measured)","use":"Chatterbox Multilingual V3 or its LatAm Spanish pack for the voice, and InfiniteTalk (Apache-2.0, Wan2.1-14B based) for lip-sync on talking heads. Neither has been run here.","status":"not served"}],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag> (publishing soon) plus the voice-runtime layer in the assemble prompt","purpose":"Pipeline, consent gates, the voice worker process, release. Binds 127.0.0.1."},{"name":"decosa-asr","port":8090,"image":"${DECOSA_REGISTRY}/decosa-asr:0.1.0","purpose":"Voxtral realtime endpoint."},{"name":"decosa-llm","port":8114,"image":"${DECOSA_REGISTRY}/decosa-llm:0.1.0","purpose":"Qwen3.8-27B OpenAI endpoint (hosted: behind our gateway)."}],"tools":[{"name":"FFmpeg / ffprobe","url":"https://ffmpeg.org/","license":"LGPL-2.1+ (GPL builds vary)","purpose":"Decoding, the video-stream duration, atempo speed-up, the preview mux."},{"name":"c2pa-python","url":"https://github.com/contentauth/c2pa-python","license":"MIT OR Apache-2.0","purpose":"Embeds the C2PA manifest in the released WAV."},{"name":"Resemble Perth","url":"https://github.com/resemble-ai/Perth","license":"MIT","purpose":"The implicit audio watermark Chatterbox adds, and its detector (run on the final track)."},{"name":"Decosa consent ledger (tool 47)","url":"https://decosa.ai/apps/consent-ledger","license":"part of decosa-api","purpose":"Signed consent entries and receipted gate decisions for every voice render."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB + CPU","fits":true,"notes":"Hosted layout on our server: Voxtral and the voice model share GPU0 (the voice loads per job, only with 8 GB free); Qwen3.8-27B is on GPU1 behind the gateway."},{"tier":"CPU only for the voice","fits":true,"notes":"Measured: 3.3x real time on 16 threads, so a 57 s video took about 6 minutes end to end. The text models still need a GPU or a remote endpoint."}],"latency":[{"lane":"57 s video to a dub ready for review (voice on GPU)","typical_ms":74000,"source":"measured on our server 2026-09-25/26: own-voice sample, 63-80 s over 6 runs, gateway route, shared GPU0"},{"lane":"43 s video to a dub ready for review (voice on GPU)","typical_ms":52000,"source":"measured on our server 2026-09-25/26: dubber-voice sample, 47-57 s over 4 runs"},{"lane":"57 s video, voice on CPU","typical_ms":365000,"source":"measured on our server 2026-09-25: one run, 16 threads, 4 lines rendered twice by the listen-back check"},{"lane":"Consent refusal","typical_ms":5,"source":"measured on our server 2026-09-25: the 422 comes back before any audio is read"},{"lane":"Approval to release (record, C2PA, bundle)","typical_ms":800,"source":"measured on our server 2026-09-26: 0.6-0.9 s over 4 runs"}],"benchmark":null,"notes":["Uploaded video is only transcribed. The voice reference is always the enrolled consent clip, so the tool cannot clone whoever is in the video.","The dub is an extra track; the original is never replaced. Nothing is published until a person approves the exact draft they reviewed (the review hash).","Hy-MT2 is a candidate dedicated translation model to test next (licence not checked yet). Nemotron 3 Diarization (OpenMDW) would be needed for more than one speaker."]},"buyer_facts":[{"label":"Consent","value":"Every voice render checks the consent ledger three times (submit, render, approval), and each decision is signed. Uploaded audio is only transcribed, never used as a voice."},{"label":"Length","value":"The dub WAV has exactly round(video length x 48,000) samples: 11 of 11 real runs and 6 of 6 awkward test files."},{"label":"Approval","value":"Nothing is released until a person approves the exact draft they reviewed. The release gets a signed record and a C2PA label: \"AI-dubbed, voice consented by <name> (consent entry <id>)\". The original track is never replaced."},{"label":"Cost per video","value":"A fraction of a cent of model time for a short video (a few receipted translation calls, measured); the voice runs on your own GPU or CPU."},{"label":"Typical time","value":"About a minute from upload to review for a short video with the voice on a GPU; several minutes with the voice on CPU (measured)."},{"label":"Data retention","value":"Jobs, uploads and files are deleted after 24 hours. Logs hold ids, counts and timings, never text."},{"label":"What leaves the box","value":"Hosted: the video, and the translation calls through our gateway. Self-host: nothing, unless you point translation at the gateway."},{"label":"For human dubbers","value":"A hand-off export (JSON and CSV) with timecodes, slot lengths, character budgets, the glossary and the machine draft, plus a subtitle QA that works on human-made files."}],"data_handling":{"page":"/data#consented-dubbing","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Jobs, uploads and files are deleted after 24 hours. Logs hold ids, counts and timings, never text.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/apps/consented-dubbing","input":"dubbing","lanes":[{"id":"consent","title":"Consent gate","kind":"list"},{"id":"track","title":"Dub track","kind":"list"},{"id":"script","title":"Script and subtitles","kind":"list"},{"id":"release","title":"Approval and release","kind":"json"}],"samples":[{"n":1,"id":"own-voice","title":"Own voice","deep_link":"/apps/consented-dubbing?sample=1&autorun=0"},{"n":2,"id":"dubber-voice","title":"Dubber voice","deep_link":"/apps/consented-dubbing?sample=2&autorun=0"},{"n":3,"id":"revoked","title":"Revoked","deep_link":"/apps/consented-dubbing?sample=3&autorun=0"},{"n":4,"id":"strike","title":"Strike","deep_link":"/apps/consented-dubbing?sample=4&autorun=0"},{"n":5,"id":"wrong-territory","title":"Wrong territory","deep_link":"/apps/consented-dubbing?sample=5&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/consented-dubbing-hosted.md","selfhost":"/prompts/consented-dubbing-selfhost.md","assemble":"/prompts/consented-dubbing-assemble.md","mac":null},"rehearsal":{"bundle":"/samples/consented-dubbing.zip","bundle_url":"https://decosa.ai/samples/consented-dubbing.zip","folder":"/samples/consented-dubbing/","expected":"/samples/consented-dubbing/expected.json","files":["/samples/consented-dubbing/expected.json","/samples/consented-dubbing/inputs/clean-es.srt","/samples/consented-dubbing/inputs/dub-request.json","/samples/consented-dubbing/inputs/glossary.json","/samples/consented-dubbing/inputs/planted-es.srt","/samples/consented-dubbing/inputs/qa-glossary.json","/samples/consented-dubbing/inputs/quill-workshop-whetstone-script.txt","/samples/consented-dubbing/inputs/quill-workshop-whetstone.mp4","/samples/consented-dubbing/inputs/source-en.srt"],"bytes":774333,"checks":["the consent gate allows Mara Quill's own voice for her channel","a dub in Theo Marsh's voice is refused because he revoked his consent","the refusal is a receipted consent-ledger decision and no voice was rendered","the subtitle QA finds the planted overlap at cue 2","the subtitle QA finds three lines in cue 6","the subtitle QA finds the end-before-start timing in cue 8","the subtitle QA finds the avoided glossary term in cue 12","the subtitle QA finds the empty cue (14), the cue past the end of the video (18) and the mixed SDH styles","the subtitle QA finds nothing in the clean file","the dub stops at human review (it is not published)","the dubbed track is exactly the video's length, to the sample (57.0 s at 48 kHz)","every glossary term is translated as the glossary says (10 of 10)","the channel name Quill Workshop is kept unchanged in the Spanish","the voice render's receipt verifies against this server's key","the render receipt points at Mara Quill's active consent entry","that consent entry is not revoked","every speech and model call has a signed receipt"],"licence":"Fictional creators and scripts, title-card videos made with ffmpeg and a hand-written Spanish SDH file, all written for Decosa; the demo voices are synthetic stock voices (Kokoro-82M, Apache-2.0; Chatterbox, MIT). No real person's voice or likeness. Part of decosa-api, AGPL-3.0-or-later.","about":"A 57-second fictional how-to video (sharpening on a whetstone) by the fictional creator Mara Quill, dubbed into Spanish in her own consented voice with a glossary. The job must stop at human review with a track of exactly the video's length and every glossary term right; a request in a voice whose consent was revoked must be refused before anything is rendered; and the subtitle QA must find the planted errors in a hand-made SDH file and nothing in the clean one. It does not approve or publish: that is the person's step. One dub takes about a minute on a free GPU; when the voice model falls back to CPU (it needs 8 GB of free GPU memory) it takes about 8 to 10 minutes, and the rehearsal waits up to 20.","run":{"containers":"docker compose exec api python scripts/rehearse.py consented-dubbing","checkout":"python scripts/rehearse.py consented-dubbing --bundle consented-dubbing.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py consented-dubbing"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=consented-dubbing","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":85.6,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":85.6,"basis":"stack","unknown":[]},{"id":"alternate-voice-lipsync","gpu_gb":2.2,"basis":"estimate","unknown":["lipsync"]}],"mac":null},"links":{"page":"/apps/consented-dubbing","json":"/use-cases/consented-dubbing.json","metrics":"/metrics/consented-dubbing","console":"/apps/consented-dubbing","console_sample":"/apps/consented-dubbing?sample=1&autorun=0","stack":"/apps/consented-dubbing#stack","try_live":"/apps/consented-dubbing#live","watch":"/apps/consented-dubbing#watch","build":"/apps/consented-dubbing#build","self_host":"/apps/consented-dubbing#self-host","prompts":{"hosted":"/prompts/consented-dubbing-hosted.md","selfhost":"/prompts/consented-dubbing-selfhost.md","assemble":"/prompts/consented-dubbing-assemble.md","mac":null}}}