Skip to content
decosa

Dub a video in your own voice

A Spanish track for a creator's video in the creator's own consented voice, or in a consented dubber's. The consent ledger is checked when the job starts, before the voice is rendered and again at approval. The track is cut to the video's exact length, the original track is never replaced, and subtitles get a QA pass. Nothing is published until a person approves. Approval seals a signed record and a C2PA label: "AI-dubbed, voice consented by" the named person.

For
Teams in film, tv and games and creative and media.
  • Time per task74 stypical (median) on the sample
  • Cost per task~$0.73 per 100 minutes of videomeasured, at list price
  • AccuracyNo accuracy eval yet
  • Self-hostRun it on your own server: nothing is sent to us or anyone else.
  • HostedUse it on our servers; see where your data goes.

Video

Mara Quill dubs her own 57-second whetstone how-to into Spanish in her consented voice. Glossary: whetstone, burr, strop, the channel name.

Voice
id_mara-quill-demo
Use
dubbing · project quill-workshop · territory WW
Target
Spanish (neutral)
Glossary
whetstone → piedra de afilar · burr → rebaba · strop → asentar · Quill Workshop (keep) · Mara (keep)

Fictional people with synthetic stock voices and self-made videos. A dub takes about a minute and each demo session gets two. The dub is an extra track: the original is never replaced, and nothing is published until a person approves it.

Result

Live

Pick a sample and dub it. The consent gate runs first; a refused voice never reaches the GPU.

Check any subtitle file

The same QA that runs on our dubs, on a file a person made: timing, reading speed, line length, glossary and SDH style. Code only, no model. It flags; a person decides.

A Spanish SDH subtitle file for the whetstone video, written by hand for this demo, with 15 errors planted on purpose in 13 edits: an overlap, a bad timecode, end before start, a line too long, three lines, a cue too fast to read, an empty cue, a missing and an avoided glossary term, mixed bracket styles, mixed speaker labels, a cue past the end of the video and a 50 ms gap.

What it does, in shortWho it's for, where it runs and the key results

A Spanish track for a creator's video in the creator's own consented voice, or in a consented dubber's. The consent ledger is checked when the job starts, before the voice is rendered and again at approval. The track is cut to the video's exact length, the original track is never replaced, and subtitles get a QA pass. Nothing is published until a person approves. Approval seals a signed record and a C2PA label: "AI-dubbed, voice consented by" the named person.

In short

Last reviewed

What it is
Your voice, your approval: a Spanish track cut to the exact video length, released only after you approve it.
Who it's for
Teams in film, tv and games and creative and media.
Where it runs
Hosted with the demo's fictional voices; self-host with your own consented voices
Key numbers
  • 11 of 11 Dub tracks exactly the video's length (sample count) (synthetic, n = 11)
  • 9 of 9 Consent gate decisions as expected (synthetic, n = 9)
  • 3.9% (own-voice), 1.0% (dubber-voice) ASR word error rate against the narration script (synthetic)
  • 74.0 s Median end-to-end run, hosted (QA sweep 2026-09-26)
All results, datasets and caveats
How we tested itEnd-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 74 s · ~$0.003 per run · 43 receipts

Loading the nightly status…

Self-host: verified 26 Sep 2026 · Fresh clone of the branch into a clean directory, docker build of the api image plus the documented voice-runtime layer, compose with named volumes, pointed at the already-running local Voxtral and Qwen3.8-27B (direct route), voice on CPU; then torn down.

Measured cost to run: about $0.73 per 100 minutes of video (hosted, 26 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The revoked sample was refused; the own-voice dub reached review in 337 s (including the first download of the voice weights) with an exact 2,736,000-sample track, glossary 10 of 10 and a render receipt; approval gave a C2PA-signed release (state Valid) and a record that verifies; no transcript or approver text in the container logs. Found on the way: chatterbox-tts pins numpy<1.26, which has no wheels for the image's Python 3.12, so the documented layer now builds the voice venv with Python 3.11 via uv.

Known limits (5)
  • One speaker, English to Spanish; no lip-sync; music under the voice is not carried into the dub track.
  • The voice keeps some English accent (cross-language cloning), and no native speaker has rated it yet: that is why approval is required.
  • About 1 line in 13 is cut short to fit its slot (12 of 160 on GPU runs).
  • Approval proves the job's key or session pressed Approve on the exact draft, not which person did.
  • C2PA credentials use a development certificate, so public validators show the issuer as untrusted.

Eval results, nightly checks and cost per run · eval not held out

Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

1 law, rule and guidance page cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels, what it is built from
Models
Voxtral Mini 4B Realtime · Qwen3.8-27B · Chatterbox Multilingual
Where
Hosted with the demo's fictional voices; self-host with your own consented voices
Checks
Receipts for every ASR segment and model call; signed consent decisions; signed record and C2PA manifest at approval
Output
Media · Signed record or verdict
Data
Personal data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Every result carries a signed record of which model produced it, so you can check it later. How that works

Questions people ask

Will the dub track be the same length as my video?

Yes, to the sample. The dub WAV has exactly round(video length x 48,000) samples: 11 of 11 real runs and 6 of 6 awkward test files (29.97 fps, audio longer or shorter than the video, 22.05 kHz, audio-only).

Can it clone whoever appears in my video?

No. Uploaded audio is only transcribed. The voice reference is always the consent clip enrolled in the consent ledger, and every voice render checks the ledger three times: at submit, before the render and at approval.

Does it replace my original audio?

No. The dub is an extra track and the original is never replaced. Nothing is released until a person approves the exact draft they reviewed; the release carries a signed record and a C2PA label reading "AI-dubbed, voice consented by <name> (consent entry <id>)".

How good is the Spanish?

Not yet rated by a native speaker, which is why approval is required. The measurable proxies: glossary terms rendered as required 98 of 98, back-translation chrF mean 73.4, and about 1 line in 13 cut short to fit its slot. The voice keeps some English accent.

What does a video cost and how long does it take?

About $0.002-0.003 of model time for a 43-57 s video (4 receipted translation calls, measured). The voice runs on your own GPU or CPU: about 75 s from upload to review for a 57 s video on a GPU, about 6 minutes on CPU.

What does it not do yet?

One speaker, English to Spanish only. No lip-sync, and music under the voice is not carried into the dub track.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Consented creator dubbing

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.