Dub a video in your own voice
A Spanish track for a creator's video in the creator's own consented voice, or in a consented dubber's. The consent ledger is checked when the job starts, before the voice is rendered and again at approval. The track is cut to the video's exact length, the original track is never replaced, and subtitles get a QA pass. Nothing is published until a person approves. Approval seals a signed record and a C2PA label: "AI-dubbed, voice consented by" the named person.
- For
- Teams in film, tv and games and creative and media.
- Time per task74 stypical (median) on the sample
- Cost per task~$0.73 per 100 minutes of videomeasured, at list price
- AccuracyNo accuracy eval yet
Result
LivePick a sample and dub it. The consent gate runs first; a refused voice never reaches the GPU.
Check any subtitle file
The same QA that runs on our dubs, on a file a person made: timing, reading speed, line length, glossary and SDH style. Code only, no model. It flags; a person decides.
A Spanish SDH subtitle file for the whetstone video, written by hand for this demo, with 15 errors planted on purpose in 13 edits: an overlap, a bad timecode, end before start, a line too long, three lines, a cue too fast to read, an empty cue, a missing and an avoided glossary term, mixed bracket styles, mixed speaker labels, a cue past the end of the video and a 50 ms gap.
Watch: consent checked, dubbed, approved
Replay · not liveRecorded on 2026-09-30 from real runs on production decosa-api (api.decosa.ai): Voxtral ASR, Qwen3.8-27B through our gateway, Chatterbox Multilingual on the shared GPU, the consent ledger, and a human Approve. Fictional people, synthetic stock voices, self-made videos. Snapshots taken while a job was still running keep its progress only (its lines and receipts are in the review snapshot), to keep the fixture small.
Mara's whetstone video: ASR, glossary translation, her consented voice, an exact-length track, subtitles with QA, then Approve seals the record and the C2PA label.
Pick a sample and dub it. The consent gate runs first; a refused voice never reaches the GPU.
Get an API key
- Call the consented creator dubbing API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and Voxtral; the voice model needs about 4 GB of GPU or runs on CPU at about 4× real time.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- consented-dubbing
Use the hosted API
# Decosa consented creator dubbing: use the hosted API
You are wiring Decosa's consented dubbing into this project. It takes a one-speaker video (English narration or a talking
head) and returns an extra Spanish audio track in a voice whose owner consented to this exact use. The track is cut to
exactly the video's length. It comes with SRT and WebVTT subtitles, a subtitle QA report and a hand-off script for
human dubbers. Nothing is released until a person approves; the approved release carries a signed record and a C2PA
label. Use only what is listed below. If you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- The hosted service dubs only with the voices listed at `GET /dubbing/voices`. In the demo they are fictional people
with synthetic stock voices. It never uses the uploaded audio as a voice. To dub with a real person's voice, that
person enrols in the consent ledger first (self-host is the right place for real voices).
- The approval must come from a person who watched the dub. Don't auto-approve in code.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in
code. Send `Authorization: Bearer $DECOSA_API_KEY`. A key gets 20 dub jobs per UTC day.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "consented-dubbing"}` returns `{"token", ...}`.
Each session gets 2 dub jobs. Over a limit you get HTTP 429 with `Retry-After`.
## Endpoints
- `POST /dubbing/uploads` (token): body = the video or audio file (`Content-Type: application/octet-stream`, up to 60 MB
and 3 minutes) → `{upload_id, probe}`. Kept for 24 hours.
- `POST /dubbing/jobs` (token): `{"upload_id", "voice", "project", "purpose"?: "dubbing"|"narration"|"accessibility",
"territory"?: "MX"|..|"WW", "variant"?: "neutral"|"es-ES"|"es-MX", "glossary"?: [{"source", "target", "keep"?, "avoid"?, "also"?}]}`
or `{"sample_id": "own-voice"}`. You get 202 `{job_id, consent, poll}`. You get 422 `{consent: {code, reason,
receipt_id}}` when the consent doesn't cover this use: `revoked`, `strike_suspended`, `territory_not_covered`,
`purpose_not_covered`, `project_not_in_scope`, `no_entry`, `voice_mismatch` and so on. Nothing is rendered then.
- `GET /dubbing/jobs/{id}` (token): poll every 3-5 s until `status` is `review`, `refused` or `failed`. `review` has
`lines` (source, Spanish, back-translation), `track` (`exact`, `samples`, `expected_samples`), `qa`, `files`
(download URLs with a preview key) and `review_hash`.
- `POST /dubbing/jobs/{id}/approve` (token): `{"approver": "name", "review_hash": "...", "confirm": true}`. This checks
consent again, seals the record and embeds the C2PA label. It returns `release.files`: the signed `dub-es.wav`,
subtitles, `record.json`, `manifest.json` and `release.zip`. If the consent was revoked since the render, you get
422 and nothing is published.
- `GET /dubbing/releases/{id}/{name}` (no token, only after approval). Verify `record.json` at `POST /record/verify`
with `{"record": ...}`.
- `POST /dubbing/qa` (token): `{"subtitles": "<SRT or VTT>", "source_subtitles"?, "glossary"?, "settings"?, "media_seconds"?}`
returns a QA report (timing, reading speed, line length, glossary, SDH style). It uses code only, no model call, and
works on human-made files.
- `POST /dubbing/gate` (token): the consent decision only. `GET /dubbing/info`, `/dubbing/samples`, `/dubbing/voices`
need no token.
## Example: dub a sample, show it to a reviewer, then approve (Python, `pip install httpx`)
```python
import httpx, os, pathlib, time
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
r = httpx.post(f"{API}/dubbing/jobs", headers=H, timeout=60, json={"sample_id": "own-voice"})
if r.status_code == 422:
raise SystemExit(f"consent refused: {r.json()['consent']['code']}")
r.raise_for_status()
job_id = r.json()["job_id"]
while True:
time.sleep(5)
job = httpx.get(f"{API}/dubbing/jobs/{job_id}", headers=H, timeout=60).json()
if job["status"] not in ("queued", "running"):
break
assert job["status"] == "review", job.get("error")
print("exact length:", job["track"]["exact"], job["track"]["samples"])
# show job["files"] (preview.mp4, dub-es.wav, subtitles) to a person; only when they approve:
ok = httpx.post(f"{API}/dubbing/jobs/{job_id}/approve", headers=H, timeout=60,
json={"approver": "Reviewer name", "review_hash": job["review_hash"], "confirm": True}).json()
wav = next(f for f in ok["release"]["files"] if f["name"] == "dub-es.wav")
pathlib.Path("dub-es.wav").write_bytes(httpx.get(API + wav["url"]).content)
print(ok["release"]["label"])
```
## Honest limits
- One speaker, English to Spanish, no lip-sync, and no music under the voice in the dub track.
- The voice keeps some English accent. A native speaker hasn't rated it; that's why approval is required.
- A 57 s video takes about 75 s on the shared GPU (about 6 minutes when the voice has to run on CPU).
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa consented creator dubbing: run it yourself (containers)
You are setting up Decosa's consented dubbing on this machine. It has Voxtral ASR, Qwen3.8-27B translation with a
glossary, the Decosa consent ledger, and Chatterbox Multilingual (MIT) speaking in a consented voice. The output is a
track cut to exactly the video's length, plus subtitles with QA, a hand-off script and a receipted human approval.
Videos, voices and consent clips stay here; nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Don't substitute other images or models.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/consented-dubbing.zip (756 KB, 17 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py consented-dubbing` (the api image carries the same bundle under /app/rehearsal/consented-dubbing/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py consented-dubbing --bundle consented-dubbing.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the consent gate allows Mara Quill's own voice for her channel", "a dub in Theo Marsh's voice is refused because he revoked his consent", "the refusal is a receipted consent-ledger decision and no voice was rendered"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Rules first
- Every voice render goes through the consent ledger. Never bypass it. Enrol a real person only with their recorded
consent and a specific description of the use.
- Keep the Perth watermark on. Never replace the original track. Release only after a person approves.
## Steps
1. Docker and the NVIDIA container toolkit: if `docker compose version` or
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi` fails, install them from the official
instructions. Voxtral and Qwen3.8-27B NVFP4 fit on one 96 GB card. The voice runs on CPU (about 3.3x real time) or
on a GPU with 8 GB free.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the
`asr`, `llm` and `api` services. The api image needs the voice runtime layer: FFmpeg, c2pa-python and a
`/opt/chatterbox` venv with `chatterbox-tts==0.1.4`. The tool's assemble prompt has the exact Dockerfile. Use
named volumes and bind every port to 127.0.0.1.
3. `docker compose up -d`, wait for the health checks, then create the C2PA signing material once:
`docker compose exec api python scripts/provenance_devcert.py && docker compose restart api`.
4. Smoke test: get a token from `POST /demo/session {"vertical":"consented-dubbing"}`. `POST /dubbing/jobs
{"sample_id":"revoked"}` must answer 422 with `consent.code` `revoked`. `POST /dubbing/jobs {"sample_id":"own-voice"}`,
then poll to `review` and check `track.exact` is true. Approve with the `review_hash`, and check that
`POST /record/verify` accepts `/dubbing/releases/{id}/record.json`.
5. Report back: the signing key id (`GET /attest/signing-key`), the time to review, and the voice device (CPU or GPU).
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Voxtral Mini 4B Realtime needs a GPU.
- GeForce RTX 4090Doesn't fit
Needs about 40 GB of GPU memory at the smallest settings; 24 GB available.
- GeForce RTX 5090Doesn't fit
Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions.
- L40SDoesn't fit
Needs about 53.6 GB of GPU memory at the smallest settings; 48 GB available.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (85.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (85.6 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBCan't tell
Memory not known for Chatterbox Multilingual (t3_23lang) has no mapped Apple Silicon build.
- Apple M5 Max, 64 GBCan't tell
Memory not known for Chatterbox Multilingual (t3_23lang) has no mapped Apple Silicon build.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API, the speech model and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with both models loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "asr": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"consented-dubbing"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py consented-dubbing
Download the mock-data bundle (756 KB, 17 checks)expected.json
A 57-second fictional how-to video (sharpening on a whetstone) by the fictional creator Mara Quill, dubbed into Spanish in her own consented voice with a glossary. The job must stop at human review with a track of exactly the video's length and every glossary term right; a request in a voice whose consent was revoked must be refused before anything is rendered; and the subtitle QA must find the planted errors in a hand-made SDH file and nothing in the clean one. It does not approve or publish: that is the person's step. One dub takes about a minute on a free GPU; when the voice model falls back to CPU (it needs 8 GB of free GPU memory) it takes about 8 to 10 minutes, and the rehearsal waits up to 20.
What the rehearsal checks
- the consent gate allows Mara Quill's own voice for her channel
- a dub in Theo Marsh's voice is refused because he revoked his consent
- the refusal is a receipted consent-ledger decision and no voice was rendered
- the subtitle QA finds the planted overlap at cue 2
- the subtitle QA finds three lines in cue 6
- the subtitle QA finds the end-before-start timing in cue 8
- the subtitle QA finds the avoided glossary term in cue 12
- the subtitle QA finds the empty cue (14), the cue past the end of the video (18) and the mixed SDH styles
- the subtitle QA finds nothing in the clean file
- the dub stops at human review (it is not published)
- the dubbed track is exactly the video's length, to the sample (57.0 s at 48 kHz)
- every glossary term is translated as the glossary says (10 of 10)
- the channel name Quill Workshop is kept unchanged in the Spanish
- the voice render's receipt verifies against this server's key
- the render receipt points at Mara Quill's active consent entry
- that consent entry is not revoked
- every speech and model call has a signed receipt
Licence: Fictional creators and scripts, title-card videos made with ffmpeg and a hand-written Spanish SDH file, all written for Decosa; the demo voices are synthetic stock voices (Kokoro-82M, Apache-2.0; Chatterbox, MIT). No real person's voice or likeness. Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa consented creator dubbing: run it yourself (containers)
You are setting up Decosa's consented dubbing on this machine. It has Voxtral ASR, Qwen3.8-27B translation with a
glossary, the Decosa consent ledger, and Chatterbox Multilingual (MIT) speaking in a consented voice. The output is a
track cut to exactly the video's length, plus subtitles with QA, a hand-off script and a receipted human approval.
Videos, voices and consent clips stay here; nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Don't substitute other images or models.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/consented-dubbing.zip (756 KB, 17 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py consented-dubbing` (the api image carries the same bundle under /app/rehearsal/consented-dubbing/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py consented-dubbing --bundle consented-dubbing.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the consent gate allows Mara Quill's own voice for her channel", "a dub in Theo Marsh's voice is refused because he revoked his consent", "the refusal is a receipted consent-ledger decision and no voice was rendered"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Rules first
- Every voice render goes through the consent ledger. Never bypass it. Enrol a real person only with their recorded
consent and a specific description of the use.
- Keep the Perth watermark on. Never replace the original track. Release only after a person approves.
## Steps
1. Docker and the NVIDIA container toolkit: if `docker compose version` or
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi` fails, install them from the official
instructions. Voxtral and Qwen3.8-27B NVFP4 fit on one 96 GB card. The voice runs on CPU (about 3.3x real time) or
on a GPU with 8 GB free.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the
`asr`, `llm` and `api` services. The api image needs the voice runtime layer: FFmpeg, c2pa-python and a
`/opt/chatterbox` venv with `chatterbox-tts==0.1.4`. The tool's assemble prompt has the exact Dockerfile. Use
named volumes and bind every port to 127.0.0.1.
3. `docker compose up -d`, wait for the health checks, then create the C2PA signing material once:
`docker compose exec api python scripts/provenance_devcert.py && docker compose restart api`.
4. Smoke test: get a token from `POST /demo/session {"vertical":"consented-dubbing"}`. `POST /dubbing/jobs
{"sample_id":"revoked"}` must answer 422 with `consent.code` `revoked`. `POST /dubbing/jobs {"sample_id":"own-voice"}`,
then poll to `review` and check `track.exact` is true. Approve with the `review_hash`, and check that
`POST /record/verify` accepts `/dubbing/releases/{id}/record.json`.
5. Report back: the signing key id (`GET /attest/signing-key`), the time to review, and the voice device (CPU or GPU).
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Doesn't fitConsented creator dubbing on GeForce RTX 5090
Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.
Better voice and lip-sync: what changesuses estimates
Nothing: it runs as listed in the stack.
Memory per component
- Alternate: Chatterbox Multilingual V3 / Single Language Pack (LatAm Spanish). ~2.2 GB, weights 1 GB, loaded while a job runs (estimate). Estimate: 0.5B parameters at 2 bytes (BF16 assumed; the quantisation is not stated) per weight is about 1 GB, plus 20% working memory and 1 GB of runtime. Not measured.
- Alternate: InfiniteTalk. Memory not known. Memory not stated in stack.json and not derivable (no parameter count).
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Consented creator dubbing, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Consented creator dubbing on my hardware Fetch https://decosa.ai/prompts/consented-dubbing-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=consented-dubbing) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Better voice and lip-sync (alternate-voice-lipsync). Fit check: can't tell, about 2.2 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Alternate: Chatterbox Multilingual V3 / Single Language Pack (LatAm Spanish) (ResembleAI/chatterbox), 2.2 GB - Alternate: InfiniteTalk (MeiGen-AI/InfiniteTalk), unknown GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Chatterbox Multilingual V3 / Single Language Pack (LatAm Spanish) (while a job runs) ~2.2 GB (7%); about 29.8 GB left Memory is not known for: InfiniteTalk (memory not known). Load them one at a time and watch memory before running everything together. During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/consented-dubbing-assemble.md
Get an API key
- Call the consented creator dubbing API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and Voxtral; the voice model needs about 4 GB of GPU or runs on CPU at about 4× real time.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Your voice, your approval: a Spanish track cut to the exact video length, released only after you approve it.
For YouTubers, educators and small localisation vendors. The service transcribes a one-speaker video and translates it with the creator's glossary. It then speaks the Spanish in the creator's own consented voice, or in a consented dubber's, and cuts the track to exactly the video's length so platforms accept it. The consent ledger is checked when the job starts, before the voice render and at approval. Subtitles get a QA pass that also works on human-made files, and a hand-off script supports paid human dubbing. Nothing is released until a person approves; the release carries a signed record and a C2PA label.
- Deployment
- Hosted or self-host
- Regulatory
- Not legal advice. Laws read on 25 Sep 2026 (the consent ledger's sources, linked there): Tennessee's ELVIS Act (in force 1 Jul 2024) makes whoever distributes a tool whose primary purpose is producing someone's voice without authorization liable. That is why every render here goes through the consent ledger. California AB 2602 (signed 17 Sep 2024) and New York S7676B (contracts from 1 Jan 2025) make a replica clause that replaces in-person work unenforceable without a reasonably specific use description and counsel or union representation; the ledger stores both. EU AI Act Art. 50(2) (applies from 2 Aug 2026) requires machine-readable marking of synthetic audio; the C2PA manifest and the Perth watermark cover it. Credentials use a development certificate, so public validators show the issuer as untrusted. YouTube's disclosure rules (Help 14328491, read 25 Sep 2026) exempt cloning your own voice for dubs; a dubber's clone of someone else is not in that exemption, so label it. Guild terms (SAG-AFTRA 2023 and 2026) were not verified against the contract text.
Text description
A creator's video enters decosa-api. Consent gate 1 checks the voice's ledger entry for this project, purpose, territory and date. Speech spans are transcribed by Voxtral (Apache-2.0) with signed receipts. Qwen3.8-27B (Apache-2.0) translates with the glossary and back-translates through our gateway. Consent gate 2 runs right before Chatterbox Multilingual (MIT) speaks each line in the consented voice with the Perth watermark. Voxtral listens back and garbled lines are rendered again. Lines are fitted and the track is cut to exactly the video length, and subtitles are written and checked. A person approves the draft; consent gate 3 runs, a signed record is sealed and a C2PA label is embedded. Outputs: an extra audio track, subtitles with a QA report, and a hand-off script for human dubbers. Everything stays on your machine when self-hosted, except the optional gateway call.
At a glance
- Consent
- Every voice render checks the consent ledger three times (submit, render, approval), and each decision is signed. Uploaded audio is only transcribed, never used as a voice.
- Length
- The dub WAV has exactly round(video length x 48,000) samples: 11 of 11 real runs and 6 of 6 awkward test files.
- Approval
- Nothing is released until a person approves the exact draft they reviewed. The release gets a signed record and a C2PA label: "AI-dubbed, voice consented by <name> (consent entry <id>)". The original track is never replaced.
- Cost per video
- A fraction of a cent of model time for a short video (a few receipted translation calls, measured); the voice runs on your own GPU or CPU.
- Typical time
- About a minute from upload to review for a short video with the voice on a GPU; several minutes with the voice on CPU (measured).
- Data retention
- Jobs, uploads and files are deleted after 24 hours. Logs hold ids, counts and timings, never text.
- What leaves the box
- Hosted: the video, and the translation calls through our gateway. Self-host: nothing, unless you point translation at the gateway.
- For human dubbers
- A hand-off export (JSON and CSV) with timecodes, slot lengths, character budgets, the glossary and the machine draft, plus a subtitle QA that works on human-made files.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
voice on CPU
The same pipeline and gates with the voice model on CPU: no GPU needed for the voice, about five times slower end to end.
- Models
- Voxtral Mini 4B Realtime
- Qwen3.8-27B (NVIDIA NVFP4)
- Chatterbox Multilingual (t3_23lang)
- ECAPA-TDNN speaker embeddings (ONNX export)
- Hardware
- A 96 GB card (or remote endpoints) for Voxtral and Qwen3.8-27B; the voice on a 16-thread CPU
- Quality evidence
- Voice real-time factor on CPU3.28measured on our server 2026-09-25, 1 run
- Exact video length1 of 1 run (3 of 18 lines cut short)decosa-api docs/evals/consented-dubbing.md, 2026-09-25
- Latency
- measured: several minutes for a short video (one run)
- Verification
- Proof: partialSelf-host onlyTranslation calls get Decosa API receipts on the gateway route; ASR and the voice render are attested by the server's key.
- In the hosted demo
Standard
the hosted demo, voice on a shared GPU
Voxtral, Qwen3.8-27B through our gateway, and Chatterbox Multilingual on GPU when 8 GB are free (CPU otherwise).
- Models
- Voxtral Mini 4B Realtime
- Qwen3.8-27B (NVIDIA NVFP4)
- Chatterbox Multilingual (t3_23lang)
- ECAPA-TDNN speaker embeddings (ONNX export)
- Hardware
- 1x RTX PRO 6000 96 GB for the text models plus about 4 GB of GPU for the voice while a job runs
- Quality evidence
- Track length equals the video's to the sample11 of 11 runs; 6 of 6 synthetic edge cases (29.97 fps, audio longer or shorter than video, audio-only)decosa-api docs/evals/consented-dubbing.md, 2026-09-25
- Consent gate: decisions as expected9 of 9 (allowed, revoked, strike, territory, purpose, project, swapped voice, no entry), all receipteddecosa-api docs/evals/consented-dubbing.md, 2026-09-25
- Subtitle QA: planted errors caught793 of 793 planted by script (16 kinds), 15 of 15 in the hand-made demo file, 0 false alarms on the clean hand-made filedecosa-api docs/evals/consented-dubbing.md, 2026-09-25; the planter and the checker have the same author
- ASR WER against the script1.0-3.9%decosa-api docs/evals/consented-dubbing.md, 2026-09-25
- Lines cut short to fit12 of 160 lines across 10 GPU runsdecosa-api docs/evals/consented-dubbing.md, 2026-09-25
- Dub quality rated by a native speakernot measured yet
- Latency
- measured on our server: about a minute from submit to review for short videos
- Verification
- Proof: partialTranslation calls get Decosa API receipts; ASR segments, the voice render and the consent decisions are signed by the server and the ledger.
Also runs on
- Better voice and lip-syncChatterbox Multilingual V3 / Single Language Pack (LatAm Spanish), InfiniteTalknot servedChatterbox Multilingual V3 or its LatAm Spanish pack for the voice, and InfiniteTalk (Apache-2.0, Wan2.1-14B based) for lip-sync on talking heads. Neither has been run here. Hardware: One 96 GB card (not measured).
We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Transcribes each speech span of the source, and listens back to every dubbed lineVoxtral Mini 4B Realtimemistralai/Voxtral-Mini-4B-Realtime-2602 on Hugging Face (opens in a new tab) 4.4B · 24 GBProof: partialIn the hosted demo | LiteStandard | 4.4B · 24 GB | Proof: partialIn the hosted demo | |
| ||||
Translation with the glossary and a length budget per line, repair of lines that break the glossary or budget, back-translation for the reviewerQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 57 GBProof: strongIn the hosted demo | LiteStandard | 27.8B · 57 GB | Proof: strongIn the hosted demo | |
| ||||
Speaks each line in the consented voice (the enrolled consent clip is the reference), with the Perth watermarkChatterbox Multilingual (t3_23lang)ResembleAI/chatterbox on Hugging Face (opens in a new tab) 0.5B · 4 GBNo proof yetIn the hosted demo | LiteStandard | 0.5B · 4 GB | No proof yetIn the hosted demo | |
| ||||
Consent ledger's speaker check: the reference clip must match the enrolled voiceprint (tool 47)ECAPA-TDNN speaker embeddings (ONNX export)speechbrain/spkrec-ecapa-voxceleb on Hugging Face (opens in a new tab) 0 GBNo proof yetIn the hosted demo | LiteStandard | 0 GB | No proof yetIn the hosted demo | |
| ||||
Alternate: the newer multilingual weights and the Latin American Spanish single-language pack from the same repositoryChatterbox Multilingual V3 / Single Language Pack (LatAm Spanish)ResembleAI/chatterbox on Hugging Face (opens in a new tab) 0.5BNo proof yetSelf-host only | Alternate | 0.5B | No proof yetSelf-host only | |
| ||||
Alternate: lip-sync for talking-head videoInfiniteTalkMeiGen-AI/InfiniteTalk on Hugging Face (opens in a new tab) No proof yetSelf-host only | Alternate | n/a | No proof yetSelf-host only | |
| ||||
Tools, services and hardware
Tools
- FFmpeg / ffprobe (opens in a new tab)LGPL-2.1+ (GPL builds vary)
Decoding, the video-stream duration, atempo speed-up, the preview mux.
- c2pa-python (opens in a new tab)MIT OR Apache-2.0
Embeds the C2PA manifest in the released WAV.
The implicit audio watermark Chatterbox adds, and its detector (run on the final track).
- Decosa consent ledger (tool 47) (opens in a new tab)part of decosa-api
Signed consent entries and receipted gate decisions for every voice render.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag> (publishing soon) plus the voice-runtime layer in the assemble promptPipeline, consent gates, the voice worker process, release. Binds 127.0.0.1.
- decosa-asr:8090
${DECOSA_REGISTRY}/decosa-asr:0.1.0Voxtral realtime endpoint.
- decosa-llm:8114
${DECOSA_REGISTRY}/decosa-llm:0.1.0Qwen3.8-27B OpenAI endpoint (hosted: behind our gateway).
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB + CPU Fits
Hosted layout on our server: Voxtral and the voice model share GPU0 (the voice loads per job, only with 8 GB free); Qwen3.8-27B is on GPU1 behind the gateway.
- CPU only for the voice Fits
Measured: 3.3x real time on 16 threads, so a 57 s video took about 6 minutes end to end. The text models still need a GPU or a remote endpoint.
Latency per lane
- 57 s video to a dub ready for review (voice on GPU)74.0 s
Measuredmeasured on our server 2026-09-25/26: own-voice sample, 63-80 s over 6 runs, gateway route, shared GPU0
- 43 s video to a dub ready for review (voice on GPU)52.0 s
Measuredmeasured on our server 2026-09-25/26: dubber-voice sample, 47-57 s over 4 runs
- 57 s video, voice on CPU365.0 s
Measuredmeasured on our server 2026-09-25: one run, 16 threads, 4 lines rendered twice by the listen-back check
- Consent refusal5 ms
Measuredmeasured on our server 2026-09-25: the 422 comes back before any audio is read
- Approval to release (record, C2PA, bundle)800 ms
Measuredmeasured on our server 2026-09-26: 0.6-0.9 s over 4 runs
Notes
- Uploaded video is only transcribed. The voice reference is always the enrolled consent clip, so the tool cannot clone whoever is in the video.
- The dub is an extra track; the original is never replaced. Nothing is published until a person approves the exact draft they reviewed (the review hash).
- Hy-MT2 is a candidate dedicated translation model to test next (licence not checked yet). Nemotron 3 Diarization (OpenMDW) would be needed for more than one speaker.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble Decosa consented creator dubbing on this machine
You are setting up consented dubbing. It takes a creator's video and gives back an extra Spanish audio track in a
voice whose owner consented to this use. The track is cut to exactly the video's length, and it comes with subtitles, a
subtitle QA report, a hand-off script for human dubbers, and a signed record plus a C2PA label once a person approves.
Work step by step. Show me each command before you run anything with `sudo`, and stop to ask if a check fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/consented-dubbing.zip (756 KB, 17 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py consented-dubbing` (the api image carries the same bundle under /app/rehearsal/consented-dubbing/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py consented-dubbing --bundle consented-dubbing.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the consent gate allows Mara Quill's own voice for her channel", "a dub in Theo Marsh's voice is refused because he revoked his consent", "the refusal is a receipted consent-ledger decision and no voice was rendered"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- Models: **Voxtral Mini 4B Realtime** (ASR, Apache-2.0) and **Qwen3.8-27B** (translation, Apache-2.0).
**Chatterbox Multilingual** (voice, MIT, `ResembleAI/chatterbox` @ `5bb1f6ee58e50c3b8d408bc82a6d3740c2db6e18`, the
`t3_23lang` weights that `chatterbox-tts==0.1.4` loads). Keep its Perth watermark on.
- **Every voice render needs consent.** This stack refuses to render any voice without an active entry in the Decosa
consent ledger. The entry must cover this project, the purpose (dubbing, narration or accessibility), the territory
and today's date. Don't bypass the gate. A voice-clone tool without one is the kind of tool Tennessee's ELVIS Act
(in force 1 Jul 2024) makes its distributor liable for.
- Enrol real people only with their recorded consent. A living performer licensing a replica to someone else needs
counsel or union representation in the entry (California AB 2602 and New York S7676B test for it). The consent
ledger's own assemble prompt explains the entry fields.
- Uploaded video is only transcribed. Its audio is never used as a voice. The voice reference is the consent clip you
enrolled.
- Nothing is published until a person presses Approve. Never replace the original track.
- Bind every port to 127.0.0.1. Logs carry ids, counts and timings, never transcript or translation text. Keep it that
way.
## 1. Check the machine
1. `nvidia-smi`. Qwen3.8-27B NVFP4 needs about 57 GB and Voxtral about 24 GB; a 96 GB card holds both. If you already
run them, just point at them. The voice model runs on CPU at about 3.3x real time (a 57 s video took about 4 min on
our box) or on a GPU with at least 8 GB free at about 0.25x real time.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, ask me, then
install them from the official repositories.
3. Disk: about 10 GB for the voice runtime and weights (CPU torch plus 3 GB of Chatterbox files), plus the model servers.
## 2. Model servers
Use the Decosa ASR and LLM images (`${DECOSA_REGISTRY}/decosa-asr:<tag>` and `decosa-llm:<tag>`, **publishing
soon**), or any vLLM that serves:
- Voxtral as a realtime WebSocket at `ws://127.0.0.1:8090/v1/realtime` (served name `voxtral-realtime`);
- Qwen3.8-27B as an OpenAI endpoint at `http://127.0.0.1:8114/v1`. Use its served name below; ours is `qwen3.8-27b`.
## 3. Build the api image with the voice runtime
`${DECOSA_REGISTRY}/decosa-api:<tag>` is **publishing soon**. Until then, build from source:
`git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required). Check out a release that contains
`decosa_api/verticals/dubbing/` and `decosa_api/verticals/consent/`, then run
`docker build -f docker/api/Dockerfile -t decosa-api:local .`. Then add the voice runtime in its own venv. Its
dependencies pin older numpy and torch than the api's, so it runs as a separate process:
```dockerfile
# ./api-dub/Dockerfile
FROM decosa-api:local
USER root
# chatterbox-tts pins numpy<1.26, which has no wheels for the image's Python 3.12: the voice venv uses Python 3.11 via uv
RUN pip install --no-cache-dir "c2pa-python>=0.37" "pillow>=10" uv && \
UV_PYTHON_INSTALL_DIR=/opt/uv-python uv venv -p 3.11 /opt/chatterbox && \
uv pip install --python /opt/chatterbox/bin/python --no-cache --index-strategy unsafe-best-match \
--extra-index-url https://download.pytorch.org/whl/cpu "torch==2.6.0+cpu" "torchaudio==2.6.0+cpu" \
"numpy>=1.24,<1.26" librosa==0.11.0 s3tokenizer transformers==4.46.3 diffusers==0.29.0 resemble-perth==1.0.1 \
conformer==0.3.2 safetensors==0.5.3 pykakasi==2.3.0 soundfile huggingface_hub "setuptools<70" && \
uv pip install --python /opt/chatterbox/bin/python --no-cache --no-deps chatterbox-tts==0.1.4 && \
chmod -R a+rX /opt/uv-python /opt/chatterbox && mkdir -p /hf && chown decosa /hf
USER decosa
ENV DECOSA_DUB_TTS_PYTHON=/opt/chatterbox/bin/python HF_HOME=/hf
```
`docker build -t decosa-api:dub ./api-dub`. Why these pins: `pkuseg` (Chinese only) doesn't build, so the package goes
in with `--no-deps`. The Perth watermarker imports `pkg_resources`, hence `setuptools<70`. For a GPU voice, build a
second venv with CUDA torch (we use `torch==2.7.1+cu128` on Blackwell), set `DECOSA_DUB_TTS_PYTHON_GPU` to it and
`DECOSA_DUB_TTS_DEVICE=auto`. The model then loads per job and only when the GPU has 8 GB free.
## 4. docker-compose.yml
Write this in `~/decosa/dubbing/`. `network_mode: host` lets the api reach your model servers on 127.0.0.1.
```yaml
services:
api:
image: decosa-api:dub
network_mode: host
environment:
DECOSA_HOST: 127.0.0.1
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_ASR_WS: ws://127.0.0.1:8090/v1/realtime
DECOSA_ASR_MODEL: voxtral-realtime
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://127.0.0.1:8114/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_DUB_TTS_DEVICE: cpu
DECOSA_PROVENANCE_DIR: /provenance
volumes: ["decosa-data:/data", "decosa-provenance:/provenance", "hf-cache:/hf"]
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/dubbing/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
decosa-data:
decosa-provenance:
hf-cache:
```
Use named volumes, never host folders: the api runs as uid 10001, and a host folder Docker creates belongs to root.
`docker compose up -d`, wait for the health check, then create the C2PA signing material once:
`docker compose exec api python scripts/provenance_devcert.py && docker compose restart api`. It is a development CA,
so public validators show the signature as valid and the issuer as untrusted. On first start the api creates this box's
Ed25519 key in `/data/attest/`; the ledger, receipts and records are signed with it. Back it up with
`docker compose cp api:/data/attest ./attest-backup` and never print it.
## 5. Smoke test
```bash
B=http://127.0.0.1:8445
curl -s $B/dubbing/info | jq '{tts: .models.tts.licence, ledger: .consent.ledger}' # "MIT", true
T=$(curl -s -XPOST $B/demo/session -H 'content-type: application/json' -d '{"vertical":"consented-dubbing"}' | jq -r .token)
curl -s -XPOST $B/dubbing/jobs -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample_id":"revoked"}' | jq '.consent.code' # "revoked"
J=$(curl -s -XPOST $B/dubbing/jobs -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample_id":"own-voice"}' | jq -r .job_id)
until curl -s $B/dubbing/jobs/$J -H "authorization: Bearer $T" | jq -e '.status!="queued" and .status!="running"' >/dev/null; do sleep 10; done
curl -s $B/dubbing/jobs/$J -H "authorization: Bearer $T" | jq '{status, exact: .track.exact, samples: .track.samples, glossary: .translation.glossary.rate, qa: .qa.target.verdict}'
H=$(curl -s $B/dubbing/jobs/$J -H "authorization: Bearer $T" | jq -r .review_hash)
curl -s -XPOST $B/dubbing/jobs/$J/approve -H "authorization: Bearer $T" -H 'content-type: application/json' \
-d "{\"approver\":\"Self-host test\",\"review_hash\":\"$H\",\"confirm\":true}" | jq '.release | {label, exact, c2pa: .c2pa.validation_state}'
curl -s $B/dubbing/releases/$J/record.json | jq '{record: .}' | curl -s -XPOST $B/record/verify -H 'content-type: application/json' -d @- | jq .ok
```
Expect: the revoked voice refused before any work; the dub in `review` with `exact: true` and 2,736,000 samples (57 s
at 48 kHz); the release labelled "AI-dubbed, voice consented by Mara Quill (consent entry ce_…)" with the C2PA state
`Valid`; the record verifying `true`. On CPU the first dub took 337 s in our clean-room test (including the 3 GB weight download). Report what you measure.
## 6. Your own voices and videos
Enrol each voice in the consent ledger (`POST /consent/entries`, with the consent clip). Then list it in
`DECOSA_DUB_VOICES_DIR/voices.json` (same shape as `decosa_api/verticals/dubbing/data/voices.json`, with the consent
clip next to it) and restart. Upload a video with `POST /dubbing/uploads` (at most 60 MB and 3 minutes). Then
`POST /dubbing/jobs` with `upload_id`, `voice`, `project`, `territory` and your `glossary`.
## 7. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`. The contract is in `API_CONTRACT.md`,
section "Consented creator dubbing".What it does, in shortWho it's for, where it runs and the key results
A Spanish track for a creator's video in the creator's own consented voice, or in a consented dubber's. The consent ledger is checked when the job starts, before the voice is rendered and again at approval. The track is cut to the video's exact length, the original track is never replaced, and subtitles get a QA pass. Nothing is published until a person approves. Approval seals a signed record and a C2PA label: "AI-dubbed, voice consented by" the named person.
In short
Last reviewed
- What it is
- Your voice, your approval: a Spanish track cut to the exact video length, released only after you approve it.
- Who it's for
- Teams in film, tv and games and creative and media.
- Where it runs
- Hosted with the demo's fictional voices; self-host with your own consented voices
- Key numbers
- 11 of 11 Dub tracks exactly the video's length (sample count) (synthetic, n = 11)
- 9 of 9 Consent gate decisions as expected (synthetic, n = 9)
- 3.9% (own-voice), 1.0% (dubber-voice) ASR word error rate against the narration script (synthetic)
- 74.0 s Median end-to-end run, hosted (QA sweep 2026-09-26)
How we tested itEnd-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 74 s · ~$0.003 per run · 43 receipts
Loading the nightly status…
Self-host: verified 26 Sep 2026 · Fresh clone of the branch into a clean directory, docker build of the api image plus the documented voice-runtime layer, compose with named volumes, pointed at the already-running local Voxtral and Qwen3.8-27B (direct route), voice on CPU; then torn down.
Measured cost to run: about $0.73 per 100 minutes of video (hosted, 26 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The revoked sample was refused; the own-voice dub reached review in 337 s (including the first download of the voice weights) with an exact 2,736,000-sample track, glossary 10 of 10 and a render receipt; approval gave a C2PA-signed release (state Valid) and a record that verifies; no transcript or approver text in the container logs. Found on the way: chatterbox-tts pins numpy<1.26, which has no wheels for the image's Python 3.12, so the documented layer now builds the voice venv with Python 3.11 via uv.
Known limits (5)
- One speaker, English to Spanish; no lip-sync; music under the voice is not carried into the dub track.
- The voice keeps some English accent (cross-language cloning), and no native speaker has rated it yet: that is why approval is required.
- About 1 line in 13 is cut short to fit its slot (12 of 160 on GPU runs).
- Approval proves the job's key or session pressed Approve on the exact draft, not which person did.
- C2PA credentials use a development certificate, so public validators show the issuer as untrusted.
Eval results, nightly checks and cost per run · eval not held out
Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
1 law, rule and guidance page cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels, what it is built from
- Models
- Voxtral Mini 4B Realtime · Qwen3.8-27B · Chatterbox Multilingual
- Where
- Hosted with the demo's fictional voices; self-host with your own consented voices
- Checks
- Receipts for every ASR segment and model call; signed consent decisions; signed record and C2PA manifest at approval
- Industry
- Film, TV and games · Creative and media
- Job
- Translate · Generate media
- Input
- Files and media
- Output
- Media · Signed record or verdict
- Data
- Personal data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Part of
- Decosa Studio: Voice
- Runs in
- Decosa hosted · Self-host
- Built from
- Live speech to text · Consent gate · Content credentials · Signed record
Every result carries a signed record of which model produced it, so you can check it later. How that works
Questions people ask
Will the dub track be the same length as my video?
Yes, to the sample. The dub WAV has exactly round(video length x 48,000) samples: 11 of 11 real runs and 6 of 6 awkward test files (29.97 fps, audio longer or shorter than the video, 22.05 kHz, audio-only).
Can it clone whoever appears in my video?
No. Uploaded audio is only transcribed. The voice reference is always the consent clip enrolled in the consent ledger, and every voice render checks the ledger three times: at submit, before the render and at approval.
Does it replace my original audio?
No. The dub is an extra track and the original is never replaced. Nothing is released until a person approves the exact draft they reviewed; the release carries a signed record and a C2PA label reading "AI-dubbed, voice consented by <name> (consent entry <id>)".
How good is the Spanish?
Not yet rated by a native speaker, which is why approval is required. The measurable proxies: glossary terms rendered as required 98 of 98, back-translation chrF mean 73.4, and about 1 line in 13 cut short to fit its slot. The voice keeps some English accent.
What does a video cost and how long does it take?
About $0.002-0.003 of model time for a 43-57 s video (4 receipted translation calls, measured). The voice runs on your own GPU or CPU: about 75 s from upload to review for a 57 s video on a GPU, about 6 minutes on CPU.
What does it not do yet?
One speaker, English to Spanish only. No lip-sync, and music under the voice is not carried into the dub track.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Consented creator dubbing
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…