Eval: consented creator dubbing (48)
Run on 25-26 Sep 2026 on our server. On a pre-release build, dev server on 127.0.0.1:8448. Voxtral ran on the live
realtime endpoint; Qwen3.8-27B ran through the model gateway, which is shared and busy. Chatterbox Multilingual used
revision 5bb1f6ee, on GPU0 when it had at least 8 GB free, and on CPU for the one CPU run. Numbers are in
decosa_api/verticals/dubbing/data/eval.json; to reproduce, run python scripts/dubbing_eval.py all <jobs dir>.
What this eval is and isn't. It uses two self-made narration videos of fictional creators with synthetic stock voices (Kokoro-82M, Apache-2.0), written for this demo. There is no human rating of the Spanish or of the voice. We report proxies we can measure honestly. The subtitle checker and its planted errors have the same author. Nothing was tuned on these numbers after they were measured. We made two changes after the first two runs, before any number here was recorded: the listen-back check and the scheduler that lets a long line borrow slack from the next one. Both came from reading those first runs, which aren't counted.
1. Exact length (the "Audio must be the same as video length" fix)
| Set | Result |
|---|---|
| Real pipeline runs (11: 10 on GPU, 1 on CPU) | 11 of 11 WAVs have exactly round(video_s × 48000) samples: 2,736,000 for the 57.0 s video, 2,073,600 for the 43.2 s one |
| Released WAVs after C2PA signing (2) | still exact: the manifest goes in its own RIFF chunk and leaves the audio untouched |
| Synthetic edge cases (6): 25 fps; 29.97 fps at 12.345 s; 23.976 fps with audio 0.7 s shorter; 30 fps with audio 1.3 s longer; 60 fps at 7.777 s with 22.05 kHz audio; audio only | 6 of 6 exact (difference 0.000 ms) |
The length comes from the video stream's duration as ffprobe reports it, not from the container or the source audio. We haven't uploaded to YouTube. Its help page asks for "roughly the same length"; an exact track removes the guesswork.
2. Consent gate
| Case | Expected | Got |
|---|---|---|
| Creator's own voice, dubbing, worldwide | allowed | allowed |
| Consented dubber, Mexico | allowed | allowed |
| Revoked entry | revoked | revoked |
| Strike suspension | strike_suspended | strike_suspended |
| Territory outside the consent (Argentina) | territory_not_covered | territory_not_covered |
| Purpose outside the consent (accessibility) | purpose_not_covered | purpose_not_covered |
| Project outside the consent | project_not_in_scope | project_not_in_scope |
| Reference clip swapped for another voice | voice_mismatch | voice_mismatch |
| Identity with no entry | no_entry | no_entry |
9 of 9 as expected, and all 9 decisions are receipted. This is deterministic code; the swapped-voice case uses a stand-in verifier in the eval and in the tests. On the dev server the ledger's real verifier (ECAPA, use case 47) matched each demo reference clip against its enrolled voiceprint at score 1.00 (threshold 0.585).
End to end over HTTP: the refused samples (revoked, strike, wrong territory) answered 422 in milliseconds with the signed decision, and no voice was rendered (a test asserts the voice renderer is never called). One scenario matters most: consent revoked after the render and before approval. The approval then answers 422, the job stays in review and nothing is published (tested).
3. Pipeline runs
Eleven runs: own-voice (57 s, Mara Quill's own voice) seven times, six on GPU and one on CPU; dubber-voice (43 s, Lucia Ferrer's voice for Mexico) four times on GPU. The last three runs (two recorded for the site, one from the browser test) include the cue-timing fix in section 4; after it, the Spanish subtitles carry 1 reading-speed warning per run instead of 2-8.
| Metric | Result | How measured |
|---|---|---|
| ASR word error rate against the narration script | 3.9% (own-voice), 1.0% (dubber-voice) | Voxtral per speech span. Most own-voice errors are "First," heard as "1." and "Test it" as "Tested". The narration is synthetic and clean, so real creators will score worse. |
| Glossary terms rendered as required | 98 of 98 | After the code check and at most two repair calls. Inflected forms count only when the glossary lists them (also). |
| Back-translation chrF against the source | mean 73.4 (range 70.0-76.0) | A proxy for meaning kept, not a quality score. Shortening a line to fit its slot costs a few points. |
| Listen-back character error (Voxtral hears each dubbed line) | mean 9.4% per job (1.5-15.7%) | Lines above 35% or far too long are rendered again with another seed: 1-4 lines per run. |
| Lines cut short to fit | GPU runs 12 of 160; CPU run 3 of 18 | After speed-up (up to 1.25x) and borrowing up to 0.4 s from the next line. |
| Speaker similarity of the Spanish dub to the English consent clip | 0.81-0.90 (5 runs) | The ledger's ECAPA verifier; its accept threshold is 0.585. Informative only, not a gate. |
| Perth watermark detected in the final, fitted track | 11 of 11 | Resemble's detector, score 1.0 in every run. |
| Voice real-time factor | GPU 0.23-0.32; CPU 3.28 | Model load of 5-6 s per job not included. |
| Submit to ready for review | own-voice 63-80 s (6 GPU runs), 365 s (CPU); dubber-voice 47-57 s (4 runs) | Queue empty; the gateway was shared with other evaluation jobs. |
| Approval to release (record, C2PA, bundle) | 0.6-0.9 s | 4 recorded runs. |
| Model calls and cost per video | 4 calls; 2,454-3,573 tokens; $0.0018-0.0027 | Gateway list price $0.30 / $1.50 per million tokens. Voice GPU time is not billed. |
| Receipts per job | 31-44 | One per ASR span, per model call and per listen-back check, plus the signed render receipt for the track. |
4. Subtitle and SDH QA
- Planted by script: 16 kinds of error, 25 tries per kind, on two clean bases: the hand-made Spanish SDH file and our generated Spanish file for the same video. A try counts only where the base had no finding of that kind at that cue. 793 of 793 caught (100% for each kind). This shows that each check fires on the error it was built for. It says nothing about how many real-world subtitle problems there are.
- The hand-made demo file (
data/qa/planted-es.srt) has 15 planted errors in 13 edits; all 15 were found. The checker also reported 2 reading-speed warnings that the planted edits caused (a lengthened line and an added speaker label). Both are real. - False alarms: 0 on the clean hand-made file (18 cues). Our own generated Spanish file (from a run before the fix below) shows 8 reading-speed warnings. One is real (19.9 characters per second, fast source speech). The other 7 read "17.0 characters per second; limit 17": the pipeline extended cues for reading time but rounded the time down, so they came out a hair over the limit. The browser test found this; the pipeline now rounds up (regression test added). The checker was right each time. The generator was wrong.
- Not evaluated:
no_speech(it needs the source's speech map, so it runs only in the pipeline).
5. Known limits
- The voice keeps some English accent: the Chatterbox README says output can inherit the accent of a reference clip in another language, and a creator's own clip is in English. A native Spanish speaker hasn't rated it.
- No lip-sync, no stem separation (music under the voice is lost from the dub track) and one speaker only.
- The demo voices are synthetic stock voices. Real voices depend on the consent clip's quality.
- The approval proves that the holder of the job's key or demo session pressed Approve on the exact draft (the review hash). It can't prove which person that was.
- C2PA uses a development certificate: public validators show the signature as valid and the issuer as untrusted.
6. Datasets and licences
- Demo narration scripts: written for this demo (
data/scripts/). - Voices: Kokoro-82M stock voices af_heart, am_michael and ef_dora (Apache-2.0), and the built-in Chatterbox voice (MIT).
- Videos: title cards made with ffmpeg (
data/media/). - The Spanish SDH file: hand-written for this demo (
data/qa/). - No real person's voice or likeness is used.
7. Expected properties of the sample runs (for the rehearse-first kit)
Checkable on any run of the bundled samples, whatever the model's exact wording:
own-voicereachesreviewwithtrack.exact == trueandtrack.samples == 2736000(57.0 s at 48 kHz);dubber-voicewith 2,073,600.revoked,strikeandwrong-territoryanswer 422 withconsent.coderevoked,strike_suspendedandterritory_not_covered, acd_decision id, and no voice render.translation.glossary.correct == translation.glossary.occurrences(10 for own-voice, 7 for dubber-voice), and "Quill Workshop" / "Sundial Notes" appear unchanged in the Spanish.- After approval:
release.labelstarts with "AI-dubbed, voice consented by Mara Quill (consent entry ce_",release.c2pa.validation_state == "Valid",release.exact == true, andrecord.jsonverifies at/record/verify. POST /dubbing/qa {"sample_id":"planted"}finds all 15 planted findings;{"sample_id":"clean"}returns no issues.track.watermark_score >= 0.5(Perth detected) andoriginal_track_replacedis false in the record.