52 · Music · Film, TV and games · live
Music video from your track
Eval results
Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)
- Lyric lines within 0.3 s / 1 s of human timing81.8% / 87.7%test splitn = 66 held-out songs; 4 of 6 songs have at least 80% of lines within 1 s. Baseline 16.8% within 1 s. Dev: 67.0% / 67.9%.
- Beat F-measure (±70 ms), artist's BPM given0.868test splitn = 40Without BPM: 0.591. Drums-only grooves: an upper bound for full mixes.
- Typed check says no to a swapped (mismatched) scene67.3%test splitn = 52The code checks (citations, anchor words) carry the grounding; the typed check lets a third through.
- Anchor words found in the cited lines96.2%test splitn = 52Citations valid: 100%
- Rendered cut offset from the beat grid (median / max)7.0-10.0 ms / 16.3 mssynthetic5 renders of one excerpt; against the detected grid, not a human one. Some cuts between similar shots are not detectable (13-21 of 20-21 found).
- Render cost, one formatabout $0.32 per minute of videosyntheticAt an assumed $1.69/h GPU rental price; the hosted demo does not bill renders.
Dataset
JamendoLyrics MultiLang (9 English songs with human word/line timings: 3 dev, 6 test), Groove MIDI Dataset (40 test grooves; tuned on the validation split), and renders of one CC BY excerpt and one generated track.
Caveats
- Visual quality is not scored; 480p upscaled clips are clearly softer than closed models.
- Beat tracking measured on synthesised drums only, not real full mixes.
- English only (the aligner is English-only).
- Small sets: 6 test songs; the cut-detector threshold was set on the same render it was measured on.
- The containerised analyzer placed one beat grid about 160 ms later than on the host.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 26 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 15 s
- Receipts
- 7
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.002
Self-host verification
Verified on 26 Sep 2026: Fresh clone of the branch into a clean directory, docker build of the api image and services/mvideo, compose with named volumes on host networking, pointed at the already-running local Qwen (direct route) and ComfyUI; then torn down.
No rights statement: 403; the sample analysed in 19 s (112.35 BPM, 18 lines, 6 grounded scenes, 20 cuts, 7 attested receipts); a 9:16 render finished with a C2PA credential, the disclosure rules met and a record that verified. Found on the way: the containerised analyzer's beat grid sat about 160 ms later than the host's on the same file (different decoder), and a clip of flickering neon fooled the cut detector (fixed: isolated spikes only).
Rehearsal bundle: music-video-studio.zip (825 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Visuals are 480p clips upscaled to 720p: a stylised visualiser, clearly below closed video models and the H3 samples on this site.
- Lyric timing is English only; 82% of held-out lines start within 0.3 s, and when it slips whole passages slip.
- Without the artist's BPM the beat tracker got the tempo right on 43% of test grooves (100% with it); which beat is "one" is a heuristic.
- The hosted demo caps a video at 60 s and renders share one GPU: 3 renders per session, 3 per API key per day, about 11 minutes per minute of video per format.
- The rights statement is not verified; the clearance pre-check covers only a small open catalogue.
- C2PA credentials are signed by a development CA: valid signature, untrusted issuer in public validators.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Beat, bar and section detection (CPU): librosa beat tracker on a full-band plus low-band onset envelope, bar phase by a kick-and-snare heuristic (4/4), sections by checkerboard novelty on chroma and MFCC self-similaritydecosa-mvideo-analyze (services/mvideo)AGPL-3.0-or-later (decosa-api)
- Lyric timing (CPU): CTC forced alignment of the artist's own lyrics on the mix, with a repair pass for lines squeezed into too little time; English letterswav2vec2-large-960h-lv60-selfApache-2.0
- Treatment writer (scenes that cite the lyric lines they show) and the typed yes/no grounding check per scene; also labels near-duplicate lyric lines in the clearance pre-checkQwen3.8-27B (NVFP4)Apache-2.0
- Sample and lyric clearance pre-check on the upload (tool 38, run in-process): audio landmarks and melody against a small open catalogue, lyric lines against a lyric setdecosa-api clearance module (decosa_api/verticals/clearance)AGPL-3.0-or-later
- Clips: text-to-video, one 5 s clip per scene per formatWan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAsApache-2.0
- The edit and the marks (CPU): cuts on bar lines on a 30 fps grid, karaoke captions (ASS, libass), the AI label on a top bar, the credit, the artist's audio; cut timing measured back from the pixels; the disclosure pre-flight's rules (tool 49) on the file; C2PA credential per file; the signed recorddecosa-api mvideo module + FFmpeg + c2pa-pythonAGPL-3.0-or-later
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · timed captions and a cited treatment, no GPU for video (3)
- Lyric line starts within 0.3 s / 1 s of human timing (6 held-out CC BY-ND songs): 81.8% / 87.7% (baseline 16.8% within 1 s)decosa-api docs/evals/music-video-studio.md, 2026-09-26
- Beat F-measure (±70 ms), 40 human-played grooves: 0.87 with the artist's BPM, 0.59 withoutdecosa-api docs/evals/music-video-studio.md, 2026-09-26; drums only
- Treatment citations valid / anchor words found (52 held-out scenes): 100% / 96%decosa-api docs/evals/music-video-studio.md, 2026-09-26
Standard · the hosted demo, clips on one 96 GB card (5)
- Cuts found in the rendered files / offset from the beat grid: 133 of 145 planned cuts found, 0 false; median 7-10 ms, max 16.3 ms (half a frame at 30 fps)decosa-api docs/evals/music-video-studio.md, 2026-09-26; 7 files from 5 renders; offsets against the detected beat grid
- GPU time per minute of video: about 680 GPU-seconds for one format, 1,360-1,740 for bothdecosa-api docs/evals/music-video-studio.md, 2026-09-26; shared card
- Disclosure rules on the delivered files (tool 49): label read on 6 of 6 sampled frames and the C2PA marking valid in all 4 files checkeddecosa-api docs/evals/music-video-studio.md, 2026-09-26
- Typed grounding check: swapped (mismatched) scenes caught: 67% (the code checks on citations and anchor words do the rest)decosa-api docs/evals/music-video-studio.md, 2026-09-26
- Visual quality: not measured; 480p upscaled, clearly below closed models
Best · MiniMax H3 clips (licence pending), self-host (1)
- clip quality: not measured yet
Wanted · MiniMax H3 on two cards, no offload (1)
- render time per clip against one card with offload: not measured yet