Eval: music video from your track (52)
Run on our server, 2026-09-26. Raw results in docs/evals/music-video-studio/ (align.json, beats.json, grounding.json,
renders.json). Scripts: scripts/mvideo_eval_audio.py (align, beats), scripts/mvideo_eval_grounding.py. Everything
below was measured; nothing was tuned on a test set except where it says so.
Data and licences
| Set | What | Licence | Used for |
|---|---|---|---|
| JamendoLyrics MultiLang (github.com/f90/jamendolyrics @ f093b7a7) | 9 English songs without an NC clause, with human word and line timings | annotations MIT; audio CC per song: 2 CC BY, 1 CC BY-SA, 6 CC BY-ND | lyric alignment (all 9), treatment grounding (lyrics only), demo excerpts (the 2 CC BY songs only) |
| Groove MIDI Dataset (Magenta, groove-v1.0.0-midionly) | human drummers playing to a click, MIDI | CC BY 4.0 | beat tracking (40 test grooves), tuning (validation split) |
| "Paper Lanterns" | a 60 s MiniMax-Music3 song made through music-gen-cleared (37) from lyrics written for Decosa | MiniMax-Music3 Community License output; lyrics CC0 | demo sample, render |
The CC BY-ND songs are only measured, never adapted or shown. Demo videos are made only from the CC BY excerpts (credited on screen and in the record) and the Music3 track.
1. Lyric caption timing (forced alignment)
facebook/wav2vec2-large-960h-lv60-self (Apache-2.0) CTC forced alignment of the artist's lyrics on the full mix, CPU.
Two settings were chosen on the three dev songs (Rxbyn, Cortez, Lower Loveday): align on the full mix rather than a
drum-removed (HPSS) version (it made Rxbyn worse: median line error 0.07 s vs 0.48 s), and the squeeze-repair
threshold (0.045 s per letter). The six CC BY-ND songs are the test set. Metric: line start error (when a caption appears)
and word onset error, against the dataset's human annotations. Baseline: lines spread evenly between the first and
last true line start (it is given the true vocal span, so it is generous).
| Split | Songs | Lines within 0.3 s | Lines within 1 s | Words within 0.3 s | Median of per-song median line error | Songs with >= 80% of lines within 1 s | Baseline: lines within 1 s |
|---|---|---|---|---|---|---|---|
| test | 6 | 81.8% | 87.7% | 82.5% | 0.042 s | 4 of 6 | 16.8% |
| dev | 3 | 67.0% | 67.9% | 66.9% | 0.070 s | 1 of 3 | 15.5% |
When it is right it is right to a few tens of milliseconds; when it is wrong, whole passages slip (Ridgway, test: 74% of lines within 1 s, the rest many seconds late; Lower Loveday, dev: only 27% placed, a doubled vocal the model did not hear). The per-line confidence flags some of these; the console marks low-confidence lines and the artist can paste LRC times instead. English only. Alignment took 18-33 s per 3-5 minute song on 16 CPU threads.
2. Beat tracking
librosa's beat tracker on a full-band plus low-band onset envelope (chosen on the GMD validation split: F-measure 0.52 with librosa's default median envelope, 0.62 with this one). Test: the first 40 4/4 grooves of 20 s or more in the GMD test split, synthesised with a small drum kit. Ground truth is the click grid the drummers played to. Drums alone are easier than a full mix; read these as an upper bound.
| Mode | Mean beat F (±70 ms) | Grooves with F >= 0.9 | Tempo within 4% | Double/half tempo errors | Median offset of matched beats | Bar lines on true downbeats |
|---|---|---|---|---|---|---|
| no BPM given | 0.591 | 35% | 42.5% | 6 of 40 | 14.9 ms | 20% |
| artist's BPM given | 0.868 | 85% | 100% | 0 | 15.4 ms | 51% |
So: ask for the BPM (the console does, and the API takes bpm). The bar phase (which beat is "one") is a kick-and-snare
heuristic and is often wrong; cuts still land on beats, but not always on the downbeat. downbeat_s fixes it.
3. Cut-on-beat accuracy (rendered files)
Cuts are planned on bar lines and snapped to the 30 fps output grid, so the planned error is at most half a frame (16.7 ms). To check the files, cuts are found again from the pixels (ffmpeg scene score) and compared with the nearest beat of the analysis grid (Rxbyn excerpt, 50 s, 20-21 planned cuts per render).
(At the first threshold, 0.3, only 7 and 9 of the 21 cuts were found, because many cuts join similar dark night shots; 0.08 added 15-25 false cuts. 0.15 was then set on this same render. The self-host check then rendered a clip of flickering neon, where a plain 0.15 threshold reported 52 false cuts, so the detector now also requires a cut to be an isolated spike: at least 2.5 times the highest score of the 6 frames either side. Re-measured with that rule, every render so far has 0 false cuts.)
| Render | File | Cuts found / planned | Unplanned | Median offset | Max offset |
|---|---|---|---|---|---|
| Rxbyn, first run | 16:9 | 20 / 21 | 0 | 9.3 ms | 16.3 ms |
| Rxbyn, first run | 9:16 | 21 / 21 | 0 | 10.0 ms | 16.3 ms |
| Rxbyn, recorded demo run | 16:9 | 17 / 20 | 0 | 7.7 ms | 15.7 ms |
| Rxbyn, recorded demo run | 9:16 | 20 / 20 | 0 | 8.1 ms | 16.3 ms |
| Rxbyn, self-host check (fresh containers) | 9:16 | 13 / 20 | 0 | 7.0 ms | 16.3 ms |
Cuts that are not found are joins between two similar shots (or a clip and its mirrored reuse), which the pixel test cannot see; they are planned on the same grid. These offsets are against the detected beat grid, not a human one; the tracker's own error (section 2) adds to them. One more caveat found in the self-host check: the containerised analyzer placed the Rxbyn beat grid about 160 ms (a third of a beat) later than the analyzer on the host, with the same tempo and library versions; the decoders differ. Both grids are self-consistent, but at most one sits on the true beats. Every render measures its own files and puts the numbers in the run and the signed record.
4. Treatment grounding
The production treatment prompt wrote scenes for each of the 9 songs' first 40 lines (sections of 4 lines, one scene each; lyrics only), then the production checks ran. A planted negative: each scene again with its citations swapped for lines from a different song (anchor removed).
| Dev (3 songs) | Test (6 songs) | |
|---|---|---|
| Scenes written | 20 (one song's reply did not parse; production retries once) | 52 |
| Citations valid (real lines, in the scene's section) | 100% | 100% |
| Anchor words found in the cited lines | 100% | 96.2% |
| Prompt screen flags | 0 | 0 |
| Typed check says yes to the scene as written | 100% | 100% |
| Typed check says no to the swapped scene | 65% | 67.3% |
The typed yes/no check alone lets a third of plainly mismatched scenes through; the code checks (citation, anchor words copied from the cited lines) are what make the grounding hold, and the typed check is a second opinion. Observation, not tuned: the swapped scenes that passed all got p(yes) = 0.94, while 57% of the real ones got 0.997; a stricter threshold might help, but it would have to be set on other data. The check says whether a scene relates to its lines, not whether it is good.
5. Render time per minute of video
Wan2.2-VACE-Fun-A14B, fp8, 4 Lightning steps, 832x480 / 480x832, 81 frames, one job at a time on the shared GPU0 (RTX PRO 6000 96 GB, other services and a sibling's renders on the same card):
| Run | Formats | Clips | Median per clip | Total job | Video | GPU seconds per minute of video |
|---|---|---|---|---|---|---|
| Rxbyn, first full run | 16:9 + 9:16 | 14 | 84 s | 1,279 s | 50 s | 1,535 |
| Rxbyn, second run | 16:9 + 9:16 | 14 | 96 s | 1,449 s | 50 s | 1,739 |
| Rxbyn, recorded demo | 16:9 + 9:16 | 12 | 141 s* | 1,359 s | 50 s | 1,631 |
| Cortez, recorded demo | 9:16 | 7 | 87 s | 628 s | 55 s | 685 |
| Paper Lanterns (Music3), recorded demo | 9:16 | 7 | 84 s | 663 s | 58.5 s | 679 |
| Self-host check | 9:16 | 6 | 141 s* | 860 s | 50 s | 1,032* |
| Cortez, rendered from the console (browser test) | 9:16 | 7 | 72 s | 520 s | 55 s | 567 |
* interleaved with another of our jobs on the same ComfyUI, so these include waiting. The total job covers the clips, the edit (7-11 s per format on CPU) and stamping. So: about 9.5-11.5 GPU-minutes per minute of video for one format and 23-29 for both. The analysis and treatment took 13-32 s per 50-60 s excerpt on the shared gateway. (The first run's files came out empty: ffmpeg could not seek the Ogg Opus sample and wrote an empty MP4 with exit code 0. The edit now trims the audio with a filter and refuses an incomplete file; clips are kept until the files pass.)
6. Cost
Model calls per analysis: about 8 (clearance lyric labels when any, one treatment, one check per scene), 2,400-2,700 prompt and 900-1,150 completion tokens: $0.002-0.0025 at the gateway's list price ($0.30/$1.50 per million). Rendering is our own GPU; priced at the list price of renting the same card (RunPod RTX PRO 6000 community cloud, $1.69/h, checked 2026-09-26): one format is about $0.32 per minute of video (a 55-60 s 9:16 video cost $0.29-0.31), both formats $0.60-0.68 for 50 s. The hosted demo does not bill renders. At that rate a 3.5-minute song would be about $1.10 in one format and $2.30 in both (extrapolated from the per-minute figure, not rendered: the hosted demo caps a video at 60 s). So "a first music video for the price of a coffee" holds at rental prices, for GPU time; it leaves out electricity margins, storage and anyone's time.
Expected properties of the sample run (for the rehearsal kit)
For rehearsal/music-video-studio (the Rxbyn CC BY excerpt, 50 s):
- No rights statement: 403.
- Tempo between 100 and 125 BPM (the tracker finds 112.35).
- All 18 sung lines placed; line 1 within 0.5 s of 0.78 s and line 9 within 0.5 s of 18.62 s (the human timings).
- Every scene cites lines and passes the code checks (verdict grounded, weak or instrumental).
- The plan is renderable with at least 10 cuts on bar lines.
- Every model call receipted and signed.
Not measured
- Visual quality against closed models or MiniMax H3: not scored. Side by side, the 480p upscaled clips are clearly softer and less coherent than the H3 samples on our site.
- Beat tracking on real full mixes with ground truth (no openly licensed beat-annotated full mixes were at hand).
- Languages other than English (the aligner is English-only).
- Same-seed render determinism for VACE.