171 · Music · Creative and media · preview
Music video starring you
Eval results
Not held outRun 29 Sep 2026Eval write-up (decosa-api, access required)
- Planned cuts found on the beat in the exports27 / 27syntheticn = 273 storyboard renders of the sample song; median 5.7 ms, max 16.3 ms from the beat grid (half a frame at 30 fps is 16.7 ms).
- Storyboard frames flagged by the frame safety check0 / 30syntheticn = 30Sample performers are synthetic adults; the check fails closed.
- Storyboard video done (p50)88.2 ssyntheticn = 386-105 s (p50 88.2 s); first picture drawn after 12-20 s; shared GPU.
- Cost per storyboard video$0.029-0.034syntheticn = 3Text model at list price plus GPU time at $1.32/h.
- MiniMax H3 time per moving shot57-65 s (one reference)syntheticn = 10Sketch tier 864x480, fp8, one 96 GB card, peak 49.6-52 GiB; recorded in a GPU window on 29 Sep.
- MiniMax H3 cost per shot$0.022-0.029syntheticn = 27GPU time at $1.32/h; 27 shots including 6 re-rolls.
- Face likeness to the consent clip (ArcFace cosine, storyboard / H3)0.52 / 0.36syntheticInternal QC with a non-commercial model, not shipped. H3 shots were often wide or turned away (faces found in 15 of 80 sampled frames); frontal re-rolls reached 0.51.
- Moving shots that kept the face (likeness gate)5 / 10syntheticn = 10One middle frame per moving shot checked against each person's reference; the rest fell back to the still. Some wrong faces still pass.
- Blind artist: posts it / would pay (storyboard; moving)yes, $20; no, $0syntheticn = 1An Opus sub-agent as an independent artist: the storyboard as a teaser and Canvas, not as the music video; the moving shots lost her face too often. Before the likeness gate on moving shots.
Dataset
Synthetic sample performers (consent clips made from designed faces and voices; no real person), one sample song made with MiniMax-Music3, 3 storyboard renders per mode, and 27 MiniMax H3 shots (10 artist, 11 couple, 6 re-rolls) rendered in a GPU window on 29 Sep 2026.
Caveats
- Synthetic performers and one song; real selfies and real tracks are not measured.
- Beat offsets are against the detected beat grid, not a human one.
- Visual quality is not scored; some H3 sketch shots show artefacts (gold squiggles in one look, a banding glitch).
- ArcFace (buffalo_l) is non-commercial and used only as internal QC.
- The same builder wrote the pipeline and the eval.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 29 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 88 s
- Receipts
- 5
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.029
Self-host verification
Not yet verified on a fresh self-host setup.
Rehearsal bundle: music-video-starring-you.zip (1 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Moving shots need a render GPU, which is off until launch: the hosted demo makes a storyboard (drawn stills with camera moves).
- No lip-sync to the vocal.
- Likeness is checked by the text model on each frame; ArcFace numbers are internal QC only.
- The adult check is a vision estimate from the consent frames, not an ID check.
- Moving H3 shots aren't ready: blind testers rejected them (faces lost or someone else's); the likeness gate keeps about half and still lets some wrong faces through.
- The likeness check is a yes/no from the text model on each picture; it caught 1 unlike picture in the couple test after the fix, and missed one before it (only the first partner was checked).
- Measured with synthetic sample performers and one sample song.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Consent read-back: hears whether the clip says the sentence and its three fresh wordsQwen3-ASR-1.7B (language pack speech service)Apache-2.0
- Face detection only (boxes): exactly one face, present and moving through the clip; its three sharpest frames become the only face referencesUltra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)MIT
- Adult check on the consent frames (vision), the shot list per song section, and the frame safety and likeness checksQwen3.8-27B (NVFP4)Apache-2.0
- Beat grid, bars and sections of the song (CPU; the music-video studio's analyzer): cuts land on bar downbeatsdecosa-mvideo-analyze (services/mvideo)AGPL-3.0-or-later (decosa-api)
- Storyboard: one still per shot drawn from the consent-clip references, then a camera move (push, pan, drift)FLUX.2 klein 4BApache-2.0
- The edit (CPU): shots cut on the bar downbeats at 30 fps, the AI video label on every frame, an end card crediting the music, exports in 16:9, 9:16 and a Spotify Canvas loop; cut timing measured back from the pixels; C2PA per filedecosa-api starring module + FFmpeg + c2pa-pythonAGPL-3.0-or-later
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Standard · the hosted demo, a storyboard cut on the beat (3)
- Planned cuts found on the beat (3 renders): 27 / 27; median 5.7 ms, max 16.3 msdecosa-api docs/evals/music-video-starring-you.md, 2026-09-29
- Frames flagged by the frame safety check: 0 / 30decosa-api docs/evals/music-video-starring-you.md, 2026-09-29
- Visual quality: not scored; drawn stills with camera moves
Best · moving shots on MiniMax H3, one 96 GB card (3)
- Time per moving shot (sketch 864x480): 57-65 s with one referencedecosa-api docs/evals/music-video-starring-you.md, 2026-09-29
- Face likeness to the consent clip (ArcFace, internal QC): 0.36 (frontal re-rolls 0.51)decosa-api docs/evals/music-video-starring-you.md, 2026-09-29; faces often small or turned
- Blind testers who would pay for moving shots: 0 of 2 (both would for the storyboard)decosa-api docs/evals/music-video-starring-you.md, 2026-09-29