Skip to content
decosa

171 · Music · Creative and media · preview

Music video starring you

Open the toolJSON

Eval results

Not held outRun 29 Sep 2026Eval write-up (decosa-api, access required)

  • Planned cuts found on the beat in the exports27 / 27syntheticn = 273 storyboard renders of the sample song; median 5.7 ms, max 16.3 ms from the beat grid (half a frame at 30 fps is 16.7 ms).
  • Storyboard frames flagged by the frame safety check0 / 30syntheticn = 30Sample performers are synthetic adults; the check fails closed.
  • Storyboard video done (p50)88.2 ssyntheticn = 386-105 s (p50 88.2 s); first picture drawn after 12-20 s; shared GPU.
  • Cost per storyboard video$0.029-0.034syntheticn = 3Text model at list price plus GPU time at $1.32/h.
  • MiniMax H3 time per moving shot57-65 s (one reference)syntheticn = 10Sketch tier 864x480, fp8, one 96 GB card, peak 49.6-52 GiB; recorded in a GPU window on 29 Sep.
  • MiniMax H3 cost per shot$0.022-0.029syntheticn = 27GPU time at $1.32/h; 27 shots including 6 re-rolls.
  • Face likeness to the consent clip (ArcFace cosine, storyboard / H3)0.52 / 0.36syntheticInternal QC with a non-commercial model, not shipped. H3 shots were often wide or turned away (faces found in 15 of 80 sampled frames); frontal re-rolls reached 0.51.
  • Moving shots that kept the face (likeness gate)5 / 10syntheticn = 10One middle frame per moving shot checked against each person's reference; the rest fell back to the still. Some wrong faces still pass.
  • Blind artist: posts it / would pay (storyboard; moving)yes, $20; no, $0syntheticn = 1An Opus sub-agent as an independent artist: the storyboard as a teaser and Canvas, not as the music video; the moving shots lost her face too often. Before the likeness gate on moving shots.

Dataset

Synthetic sample performers (consent clips made from designed faces and voices; no real person), one sample song made with MiniMax-Music3, 3 storyboard renders per mode, and 27 MiniMax H3 shots (10 artist, 11 couple, 6 re-rolls) rendered in a GPU window on 29 Sep 2026.

Caveats

  • Synthetic performers and one song; real selfies and real tracks are not measured.
  • Beat offsets are against the detected beat grid, not a human one.
  • Visual quality is not scored; some H3 sketch shots show artefacts (gold squiggles in one look, a banding glitch).
  • ArcFace (buffalo_l) is non-commercial and used only as internal QC.
  • The same builder wrote the pipeline and the eval.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
29 Sep 2026
Latency, this run
n/a
p50 over passed runs
88 s
Receipts
5
Model calls
n/a
Tokens
n/a
Cost per run
$0.029

Self-host verification

Not yet verified on a fresh self-host setup.

Rehearsal bundle: music-video-starring-you.zip (1 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Moving shots need a render GPU, which is off until launch: the hosted demo makes a storyboard (drawn stills with camera moves).
  • No lip-sync to the vocal.
  • Likeness is checked by the text model on each frame; ArcFace numbers are internal QC only.
  • The adult check is a vision estimate from the consent frames, not an ID check.
  • Moving H3 shots aren't ready: blind testers rejected them (faces lost or someone else's); the likeness gate keeps about half and still lets some wrong faces through.
  • The likeness check is a yes/no from the text model on each picture; it caught 1 unlike picture in the couple test after the fix, and missed one before it (only the first partner was checked).
  • Measured with synthetic sample performers and one sample song.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Consent read-back: hears whether the clip says the sentence and its three fresh wordsQwen3-ASR-1.7B (language pack speech service)Apache-2.0
  • Face detection only (boxes): exactly one face, present and moving through the clip; its three sharpest frames become the only face referencesUltra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)MIT
  • Adult check on the consent frames (vision), the shot list per song section, and the frame safety and likeness checksQwen3.8-27B (NVFP4)Apache-2.0
  • Beat grid, bars and sections of the song (CPU; the music-video studio's analyzer): cuts land on bar downbeatsdecosa-mvideo-analyze (services/mvideo)AGPL-3.0-or-later (decosa-api)
  • Storyboard: one still per shot drawn from the consent-clip references, then a camera move (push, pan, drift)FLUX.2 klein 4BApache-2.0
  • The edit (CPU): shots cut on the bar downbeats at 30 fps, the AI video label on every frame, an end card crediting the music, exports in 16:9, 9:16 and a Spotify Canvas loop; cut timing measured back from the pixels; C2PA per filedecosa-api starring module + FFmpeg + c2pa-pythonAGPL-3.0-or-later

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Standard · the hosted demo, a storyboard cut on the beat (3)
  • Planned cuts found on the beat (3 renders): 27 / 27; median 5.7 ms, max 16.3 msdecosa-api docs/evals/music-video-starring-you.md, 2026-09-29
  • Frames flagged by the frame safety check: 0 / 30decosa-api docs/evals/music-video-starring-you.md, 2026-09-29
  • Visual quality: not scored; drawn stills with camera moves
Best · moving shots on MiniMax H3, one 96 GB card (3)
  • Time per moving shot (sketch 864x480): 57-65 s with one referencedecosa-api docs/evals/music-video-starring-you.md, 2026-09-29
  • Face likeness to the consent clip (ArcFace, internal QC): 0.36 (frontal re-rolls 0.51)decosa-api docs/evals/music-video-starring-you.md, 2026-09-29; faces often small or turned
  • Blind testers who would pay for moving shots: 0 of 2 (both would for the storyboard)decosa-api docs/evals/music-video-starring-you.md, 2026-09-29

How we measure · All tools