Eval: music video starring you (171) and our story film (172)
Run on our server, 2026-09-29, against a pre-release test instance of this branch (text model on the direct route, priced
at list; storyboard stills on the shared GPU0). MiniMax H3 shots were rendered in a GPU window on GPU0 the same morning
(09:10-09:53, Voxtral and Hy-MT2 paused and restored, logged) and later replayed through the product's H3 path. Raw
results in docs/evals/music-video-starring-you/ (storyboard_runs.json, h3_move.json). Everything below was
measured. The same builder wrote the pipeline and this eval; there is no held-out split.
Data and licences
| Set | What | Licence | Used for |
|---|---|---|---|
| Sample performers Juno and Theo | consent clips made from designed faces and voices; no real person | ours (synthetic) | every run |
| "Paper Lanterns" | a sample song made with MiniMax-Music3 through music-gen-cleared (37), 95.7 BPM | ours, with its licence certificate | every run |
1. Consent clip checks (tests)
tests/test_studio_starring.py: a still picture held to the camera is not a live face; no face fails; the age check
fails closed; a used or expired sentence is refused; a minor can't be enrolled for the music_video purpose; memories
that describe a child are refused; a performer's own link deletes their videos; media is refused once consent is gone.
All pass (60 targeted tests with the family film and the studio packs). The sample performers pass the live checks
(sentence heard 0.955, one moving face).
2. Storyboard runs (the hosted path today)
Three runs per mode, sample performers, sample song, look cinematic, 16:9 plus 9:16 and the Canvas loop.
| Run | Mode | Shots | First shot | Done | Cuts found / planned | Median / max off the beat | Frames flagged | Cost |
|---|---|---|---|---|---|---|---|---|
| 1 | artist | 10 | 20.1 s | 104.7 s | 9 / 9 | 5.7 / 16.3 ms | 0 / 10 | $0.0342 |
| 2 | artist | 10 | 11.9 s | 86.5 s | 9 / 9 | 5.7 / 16.3 ms | 0 / 10 | $0.0287 |
| 3 | artist | 10 | 12.1 s | 88.2 s | 9 / 9 | 5.7 / 16.3 ms | 0 / 10 | $0.0291 |
| 4 | couple | 11 | 17.0 s | 274.9 s | 10 / 10 | 5.2 / 16.0 ms | 0 / 11 | $0.0739 |
| 5 | couple | 11 | 15.5 s | 142.5 s | 10 / 10 | 5.2 / 16.0 ms | 0 / 11 | $0.0399 |
| 6 | couple | 11 | 14.9 s | 122.0 s | 10 / 10 | 5.2 / 16.0 ms | 0 / 11 | $0.0405 |
- p50 done: artist 88.2 s, couple 142.5 s. Run 4 waited behind other jobs on the shared card (191 GPU-seconds against 96-100 in runs 5-6).
- Cuts are found in the delivered 16:9 file's own pixels (isolated FFmpeg scene-score spikes) and compared with the analyzer's beat grid, not a human one. Half a frame at 30 fps is 16.7 ms, so every cut sits on the frame nearest the beat. The same song and plan give the same offsets in every run.
- 5 receipts per video (text model and ASR); cost is the text model at $0.30/$1.50 per million tokens plus GPU time at $1.32/h (a shared card, so an upper bound) plus CPU at $0.10/h.
3. MiniMax H3 shots (GPU window, then the product path)
27 shots on GPU0: 10 artist, 11 couple, 6 re-rolls. Turbo v4 LoRA, sketch tier (864x480, 4 forwards), fp8 rowwise transformer, SDPA, no compile; the text encoder (fp8 weights, 37 GiB) encoded all prompts first and was unloaded before the transformer loaded.
| Value | |
|---|---|
| Time per shot | 57-65 s with one reference (transformer 47-55 s); about 80 s with two |
| Peak GPU memory | 49.6-52 GiB |
| Load | 44 s text encoder, about 80 s transformer |
| GPU cost | $0.022-0.029 a shot at $1.32/h; $0.24 for the 10-shot artist video, $0.33 for the 11-shot couple film |
Then POST /studio/videos/{vid}/move on the two videos those shot lists came from, with a replay worker serving the
recorded clips over the real worker protocol (labelled a replay): both videos finished in 27 s. Artist: 8 of 9 cuts
found (one cut between two similar dusk shots is not detectable), median 6.9 ms, max 16.3 ms. Couple: 10 of 10, median
5.2 ms, max 16.0 ms. The every-frame gate checked 3 frames of every moving shot (30 and 33 frames): 0 flagged. Credits
name MiniMax H3. Visible artefacts on some sketch shots: gold squiggles in the super 8 look (the paint scene) and one
banding glitch.
4. Likeness (internal QC only)
ArcFace (insightface buffalo_l, a non-commercial model, so internal measurement only, never shipped) cosine between
each performer's consent-clip reference and faces found in the output:
| Artist (Juno) | Couple (Juno / Theo) | |
|---|---|---|
| Storyboard stills | 0.52 | 0.43 (0.40 / 0.45) |
| H3 shots | 0.36 | 0.40 |
| H3 re-rolls (frontal prompts) | 0.51 |
H3 shots were often wide or turned away: faces were found in only 15 of 80 sampled artist frames. Page 74's 0.65-0.68 was a frontal prompt. The product's own likeness check is the text model's yes/no per frame: 0 "unlike" in the 63 frames of the six storyboard runs, 1 in the 10 storyboard frames of the window's artist video.
5. Blind tests (sub-agents) and what changed
A blind Claude Code sub-agent (Opus 5.5) played two buyers from frame sheets and shot lists only
(music-video-starring-you/blind_verdict.json): Juno, an independent artist, on the storyboard and H3 artist videos;
Theo, buying an anniversary gift, on the storyboard and H3 couple films.
| Storyboard | Moving (H3, before the gate) | |
|---|---|---|
| Juno (artist) | would post it as a teaser and Canvas, $20 | would not post it, $0 ("her face is never clearly on screen", melting shots, a different outfit on a beach) |
| Theo (partner) | would give it after one fix, $29 | would not give it, $0 (strangers in shot 1, two men on the train, melting paint faces) |
Both said one wrong face is worse than no video, and both wanted a preview and free single-shot re-rolls. Changes:
- Couples likeness bug, fixed: the likeness check compared only the first performer; shot 9 of a couple storyboard gave Theo long locs and another face, and passed. Each performer is now compared with their own reference in the shots they're in, and the prompt names hair length, texture and facial hair (test added). On the next couple run the check caught 1 unlike picture, redrew it, and marked it for re-roll.
- Moving shots get a likeness gate: the middle frame of every H3 shot is checked against each person's reference; an unlike shot falls back to its checked still. On the window's videos, 5 of 10 artist and 6 of 11 couple moving shots fell back. Some still get through (the couple opening shot shows a stranger: one frame per shot is checked).
- End card: "performers" for one person is now "a synthetic sample performer"; the doubled "(sample song)" was fixed earlier.
- Not changed: the "AI video" label on every frame (both disliked it; it stays: the disclosure is the point), lyric-free memory captions cut into fragments in the H3 couple film, light that jumps between dusk and day across shots.
Verdict: the storyboard is the product that sells today (a keepsake or a teaser); moving shots are not ready.
Fixed during this eval
- The frame check ran out of tokens at 8 pictures and failed closed (shared
packs/safety.py: max(260, 48n+60)). - The sample song's credit read "by Decosa Studio (sample song) (a sample song made with AI)": the sample track's artist is now "Decosa Studio".
- "Famous" false positives on the consented performer's own photoreal shots; borrowed-picture chains between shots.
Expected properties of the sample run (for the rehearsal kit)
Sample performers (artist) + the sample song, video: stills:
- a video that asks for
h3-sketchwhile H3 is off is refused (409, codeh3_off); - the video reaches
donewith 10 shots; - every planned cut is found in the file (
sync.found=sync.planned), each within 17 ms of the beat; - no frame is flagged (
safety[0].flagged= 0); - the 16:9, 9:16 and canvas exports each carry
credential: "c2pa"; - the credits name the music and say the video is AI.
Not measured
- Real selfies (lighting, glasses, beards, darker rooms) and real tracks; lip-sync (not built).
- Visual quality, and whether people would post the storyboard cut.
- A warm H3 worker on a 96 GB card (the window used encode-first, then the transformer).