Make a music video starring you
Record a 15-second consent clip and add your track. Get 8 to 12 shots of you, planned per section of the song and cut on the beat, in 16:9 and 9:16, with a Spotify Canvas loop. Change or re-roll any shot.
About 2 minutes to a storyboard cut of a 40-second video (measured)Median 5-13 ms from the beat for every cut (measured on the delivered files)About 3 cents a storyboard video at list pricesDetails, API and self-host
Try it in one click
Juno's video for “Paper Lanterns”
A consent clip, a song and one line of direction: 8 to 12 shots of the performer, cut on the beat, with a tall version and a Spotify Canvas loop.
Juno and Theo are synthetic sample performers (faces and voices made by AI, not people). The sample song was made with AI too.
You
Record your consent clip
A 15-second selfie video reading a sentence that ends with three words Studio picks for you. It is the only place your face comes from: there is no photo upload. Adults only.
The rules
- Your face comes only from your own live consent clip: a selfie video reading a sentence that ends with three words Studio picks for you. There is no photo upload.
- Adults only. The clip must show one live, moving face and an adult; anything that could be someone under 18 stops, and the clip is deleted.
- Every frame says “AI video”, the end card credits the music to the artist, and every file carries a content credential.
- Every picture passes a safety check (no children, nudity or gore; it fails closed) and a likeness check.
- Private: share with people you name, or download it and post it yourself. Your own link takes your consent back and deletes your clip and every video with your face.
- No famous people, brands or franchise characters. The performer doesn't lip-sync yet: shots are cinematic, not singing.
What H3 adds, and what it costs
Moving shots are not ready yet. Recorded on 29 Sep 2026 in a render window on our own 96 GB card: the same shot lists as the storyboards, each shot a MiniMax H3 clip from the sample performers' consent-clip frames. In a blind test, an artist and a partner both said they'd pay for the storyboard and not for these moving shots: too often the face was lost, turned away or someone else's. The tool now keeps a moving shot only when it passes the likeness check (5 of 10 and 6 of 11 fell back to the still), and some wrong faces still get through. The performers are synthetic adults.
- Time per moving shot
- 57-65 s (one face), about 80 s (two)
- Peak GPU memory
- 49.6-52 GiB (fp8)
- GPU cost per shot
- $0.022-0.029 at $1.32/h
- Moving shots that kept the face (likeness gate)
- 5 of 10 and 5 of 11
- Cuts found on the beat
- 8 of 9 and 10 of 10; median 5.2-6.9 ms
- Blind testers who'd pay for the moving version
- 0 of 2 (both would for the storyboard)
Recorded, not live: the hosted tool makes a storyboard until a render GPU is attached. MiniMax H3 is used under our commercial licence.
Models and where it runs
Moving shots: MiniMax H3 in reference mode with the Turbo v4 adapter (Decosa's MiniMax licence), when a render GPU is attached. Storyboard: FLUX.2 klein 4B (Apache-2.0) from your consent-clip references. Shot list and checks: Qwen3.8-27B. Consent read-back: Qwen3-ASR-1.7B. Face detection (boxes only): UltraFace (MIT). Beats and sections: the music-video studio's analyzer on CPU. Hosted on Decosa's GPUs; nothing goes to a third-party video service.