Skip to content
decosa

Make a music video starring you

Shots of you from a 15-second consent clip, planned per section of your song and cut on the beat, in 16:9, 9:16 and a Canvas loop.

Moving MiniMax H3 shots need a render GPU, which is switched on at launch. Until then each shot is a drawn still with a camera move, cut on the beat.

For
Independent artists releasing a single
Instead of
A first video is a static visualiser or a shoot that costs a thousand dollars or more.
  • Time per task88 stypical (median) on the sample; slowest 1 in 20: 105 s
  • Cost per task~$0.024 per videomeasured, at list price
  • Accuracy27 / 27Planned cuts found on the beat (synthetic set)All results and caveats
  • Self-hostRun it on your own server: nothing is sent to us or anyone else.
  • HostedUse it on our servers; see where your data goes.

Try it in one click

Juno's video for “Paper Lanterns”

A consent clip, a song and one line of direction: 8 to 12 shots of the performer, cut on the beat, with a tall version and a Spotify Canvas loop.

Juno and Theo are synthetic sample performers (faces and voices made by AI, not people). The sample song was made with AI too.

You

Record your consent clip

A 15-second selfie video reading a sentence that ends with three words Studio picks for you. It is the only place your face comes from: there is no photo upload. Adults only.

What it does, in shortWho it's for, where it runs and the key results

Record a 15-second consent clip (a selfie reading a sentence that ends in three fresh words), add your track and a line of direction, and get 8 to 12 shots of you cut on the beat: the beat grid and sections of the song, a shot list per section you can edit, a re-roll on any shot, a 16:9 and a 9:16 export and a Spotify Canvas loop. Your face comes only from your consent clip: there is no photo upload, and adults only. Every frame says AI video and the end card credits the music to you. Moving shots run on MiniMax H3 when a render GPU is attached; otherwise the shots are a storyboard of drawn stills with camera moves.

In short

Last reviewed

What it is
Record a 15-second consent clip, add your track, and get a music video of you cut on the beat, in 16:9, 9:16 and a Canvas loop.
Who it's for
Independent artists releasing a single who want a first video with themselves in it.
Where it runs
Hosted (moving shots when a render GPU is attached); self-host on a 96 GB card with your own H3 licence
Key numbers
  • 27 / 27 Planned cuts found on the beat in the exports (synthetic, n = 27)
  • 0 / 30 Storyboard frames flagged by the frame safety check (synthetic, n = 30)
  • 88.2 s Storyboard video done (p50) (synthetic, n = 3)
  • 88.2 s Median end-to-end run, hosted (QA sweep 2026-09-29)
All results, datasets and caveats
How we tested itEnd-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 29 Sep 2026 · measured 29 Sep 2026: · p50 88 s · p95 105 s (3 runs) · ~$0.029 per run · 5 receipts

Loading the nightly status…

Self-host: not yet verified

Measured cost to run: about $0.024 per video (hosted, 29 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Known limits (7)
  • Moving shots need a render GPU, which is off until launch: the hosted demo makes a storyboard (drawn stills with camera moves).
  • No lip-sync to the vocal.
  • Likeness is checked by the text model on each frame; ArcFace numbers are internal QC only.
  • The adult check is a vision estimate from the consent frames, not an ID check.
  • Moving H3 shots aren't ready: blind testers rejected them (faces lost or someone else's); the likeness gate keeps about half and still lets some wrong faces through.
  • The likeness check is a yes/no from the text model on each picture; it caught 1 unlike picture in the couple test after the fix, and missed one before it (only the first partner was checked).
  • Measured with synthetic sample performers and one sample song.

Eval results, nightly checks and cost per run · eval not held out

Technical detailsModels, where it runs, labels, what it is built from
Models
MiniMax H3 (reference mode, Turbo v4) · FLUX.2 klein 4B (storyboard) · Qwen3.8-27B (shot list, checks) · Qwen3-ASR-1.7B (consent read-back)
Where
Hosted (moving shots when a render GPU is attached); self-host on a 96 GB card with your own H3 licence
Checks
Consent clip checked (words heard, one live face, an adult); frame safety and likeness checks; cuts measured on the beat; C2PA in every export
Output
Media
Data
Personal data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Every result carries a signed record of which model produced it, so you can check it later. How that works

Questions people ask

Can I upload a photo instead of recording a clip?

No. Your face comes only from a live 15-second consent clip that reads three fresh words, so nobody can make a video of someone else. Adults only.

Are the shots moving video?

Not yet. The hosted tool makes a storyboard: a drawn still of you per shot with a camera move, cut on the beat. MiniMax H3 moving shots (57-65 s a shot, measured) lost the face too often in a blind test, so they stay off until they hold it.

How well does it cut on the beat?

In 3 renders every planned cut was found in the file (27 / 27), median 5.7 ms and at most 16.3 ms from the beat, under half a frame at 30 fps.

Is it labelled as AI?

Yes: an AI video label on every frame, an end card crediting the music to you, and a C2PA credential in every file. Use each platform's own AI toggle as well.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Music video starring you

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.