Make a music video starring you
Shots of you from a 15-second consent clip, planned per section of your song and cut on the beat, in 16:9, 9:16 and a Canvas loop.
Moving MiniMax H3 shots need a render GPU, which is switched on at launch. Until then each shot is a drawn still with a camera move, cut on the beat.
- For
- Independent artists releasing a single
- Instead of
- A first video is a static visualiser or a shoot that costs a thousand dollars or more.
- Time per task88 stypical (median) on the sample; slowest 1 in 20: 105 s
- Cost per task~$0.024 per videomeasured, at list price
- Accuracy27 / 27Planned cuts found on the beat (synthetic set)All results and caveats
Try it in one click
Juno's video for “Paper Lanterns”
A consent clip, a song and one line of direction: 8 to 12 shots of the performer, cut on the beat, with a tall version and a Spotify Canvas loop.
Juno and Theo are synthetic sample performers (faces and voices made by AI, not people). The sample song was made with AI too.
You
Record your consent clip
A 15-second selfie video reading a sentence that ends with three words Studio picks for you. It is the only place your face comes from: there is no photo upload. Adults only.
Watch
Recorded sessions are for the live audio tools. For the studio, the gallery under Try it shows finished renders.
Get an API key
- Call the music video starring you API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1x 96 GB card for MiniMax H3 shots (fp8, about 50 GB); the storyboard runs on FLUX.2 klein (about 6 GB) and the edit on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- music-video-starring-you
Use the hosted API
# Decosa music video starring you: use the hosted API
You are wiring Decosa's "music video starring you" into this project. An artist records a consent clip, adds a track and a line of direction, and gets 8 to 12 shots planned per section of the song, cut on the beat, in 16:9, 9:16 and a Spotify Canvas loop. Faces come only from a live 15-second consent clip that reads
a sentence ending in three fresh words: there is no photo upload, and adults only. Every frame says AI video, the end
card credits the music, and every export carries a C2PA credential. Use only what is listed below. If you need something
else, stop and ask me.
- Base URL: `https://api.decosa.ai`. Health check: `GET https://api.decosa.ai/healthz`.
- Moving MiniMax H3 shots need a render GPU, which is switched on at launch. Until then `video` is `stills`: a drawn still
per shot with a camera move, cut on the beat. Asking for `h3-sketch` while it's off returns 409.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool's page. Keep it in `DECOSA_API_KEY`, never in code.
Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "studio"}` returns `{"token", ...}` (30 minutes).
3. Every tool call works inside a Studio project: `POST /studio/projects {"name": "..."}` returns `project.id` and a
`project_key`. Send the key as `?key=` on every project URL (the API's CORS allows only two headers). Keep it private:
the project key alone opens the project.
## Endpoints (all under a project; `?key=<project_key>`)
- `GET /studio/starring/info`: looks, limits, models, whether H3 is attached (no token).
- `POST /studio/projects/{pid}/performers/challenge {"name", "couple"?}` → `{nonce, sentence, words, expires_in_s}`; record the
person reading it, then `POST /studio/projects/{pid}/performers/clip {"nonce", "clip_b64"}` (webm or mp4, about 15 s).
Refused: no face, two faces, a still picture, the wrong words, an expired sentence, anyone who isn't an adult.
`POST /studio/projects/{pid}/performers/sample {"which": "artist"|"couple"}`: synthetic sample performers.
- `POST /studio/projects/{pid}/tracks {"audio_b64", "title", "artist", "rights": true}` (up to 20 MB) or `{"sample": "paper-lanterns"}`.
- `POST /studio/projects/{pid}/videos {"mode": "artist", "performers": [ids], "track_id", "look", "aspect": "16:9", "direction": "a line of direction"}`
- `GET /studio/videos/{vid}`: poll every 3-5 s until `status` is `done`: `shots[]`, `sync` (cuts measured on the beat),
`safety`, `credits`, `exports` (`16:9`, `9:16`, `canvas`: `url` relative to the base, `credential`).
- `POST /studio/videos/{vid}/shots {"n", "picture"?}`: change a shot's picture, or re-roll it (no picture); re-cuts.
- `POST /studio/videos/{vid}/move {"video": "h3-sketch"}`: the same shots as MiniMax H3 clips (when attached).
- `POST /studio/videos/{vid}/share` / `unshare`; `DELETE /studio/videos/{vid}`. `POST /studio/starring/revoke {"token"}` (the
performer's own link) deletes their clip, references and videos.
## Honest limits
- Measured with synthetic sample performers and one sample song: every planned cut found within 17 ms of the beat in 3
storyboard renders; nothing flagged by the frame check.
- The storyboard is drawn stills with camera moves, not moving video. No lip-sync to the vocal.
- The adult check is a vision estimate from the consent frames, not an ID check.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa music video starring you: run it yourself (containers)
You are setting up Decosa's "music video starring you" on this machine: a consent-clip check (Qwen3-ASR-1.7B, a face detector and
Qwen3.8-27B), the music-video analyzer (CPU), a shot list from Qwen3.8-27B, storyboard stills from FLUX.2 klein 4B
through ComfyUI, optional MiniMax H3 moving shots, and an FFmpeg edit with the AI label and a C2PA credential. Clips,
tracks and videos stay here; nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images or models.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/music-video-starring-you.zip (1 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py music-video-starring-you` (the api image carries the same bundle under /app/rehearsal/music-video-starring-you/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py music-video-starring-you --bundle music-video-starring-you.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "moving H3 shots are refused while no render GPU is attached", "the video finishes", "the shot list has 10 shots"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Licences first
Qwen3.8-27B, Qwen3-ASR-1.7B and FLUX.2 klein 4B are Apache-2.0; the face detector is MIT; librosa ISC; FFmpeg LGPL/GPL;
ComfyUI GPL-3.0. MiniMax H3 needs a licence that covers you (the public community licence excludes the US, EU, UK and
Korea); without one, leave the H3 worker out and the tool makes storyboards.
## Steps
1. Docker and the NVIDIA container toolkit: if `docker compose version` or
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi` fails, install them from the official
instructions. Storyboards need about 6 GB of GPU plus the text model (about 20 GB, or an existing server); H3 needs a
96 GB card with about 60 GB free (peak 49.6-52 GiB measured, fp8).
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the
`llm`, `comfyui`, `speech`, `mvideo-analyze` and `api` services. For `api` set `DECOSA_LLM_ROUTE=direct`,
`DECOSA_STUDIO_WORKER=command`, `DECOSA_LANG_SPEECH_URL`, `DECOSA_MVIDEO_ANALYZE_URL`, and, only with an H3 licence,
`DECOSA_STUDIO_H3_FAST=1`, `DECOSA_STUDIO_H3_URL` and `DECOSA_STUDIO_H3_TOKEN`; use named volumes; bind every port to
127.0.0.1.
3. `docker compose pull && docker compose up -d`; wait for the health checks.
4. Smoke test: a token from `POST /demo/session {"vertical":"studio"}`, a project, `performers/sample {"which":"artist"}`,
`tracks {"sample":"paper-lanterns"}`, then `POST /studio/projects/{pid}/videos` and poll `GET /studio/videos/{vid}` to
`done`: expect every planned cut found (`sync.found` = `sync.planned`), nothing flagged, and `c2pa` on every export.
5. Report back: the signing key id (`GET /attest/signing-key`), the time to the first shot and to the finished video.
Off. Renders never go to the network; only the text model could, and only if I say yes.
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3-ASR-1.7B (language pack speech service) needs a GPU.
- GeForce RTX 4090Doesn't fit
FLUX.2 klein 4B: This build is NVIDIA NVFP4, which needs a Blackwell GPU. No replacement is listed.
- GeForce RTX 5090Doesn't fit
Needs about 39.1 GB of GPU memory at the smallest settings; 32 GB available.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- L40SDoesn't fit
FLUX.2 klein 4B: This build is NVIDIA NVFP4, which needs a Blackwell GPU. No replacement is listed.
- H100 80 GB (SXM)Doesn't fit
FLUX.2 klein 4B: This build is NVIDIA NVFP4, which needs a Blackwell GPU. No replacement is listed.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (68.7 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns
The standard tier fits (68.7 of 192 GB). The best tier fits too.
- Apple M3 Ultra (Mac Studio), 96 GBCan't tell
Memory not known for Qwen3-ASR-1.7B (language pack speech service) has no mapped Apple Silicon build; FLUX.2 klein 4B has no mapped Apple Silicon build.
- Apple M5 Max, 64 GBCan't tell
Memory not known for Qwen3-ASR-1.7B (language pack speech service) has no mapped Apple Silicon build; FLUX.2 klein 4B has no mapped Apple Silicon build.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API, the speech model and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with both models loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "asr": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"music-video-starring-you"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py music-video-starring-you
Download the mock-data bundle (1 KB, 10 checks)expected.json
A synthetic sample performer (an adult, consent clip on the server) and a sample song made with MiniMax-Music3. Asking for moving H3 shots while no render GPU is attached must be refused; the storyboard must finish with 10 shots, every planned cut found on the beat in the file, no frame flagged, and a C2PA credential on every export. About 1.5-2 minutes on a shared GPU.
What the rehearsal checks
- moving H3 shots are refused while no render GPU is attached
- the video finishes
- the shot list has 10 shots
- all 9 planned cuts are found on the beat in the file
- no cut appears that wasn't planned
- the worst cut is within 17 ms of the beat
- no frame is flagged by the safety check
- the 16:9 export carries a C2PA credential
- the credits say the video is AI
- every model call has a signed receipt
Licence: No local inputs. The sample performers are synthetic (designed faces and voices, no real person); the sample song was made with MiniMax-Music3 through music-gen-cleared and carries its licence certificate.
Prompt for your coding agent
# Decosa music video starring you: run it yourself (containers)
You are setting up Decosa's "music video starring you" on this machine: a consent-clip check (Qwen3-ASR-1.7B, a face detector and
Qwen3.8-27B), the music-video analyzer (CPU), a shot list from Qwen3.8-27B, storyboard stills from FLUX.2 klein 4B
through ComfyUI, optional MiniMax H3 moving shots, and an FFmpeg edit with the AI label and a C2PA credential. Clips,
tracks and videos stay here; nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images or models.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/music-video-starring-you.zip (1 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py music-video-starring-you` (the api image carries the same bundle under /app/rehearsal/music-video-starring-you/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py music-video-starring-you --bundle music-video-starring-you.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "moving H3 shots are refused while no render GPU is attached", "the video finishes", "the shot list has 10 shots"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Licences first
Qwen3.8-27B, Qwen3-ASR-1.7B and FLUX.2 klein 4B are Apache-2.0; the face detector is MIT; librosa ISC; FFmpeg LGPL/GPL;
ComfyUI GPL-3.0. MiniMax H3 needs a licence that covers you (the public community licence excludes the US, EU, UK and
Korea); without one, leave the H3 worker out and the tool makes storyboards.
## Steps
1. Docker and the NVIDIA container toolkit: if `docker compose version` or
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi` fails, install them from the official
instructions. Storyboards need about 6 GB of GPU plus the text model (about 20 GB, or an existing server); H3 needs a
96 GB card with about 60 GB free (peak 49.6-52 GiB measured, fp8).
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the
`llm`, `comfyui`, `speech`, `mvideo-analyze` and `api` services. For `api` set `DECOSA_LLM_ROUTE=direct`,
`DECOSA_STUDIO_WORKER=command`, `DECOSA_LANG_SPEECH_URL`, `DECOSA_MVIDEO_ANALYZE_URL`, and, only with an H3 licence,
`DECOSA_STUDIO_H3_FAST=1`, `DECOSA_STUDIO_H3_URL` and `DECOSA_STUDIO_H3_TOKEN`; use named volumes; bind every port to
127.0.0.1.
3. `docker compose pull && docker compose up -d`; wait for the health checks.
4. Smoke test: a token from `POST /demo/session {"vertical":"studio"}`, a project, `performers/sample {"which":"artist"}`,
`tracks {"sample":"paper-lanterns"}`, then `POST /studio/projects/{pid}/videos` and poll `GET /studio/videos/{vid}` to
`done`: expect every planned cut found (`sync.found` = `sync.planned`), nothing flagged, and `c2pa` on every export.
5. Report back: the signing key id (`GET /attest/signing-key`), the time to the first shot and to the finished video.
Off. Renders never go to the network; only the text model could, and only if I say yes.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Doesn't fitMusic video starring you on GeForce RTX 5090
Needs about 39.1 GB of GPU memory at the smallest settings; 32 GB available.
Standard · the hosted demo, a storyboard cut on the beat: what changesuses estimates
- Needs about 39.1 GB of GPU memory at the smallest settings; 32 GB available.
Memory per component
- Consent read-back: Qwen3-ASR-1.7B (language pack speech service). ~5.1 GB, weights 3.4 GB (estimate). Estimate: 1.7B parameters at 2 bytes (BF16) per weight is about 3.4 GB, plus 20% working memory and 1 GB of runtime. Not measured.
- Face detection only: Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Adult check on the consent frames: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
- Beat grid, bars and sections of the song: decosa-mvideo-analyze (services/mvideo). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Storyboard: FLUX.2 klein 4B. ~6 GB, loaded while a job runs (from stack.json). vram_gb 6 in stack.json.
- The edit: decosa-api starring module + FFmpeg + c2pa-python. CPU. Runs on CPU (vram_gb 0 in stack.json).
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Music video starring you, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Music video starring you on my hardware Fetch https://decosa.ai/prompts/music-video-starring-you-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=music-video-starring-you) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · the hosted demo, a storyboard cut on the beat (standard). Fit check: doesn't fit; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Consent read-back: Qwen3-ASR-1.7B (language pack speech service) (Qwen/Qwen3-ASR-1.7B), 5.1 GB - Face detection only: Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320), CPU - Adult check on the consent frames: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB - Beat grid, bars and sections of the song: decosa-mvideo-analyze (services/mvideo), CPU - Storyboard: FLUX.2 klein 4B (black-forest-labs/FLUX.2-klein-4B), 6 GB - The edit: decosa-api starring module + FFmpeg + c2pa-python, CPU Warning: the fit check says this tier does not fit: Needs about 39.1 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further. During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/music-video-starring-you-assemble.md
Get an API key
- Call the music video starring you API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1x 96 GB card for MiniMax H3 shots (fp8, about 50 GB); the storyboard runs on FLUX.2 klein (about 6 GB) and the edit on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Record a 15-second consent clip, add your track, and get a music video of you cut on the beat, in 16:9, 9:16 and a Canvas loop.
An independent artist records a 15-second selfie reading a sentence that ends in three fresh words, adds a track and a line of direction, and gets 8 to 12 shots of themselves planned per section of the song: edit or re-roll any shot, then export 16:9, 9:16 and a Spotify Canvas loop, cut on the beat. The face comes only from the consent clip: there is no photo upload, and adults only. Every frame says AI video and the end card credits the music to the artist.
- Deployment
- Hosted or self-host
- Regulatory
- Not legal advice; written 29 Sep 2026. Consent: a live consent clip with three fresh words, one moving face and an adult check (fails closed) goes into the consent ledger for the purpose music_video only, with the performer's own revocation link, which deletes their references and videos. There is no photo upload anywhere in the tool, so it can't be pointed at someone else's face. Children: Studio keeps children out of every video: a minor can't be enrolled for this purpose (tested), and memories that describe a child are refused. Disclosure: every frame carries a visible AI video label and every export a C2PA credential; EU AI Act (Regulation (EU) 2024/1689) Art. 50(2) asks providers to mark generated video in a machine-readable way from 2 Aug 2026 (EUR-Lex link under Tools). Use each platform's own AI-content toggle as well. Music: you confirm you have the rights to the track, and the end card credits it; the statement is not verified.
Text description
The artist records a 15-second consent clip reading a sentence that ends in three fresh words; Qwen3-ASR-1.7B (Apache-2.0) hears the words, a face detector (UltraFace, MIT) finds one moving face, and Qwen3.8-27B (Apache-2.0) checks for an adult; the consent goes into the ledger for music video only. A CPU analyzer finds the song's beats, bars and sections. Qwen3.8-27B plans the shot list per section. Shots are MiniMax H3 reference-mode clips when a render GPU is attached, otherwise FLUX.2 klein (Apache-2.0) stills from the consent-clip frames with a camera move. Every frame is checked for safety and likeness. FFmpeg cuts on the bar downbeats, adds the AI video label and an end card crediting the music, and exports 16:9, 9:16 and a Canvas loop, each with a C2PA credential. Receipts on every model call. Self-hosted, everything stays on your machine.
At a glance
- Data retention
- The consent clip, face references, track and videos stay in your private Studio project until you delete it or the performer takes consent back with their own link (which deletes their references and videos). The consent entry expires after one year. Logs hold ids, counts and hashes.
- What leaves the box
- On the hosted route, speech, the text model and renders run on Decosa's hosted service. Self-hosted, nothing leaves your machine.
- Whose face
- Only your face, from your own live consent clip, adults only. There is no photo upload anywhere, so it can't be pointed at someone else.
- Cost per video
- A few cents per storyboard video in the eval; a moving MiniMax H3 shot costs a few cents of GPU time.
- Typical time
- Every picture is checked before you see it, so the shots appear together with the storyboard video, a minute or two after you start (measured); each moving H3 shot takes about a minute more.
- Outputs
- MP4s in 16:9 and 9:16 and a Spotify Canvas loop, with an AI video label on every frame, an end card crediting the music and a C2PA credential; private share links you can turn off.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
- In the hosted demo
Standard
the hosted demo, a storyboard cut on the beat
A drawn still per shot from your consent clip, with a camera move, cut on the beat: a planning cut and a keepsake, not yet a moving video.
- Models
- Qwen3-ASR-1.7B (language pack speech service)
- Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)
- Qwen3.8-27B (NVFP4)
- decosa-mvideo-analyze (services/mvideo)
- FLUX.2 klein 4B
- decosa-api starring module + FFmpeg + c2pa-python
- Hardware
- One card with about 6 GB free for the stills, the text model (about 20 GB, or remote), CPU for the edit
- Quality evidence
- Planned cuts found on the beat (3 renders)27 / 27; median 5.7 ms, max 16.3 msdecosa-api docs/evals/music-video-starring-you.md, 2026-09-29
- Frames flagged by the frame safety check0 / 30decosa-api docs/evals/music-video-starring-you.md, 2026-09-29
- Visual qualitynot scored; drawn stills with camera moves
- Latency
- measured on our server: done in a minute or two
- Verification
- Proof: partial
Best
moving shots on MiniMax H3, one 96 GB card
The same shot list with each shot moving, from your consent-clip references. Not ready: blind testers rejected the moving shots, and the likeness gate keeps about half. Hosted once a render GPU is attached; self-host with your own H3 licence.
- Models
- Qwen3-ASR-1.7B (language pack speech service)
- Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)
- Qwen3.8-27B (NVFP4)
- decosa-mvideo-analyze (services/mvideo)
- MiniMax-H3 (reference mode, Turbo v4)
- decosa-api starring module + FFmpeg + c2pa-python
- Hardware
- 1x RTX PRO 6000 96 GB with about 60 GB free (fp8 H3, peak 49.6-52 GiB measured)
- Quality evidence
- Time per moving shot (sketch 864x480)57-65 s with one referencedecosa-api docs/evals/music-video-starring-you.md, 2026-09-29
- Face likeness to the consent clip (ArcFace, internal QC)0.36 (frontal re-rolls 0.51)decosa-api docs/evals/music-video-starring-you.md, 2026-09-29; faces often small or turned
- Blind testers who would pay for moving shots0 of 2 (both would for the storyboard)decosa-api docs/evals/music-video-starring-you.md, 2026-09-29
- Latency
- measured in a GPU window on 29 Sep 2026
- Verification
- Proof: partialSelf-host onlyRender receipts signed by the box; not a community-provider model.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Consent read-back: hears whether the clip says the sentence and its three fresh wordsQwen3-ASR-1.7B (language pack speech service)Qwen/Qwen3-ASR-1.7B on Hugging Face (opens in a new tab) 1.7BProof: partial | StandardBest | 1.7B | Proof: partial | |
| ||||
Face detection only (boxes): exactly one face, present and moving through the clip; its three sharpest frames become the only face referencesUltra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320) 0 GBNo proof yet | StandardBest | 0 GB | No proof yet | |
| ||||
Adult check on the consent frames (vision), the shot list per song section, and the frame safety and likeness checksQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | StandardBest | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Beat grid, bars and sections of the song (CPU; the music-video studio's analyzer): cuts land on bar downbeatsdecosa-mvideo-analyze (services/mvideo) 0 GBNo proof yet | StandardBest | 0 GB | No proof yet | |
| ||||
Storyboard: one still per shot drawn from the consent-clip references, then a camera move (push, pan, drift)FLUX.2 klein 4Bblack-forest-labs/FLUX.2-klein-4B on Hugging Face (opens in a new tab) 4B · 6 GBProof: partialIn the hosted demo | Standard | 4B · 6 GB | Proof: partialIn the hosted demo | |
| ||||
Moving shots: MiniMax H3 reference mode with the performer's consent-clip frames, Turbo v4 LoRA, sketch tier 864x480MiniMax-H3 (reference mode, Turbo v4)MiniMaxAI/MiniMax-H3 on Hugging Face (opens in a new tab) 33.1B (transformer) + 33.4B (text encoder) · 52 GBProof: partialSelf-host only | Best | 33.1B (transformer) + 33.4B (text encoder) · 52 GB | Proof: partialSelf-host only | |
| ||||
The edit (CPU): shots cut on the bar downbeats at 30 fps, the AI video label on every frame, an end card crediting the music, exports in 16:9, 9:16 and a Spotify Canvas loop; cut timing measured back from the pixels; C2PA per filedecosa-api starring module + FFmpeg + c2pa-python 0 GBProof: partial | StandardBest | 0 GB | Proof: partial | |
| ||||
Tools, services and hardware
Tools
- FFmpeg (opens in a new tab)LGPL-2.1+ / GPL for some builds
Cuts, camera moves, labels, the end card, and measuring the cuts back from the pixels.
- c2pa-python (opens in a new tab)MIT OR Apache-2.0
The C2PA content credential on every export.
- ComfyUI (opens in a new tab)GPL-3.0
Runs the FLUX.2 klein storyboard graph for the studio worker.
Primary source for the machine-readable marking duty the C2PA credential answers.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /studio/starring/info; POST /studio/projects/{pid}/performers/challenge, /performers/clip, /tracks, /videos; GET /studio/videos/{vid}; POST /studio/videos/{vid}/shots (edit or re-roll), /move (MiniMax H3 shots), /share, /unshare; POST /studio/starring/revoke. Needs FFmpeg and c2pa-python in the image.
- decosa-lang (speech):8492
Qwen3-ASR hears the consent sentence.
- comfyui:8188
FLUX.2 klein storyboard stills for the studio worker, one job at a time.
- H3 render worker
MiniMax H3 reference-mode shots on a 96 GB card; off until a render GPU is attached (the tool then makes a storyboard).
- vLLM:8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4, behind our gateway (hosted) or called directly (self-host).
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB, shared Fits
Measured on our server 2026-09-29: storyboard stills on GPU0 beside other services; MiniMax H3 shots in a GPU window with the speech and translation servers paused (peak 49.6-52 GiB); the text model on GPU1.
- 1x 24-32 GB card
Not tested. The storyboard path (FLUX.2 klein, about 6 GB) should fit with a remote text model; MiniMax H3 does not.
Latency per lane
- First picture drawn (shown once every picture is checked)12.1 s
Measuredmeasured on our server 2026-09-29: 12-20 s over 3 runs
- Storyboard video done (10 shots, 40 s)88.2 s
Measuredmeasured on our server 2026-09-29: 86-105 s (p50 88.2 s)
- One MiniMax H3 moving shot (sketch tier)61.0 s
Measuredmeasured on our server 2026-09-29 in a GPU window: 57-65 s with one reference
Notes
- Moving MiniMax H3 shots need a render GPU, which is switched on at launch; until then each shot is a drawn still with a camera move, cut on the beat.
- No lip-sync to the vocal yet.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa "music video starring you" on this machine
You are setting up a music-video maker that puts the artist in their own video: the artist records a 15-second consent
clip reading a sentence that ends in three fresh words; the tool checks the words, one live face and an adult, keeps the
consent in a ledger, finds the beats, bars and sections of their track, plans 8 to 12 shots per section, and cuts them on
the beat into 16:9, 9:16 and a Spotify Canvas loop, with an AI video label on every frame, an end card crediting the
music and a C2PA credential in every file. Shots are MiniMax H3 moving clips when an H3 worker is attached, otherwise
drawn stills (FLUX.2 klein) with a camera move. Work step by step, show me each command before you run anything with
`sudo`, and stop to ask if a check fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/music-video-starring-you.zip (1 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py music-video-starring-you` (the api image carries the same bundle under /app/rehearsal/music-video-starring-you/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py music-video-starring-you --bundle music-video-starring-you.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "moving H3 shots are refused while no render GPU is attached", "the video finishes", "the shot list has 10 shots"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- Apache-2.0: Qwen3.8-27B (adult check, shot list, frame checks), Qwen3-ASR-1.7B (consent read-back), FLUX.2 klein 4B
with its Qwen3-4B text encoder and VAE (storyboard stills); MIT: the Ultra-Light face detector (bundled in the api
image); librosa (ISC) for the beat grid.
- MiniMax H3 is NOT open for every use: the public community licence excludes the US, EU, UK and Korea. Attach an H3
worker only if you hold a licence that covers you; keep `LICENSE-MiniMax-H3.txt` next to the weights and "MiniMax H3"
in the credits (the api does this). Without it, the tool makes storyboards.
- Faces come only from a live consent clip: there is no photo upload anywhere, adults only, and a minor can't be enrolled
for the music-video purpose. Don't add a photo path.
- Only use tracks you have the rights to put in a video; the end card credits the music to the artist you name.
- Bind every port to 127.0.0.1. Logs carry ids, counts and hashes.
## 1. Check the machine
1. `nvidia-smi`: the storyboard needs about 6 GB above your baseline for FLUX.2 klein (fp8 transformer, NVFP4 text
encoder); the text model about 20 GB more (Qwen3.8-27B NVFP4 needs a Blackwell card; use a server you already run
otherwise). H3 moving shots need a 96 GB card with about 60 GB free: we measured a peak of 49.6-52 GiB with the fp8
transformer after the text encoder had run and been unloaded.
2. `docker --version`, `docker compose version`; if Docker or the NVIDIA container toolkit is missing, ask me, then
install them from the official repositories and check `docker run --rm --gpus all ubuntu nvidia-smi`.
3. Disk: Qwen3.8 NVFP4 is about 20 GB; FLUX.2 klein, its encoder and VAE and the ASR model add a few more; H3 adds
about 200 GB (our tree of transformers, text encoder, VAEs and the Turbo v4 adapter is 197 GB).
## 2. Images, weights and services
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
`git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out the newest release that contains
`decosa_api/studio_core/starring/` (`main` until one does), and `docker build -f docker/api/Dockerfile -t decosa-api:starring .`.
- ComfyUI for the storyboard: build `decosa-comfyui:local` as in the Decosa **characters** assemble prompt (ComfyUI at
commit `30bdda1ef13a3a34fce2cd2fec633f15d832122a`) and put `flux-2-klein-4b.safetensors`
(`black-forest-labs/FLUX.2-klein-4B`, revision `e7b7dc27f91deacad38e78976d1f2b499d76a294`),
`qwen_3_4b_fp4_klein.safetensors` and `flux2-klein-vae.safetensors` in its `diffusion_models`, `text_encoders` and
`vae` folders.
- The speech service: `services/lang/server.py` from the checkout in its own Python 3.12 venv (we run torch 2.8.0+cu128,
`qwen-asr` 0.0.6, fastapi, uvicorn, soundfile, librosa 1.0.0) with `Qwen/Qwen3-ASR-1.7B` (revision
`7278e1e70fe206f11671096ffdd38061171dd6e5`) in `LANG_MODELS/Qwen3-ASR-1.7B/`; run it with `LANG_PORT=8492`.
- The beat analyzer: `docker build -f services/mvideo/Dockerfile -t decosa-mvideo-analyze:local .` (CPU).
- Optional, H3: on the licensed box run `services/studio_core/h3/setup_box.sh`, then
`H3_FAMILY=turbo-v4 H3_MODE=ref H3_FP8=rowwise H3_TE=swap H3_WORKER_TOKEN=<a long random string> python
services/studio_core/h3/worker.py --port 8740` (127.0.0.1 only; reach another box over an SSH tunnel, never an open
port). We measured shots with the same fp8 transformer; a warm worker with `H3_TE=swap` on a 96 GB card is untested.
## 3. docker-compose.yml
Write this in `~/decosa/starring/`. `network_mode: host` lets the api reach ComfyUI, the speech service, the analyzer,
your model server and the H3 worker on 127.0.0.1.
```yaml
services:
analyze:
image: decosa-mvideo-analyze:local
network_mode: host
environment: { MVIDEO_ANALYZE_HOST: 127.0.0.1, MVIDEO_ANALYZE_PORT: "8488", MVIDEO_ANALYZE_THREADS: "8" }
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8488/health', timeout=4)"], interval: 30s, retries: 20, start_period: 600s }
api:
image: decosa-api:starring
network_mode: host
environment:
DECOSA_HOST: 127.0.0.1
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://127.0.0.1:8114/v1 # your Qwen3.8-27B server (vLLM, served name qwen3.8-27b)
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LANG_SPEECH_URL: http://127.0.0.1:8492
DECOSA_MVIDEO_ANALYZE_URL: http://127.0.0.1:8488
DECOSA_STUDIO_WORKER: command
DECOSA_STUDIO_WORKER_CMD: python /app/scripts/studio_worker.py
DECOSA_STUDIO_COMFY_URL: http://127.0.0.1:8188
DECOSA_STUDIO_WORK_DIR: /work
DECOSA_STUDIO_GPU_USD_PER_H: "1.32" # what your GPU hour costs you, for the estimates
DECOSA_PROVENANCE_DIR: /provenance
# H3 (only with a licence that covers you): uncomment these three
# DECOSA_STUDIO_H3_FAST: "1"
# DECOSA_STUDIO_H3_URL: http://127.0.0.1:8740
# DECOSA_STUDIO_H3_TOKEN: <the worker's token>
volumes: ["decosa-data:/data", "decosa-work:/work", "decosa-provenance:/provenance"]
depends_on: { analyze: { condition: service_healthy } }
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/studio/starring/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
decosa-data:
decosa-work:
decosa-provenance:
```
Use named volumes (the api runs as uid 10001). `docker compose up -d` and wait for both health checks. Create the C2PA
signing material once, then restart: `docker compose exec api python scripts/provenance_devcert.py && docker compose
restart api` (a development CA: valid signature, untrusted issuer in public validators). Back up the box's Ed25519 key:
`docker compose cp api:/data/attest ./attest-backup`. Keep the H3 token out of logs and never print it.
## 4. Smoke test (sample performers, sample song)
```bash
B=http://127.0.0.1:8445
T=$(curl -s -XPOST $B/demo/session -H 'content-type: application/json' -d '{"vertical":"studio"}' | jq -r .token); A="authorization: Bearer $T"
P=$(curl -s -XPOST $B/studio/projects -H "$A" -H 'content-type: application/json' -d '{"name":"[TEST] starring"}')
PID=$(echo "$P" | jq -r .project.id); K=$(echo "$P" | jq -r .project_key)
PERF=$(curl -s -XPOST "$B/studio/projects/$PID/performers/sample?key=$K" -H "$A" -H 'content-type: application/json' -d '{"which":"artist"}' | jq -r '.performers[0].id')
TR=$(curl -s -XPOST "$B/studio/projects/$PID/tracks?key=$K" -H "$A" -H 'content-type: application/json' -d '{"sample":"paper-lanterns"}' | jq -r .id)
V=$(curl -s -XPOST "$B/studio/projects/$PID/videos?key=$K" -H "$A" -H 'content-type: application/json' \
-d "{\"mode\":\"artist\",\"performers\":[\"$PERF\"],\"track_id\":\"$TR\",\"look\":\"cinematic\",\"aspect\":\"16:9\",\"direction\":\"warm dusk by a harbour\"}" | jq -r .id)
until curl -s "$B/studio/videos/$V?key=$K" -H "$A" | jq -e '.status=="done" or .status=="failed"' >/dev/null; do sleep 5; done
curl -s "$B/studio/videos/$V?key=$K" -H "$A" | jq '{status, shots: (.shots|length), sync: {found: .sync.found, planned: .sync.planned, ms: .sync.median_abs_ms}, safety: .safety[0], exports: (.exports|map_values(.credential))}'
```
Expect `done` with 10 shots, every planned cut found within about 17 ms of the beat, nothing flagged, and `c2pa` on the
16:9, 9:16 and canvas exports, in about 90 s on our shared card. With H3 attached, `POST /studio/videos/$V/move?key=$K`
with `{"video":"h3-sketch"}` makes the same shots move (about a minute a shot). Report the times you measure.
## 5. Your own clip and track
In the app: get a sentence (`/performers/challenge`), record the clip within its time limit (the three words are fresh
each time, so a stored clip fails), upload your track (up to 20 MB) with its title and artist, write a line of
direction, then edit or re-roll any shot (`POST /studio/videos/{vid}/shots`). The performer's revocation link deletes
their references and videos.
## 6. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`. Contract: `API_CONTRACT.md`, the Changes
entry "Studio: family interview film, music video starring you, our story film" (29 Sep 2026).What it does, in shortWho it's for, where it runs and the key results
Record a 15-second consent clip (a selfie reading a sentence that ends in three fresh words), add your track and a line of direction, and get 8 to 12 shots of you cut on the beat: the beat grid and sections of the song, a shot list per section you can edit, a re-roll on any shot, a 16:9 and a 9:16 export and a Spotify Canvas loop. Your face comes only from your consent clip: there is no photo upload, and adults only. Every frame says AI video and the end card credits the music to you. Moving shots run on MiniMax H3 when a render GPU is attached; otherwise the shots are a storyboard of drawn stills with camera moves.
In short
Last reviewed
- What it is
- Record a 15-second consent clip, add your track, and get a music video of you cut on the beat, in 16:9, 9:16 and a Canvas loop.
- Who it's for
- Independent artists releasing a single who want a first video with themselves in it.
- Where it runs
- Hosted (moving shots when a render GPU is attached); self-host on a 96 GB card with your own H3 licence
- Key numbers
- 27 / 27 Planned cuts found on the beat in the exports (synthetic, n = 27)
- 0 / 30 Storyboard frames flagged by the frame safety check (synthetic, n = 30)
- 88.2 s Storyboard video done (p50) (synthetic, n = 3)
- 88.2 s Median end-to-end run, hosted (QA sweep 2026-09-29)
How we tested itEnd-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 29 Sep 2026 · measured 29 Sep 2026: · p50 88 s · p95 105 s (3 runs) · ~$0.029 per run · 5 receipts
Loading the nightly status…
Self-host: not yet verified
Measured cost to run: about $0.024 per video (hosted, 29 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
Known limits (7)
- Moving shots need a render GPU, which is off until launch: the hosted demo makes a storyboard (drawn stills with camera moves).
- No lip-sync to the vocal.
- Likeness is checked by the text model on each frame; ArcFace numbers are internal QC only.
- The adult check is a vision estimate from the consent frames, not an ID check.
- Moving H3 shots aren't ready: blind testers rejected them (faces lost or someone else's); the likeness gate keeps about half and still lets some wrong faces through.
- The likeness check is a yes/no from the text model on each picture; it caught 1 unlike picture in the couple test after the fix, and missed one before it (only the first partner was checked).
- Measured with synthetic sample performers and one sample song.
Eval results, nightly checks and cost per run · eval not held out
Technical detailsModels, where it runs, labels, what it is built from
- Models
- MiniMax H3 (reference mode, Turbo v4) · FLUX.2 klein 4B (storyboard) · Qwen3.8-27B (shot list, checks) · Qwen3-ASR-1.7B (consent read-back)
- Where
- Hosted (moving shots when a render GPU is attached); self-host on a 96 GB card with your own H3 licence
- Checks
- Consent clip checked (words heard, one live face, an adult); frame safety and likeness checks; cuts measured on the beat; C2PA in every export
- Industry
- Music · Creative and media
- Input
- Files and media · Voice, live audio
- Output
- Media
- Data
- Personal data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Part of
- Decosa Studio: Music
- Runs in
- Decosa hosted · Self-host
- Built from
- Live speech to text · Studio render · Content credentials
Every result carries a signed record of which model produced it, so you can check it later. How that works
Questions people ask
Can I upload a photo instead of recording a clip?
No. Your face comes only from a live 15-second consent clip that reads three fresh words, so nobody can make a video of someone else. Adults only.
Are the shots moving video?
Not yet. The hosted tool makes a storyboard: a drawn still of you per shot with a camera move, cut on the beat. MiniMax H3 moving shots (57-65 s a shot, measured) lost the face too often in a blind test, so they stay off until they hold it.
How well does it cut on the beat?
In 3 renders every planned cut was found in the file (27 / 27), median 5.7 ms and at most 16.3 ms from the beat, under half a frame at 30 fps.
Is it labelled as AI?
Yes: an AI video label on every frame, an end card crediting the music to you, and a C2PA credential in every file. Use each platform's own AI toggle as well.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Music video starring you
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…