Make a music video for your track
A first music video for an independent artist's own track. It finds the beats, bars and sections, times the artist's lyrics to the audio, runs the sample-clearance pre-check, and has an open model write scenes that each cite the lines they illustrate; code and a receipted yes/no check test every citation. Clips render with an open video model, cut on the bar lines, with the lyrics as karaoke captions, in 16:9 and 9:16. Each file carries a C2PA credential and a burned-in AI label, and a signed record holds the rights statement, the checks and the measured cut timing. The visuals are 480p upscaled: a stylised first video, not a studio shoot.
- For
- Teams in music and film, tv and games.
- Time per task15 stypical (median) on the sample
- Cost per task~$0.19 per minute of videomeasured, at list price
- Accuracy81.8% / 87.7%Lyric lines within 0.3 s / 1 s of human timing (held-out test)All results and caveats
Result
LivePick a sample or upload your own track, confirm the rights, and analyse. Nothing renders until you ask.
Watch: a track analysed, written, rendered and labelled
Replay · not liveRecorded from real runs on the pre-release server on 26 Sep 2026: Qwen3.8-27B through our gateway (receipts are real), the CPU analyzer, and renders on the shared studio GPU. Replay plays the events faster than they happened and shows re-encoded, smaller copies of the videos (their C2PA credential is not in the copy; the originals carry it). The CC BY songs are credited on screen.
Look for: the beat grid and 18 lines landing on the timeline, each scene quoting the lines it shows, cuts on the bar lines, captions filling word by word, and the disclosure rules met on the file.
Pick a sample or upload your own track, confirm the rights, and analyse. Nothing renders until you ask.
Get an API key
- Call the music video from your track API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) shared with the text model, or a 48 GB card for the video model plus a remote text model; the analysis runs on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- music-video-studio
Use the hosted API
# Decosa music video from your track: use the hosted API
You are wiring Decosa's music-video studio into this project. It takes an artist's own track and lyrics, finds the beat,
bars and sections, times the lyrics to the audio, runs a sample-clearance pre-check, writes a treatment whose scenes cite
the lyric lines they show (each scene checked by code and by a receipted yes/no call), and on request renders a video in
16:9 and 9:16: clips from an open video model, cut on the bar lines, with karaoke captions, an AI label on every frame and
a C2PA credential. Every text-model call has its own signed receipt. Use only what is listed below. If you need something
else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- Only send tracks the user owns or has the rights to use; the API refuses a request without the rights statement (403)
and signs the statement into the record. It does not verify it.
- The visuals are 480p clips upscaled to 720p: a stylised first video. Keep the burned-in label; also use each platform's
AI-content toggle when publishing.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in
code. Send `Authorization: Bearer $DECOSA_API_KEY`. A key renders 3 videos per UTC day on the shared GPU.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "music-video-studio"}` returns `{"token", ...}`.
Sessions last 30 minutes; a render takes longer, so poll it with the run's `watch` key (below).
## Endpoints
- `POST /mvideo/analyze` (token). Body: `{"audio_b64": "<file as base64>" | "sample": "<id>", "lyrics"?: "one sung line per line",
"bpm"?: 40-240, "title"?, "artist"?, "look"?, "start_s"?, "end_s"?, "aspects"?: ["16:9","9:16"],
"rights": {"confirmed": true, "role": "artist"|"songwriter"|"label"|"publisher"|"licensee"|"open-licence"}, "stream"?: false}`.
- SSE by default (`ready`, `stage`, `analysis`, `lyrics`, `clearance`, `treatment`, `scene`…, `plan`, `receipt`…, `done`);
`"stream": false` returns one JSON object with `run_id`, `watch`, `analysis`, `lyrics`, `clearance`, `treatment`, `scenes`,
`plan`, `receipts`.
- Send `bpm` when you know it: without it the beat tracker can land on half or double time.
- Limits: 20 MB of audio, 8 minutes, 12,000 characters and 250 lines of lyrics (English), 60 s of video per run.
- Errors: 403 no rights statement, 400 bad input, 422 no steady beat or no alignable words, 429 busy (`Retry-After`).
- `POST /mvideo/runs/{run_id}/render` (token) `{"aspects"?: [...]}` → 202; about 20-25 minutes for 50 s in both formats.
- `GET /mvideo/runs/{run_id}?watch=<watch>` (no token): poll every 20-30 s until `render.status` is `done` or `failed`.
`done` has `render.outputs["16:9"|"9:16"]` with `url` (relative: prefix `https://api.decosa.ai`), `file_sha256`, `credential`,
and `cut_accuracy` (measured from the pixels), plus `disclosure` (the rules run on the file) and `certificate`.
- `GET /mvideo/runs/{run_id}/captions?format=srt|lrc|ass` (token): the timed lyrics on their own.
- `GET /mvideo/certificates/{run_id}` (no token): the signed record; verify with `POST /record/verify {"record": ...}`.
- `GET /mvideo/info`, `GET /mvideo/samples` (no token).
## Example: analyse, render, save both videos (Python, `pip install httpx`)
```python
import base64, httpx, os, pathlib, time
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {"audio_b64": base64.b64encode(pathlib.Path("track.mp3").read_bytes()).decode(),
"lyrics": pathlib.Path("lyrics.txt").read_text(), "bpm": 96, "title": "My Song", "artist": "Me",
"rights": {"confirmed": True, "role": "artist"}, "stream": False}
a = httpx.post(f"{API}/mvideo/analyze", headers=H, json=body, timeout=300).raise_for_status().json()
print([(s["id"], s["verdict"], s["lines"]) for s in a["scenes"]], a["plan"]["estimate"]["render_minutes"], "min")
httpx.post(f"{API}/mvideo/runs/{a['run_id']}/render", headers=H, json={}, timeout=60).raise_for_status()
while True:
time.sleep(30)
run = httpx.get(f"{API}/mvideo/runs/{a['run_id']}", params={"watch": a["watch"]}, timeout=600).json()
if run["render"]["status"] in ("done", "failed"):
break
for aspect, o in run["render"]["outputs"].items():
pathlib.Path(f"video-{aspect.replace(':', 'x')}.mp4").write_bytes(httpx.get(API + o["url"], timeout=120).content)
print(aspect, o["credential"], o["cut_accuracy"].get("median_abs_ms"), "ms from the beat")
```
## Honest limits
- Lyric timing: 82% of held-out lines start within 0.3 s of human timing; when it slips, whole passages slip. Check the
low-confidence lines, or send LRC timestamps. English only.
- Beats: without the BPM, the tracker got the tempo right on 43% of held-out grooves; with it, 100%. Which beat is "one"
is a heuristic and often wrong.
- The grounding check alone catches about two thirds of mismatched scenes; the code checks on citations do the rest.
- Visual quality is well below closed video models. The clearance pre-check knows only a small open catalogue.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa music video from your track: run it yourself (containers)
You are setting up Decosa's music-video studio on this machine: a CPU analyzer (beats, bars, sections, lyric timing),
a treatment from Qwen3.8-27B with checked citations, clips from Wan2.2-VACE-Fun-A14B through ComfyUI, and an ffmpeg edit
with captions, an AI label and a C2PA credential. Tracks, lyrics and renders stay here; nothing is sent to Decosa's
hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images or models.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/music-video-studio.zip (825 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py music-video-studio` (the api image carries the same bundle under /app/rehearsal/music-video-studio/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py music-video-studio --bundle music-video-studio.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "a request without the rights statement is refused", "a text file sent as audio is refused", "a steady beat near 112 BPM"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Licences first
All models are Apache-2.0: Qwen3.8-27B, Wan2.2-VACE-Fun-A14B with the Wan2.2-Lightning LoRAs, the Wan2.1 VAE and umt5
encoder, and wav2vec2-large-960h-lv60-self. librosa is ISC, FFmpeg LGPL/GPL, ComfyUI GPL-3.0. Only process tracks you
own or have the rights to use.
## Steps
1. Docker and the NVIDIA container toolkit: if `docker compose version` or
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi` fails, install them from the official
instructions. Plan on a 96 GB card (a 48 GB card is untested) plus about 20 GB for the text model or an existing
Qwen3.8-27B server, 8+ CPU cores and about 110 GB of disk.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the
`llm`, `comfyui`, `mvideo-analyze` and `api` services. For `api` set `DECOSA_LLM_ROUTE=direct`,
`DECOSA_STUDIO_WORKER=command`, `DECOSA_MVIDEO_ANALYZE_URL=http://mvideo-analyze:8488`; use named volumes; bind every
port to 127.0.0.1. The api image needs FFmpeg, fonts-dejavu-core and c2pa-python.
3. `docker compose pull && docker compose up -d`; wait for the health checks (the first start downloads the weights).
4. Smoke test: a token from `POST /demo/session {"vertical":"music-video-studio"}`; `POST /mvideo/analyze
{"sample":"rxbyn-bad-side","stream":false}` must be refused (403, no rights statement); with
`"rights":{"confirmed":true,"role":"open-licence"}` it must return about 112 BPM, 18 timed lines and scenes that cite
them; then `POST /mvideo/runs/{id}/render {"aspects":["9:16"]}`, poll `GET /mvideo/runs/{id}?watch=<watch>` to
`done`, and check `credential: "c2pa"`, the cut timing, and that `POST /record/verify` accepts `GET /mvideo/certificates/{id}`.
5. Report back: the signing key id (`GET /attest/signing-key`), the time per clip and the cut timing.
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVFP4) needs a GPU.
- GeForce RTX 4090lite tierRuns with a smaller tier
The standard tier does not fit: Needs about 54 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes.
- GeForce RTX 5090lite tierRuns with a smaller tier
The standard tier does not fit: Needs about 62 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.
- 2x GeForce RTX 5090lite tierRuns with a smaller tier
The standard tier does not fit: Wan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAs needs about 34 GB on one GPU; each GPU here has 32 GB. The lite tier fits with changes.
- L40Slite tierRuns with a smaller tier
The standard tier does not fit: Needs about 67.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (91.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns
The standard tier fits (91.6 of 192 GB). The best tier fits too.
- Apple M3 Ultra (Mac Studio), 96 GBlite tierRuns with a smaller tier
The standard tier can't be checked: Wan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAs has no mapped Apple Silicon build The lite tier fits with changes.
- Apple M5 Max, 64 GBlite tierRuns with a smaller tier
The standard tier can't be checked: Wan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAs has no mapped Apple Silicon build The lite tier fits with changes.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"music-video-studio"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py music-video-studio
Download the mock-data bundle (825 KB, 10 checks)expected.json
A 50 s excerpt of a CC BY song and its lyrics. A request without the rights statement must be refused; with it, the analysis must find a steady beat near 112 BPM, place all 18 sung lines with the first and ninth line within half a second of the dataset's human timing, write a treatment whose scenes each cite lines and pass the checks, and plan a renderable edit, with every model call receipted. A text file sent as audio must be refused. No render: a render takes about 20 minutes of GPU; run it from the console or with POST /mvideo/runs/{id}/render when you are ready.
What the rehearsal checks
- a request without the rights statement is refused
- a text file sent as audio is refused
- a steady beat near 112 BPM
- all 18 sung lines are placed on the audio
- the first line starts within 0.5 s of the human timing (0.78 s)
- the ninth line starts within 0.5 s of the human timing (18.62 s)
- every scene cites lines and passes the code checks
- the edit plan cuts on the bar lines at least 10 times
- the plan can be rendered
- every model call is receipted and signed
Licence: inputs/bad-side-excerpt.ogg: "Bad Side" by Rxbyn (Jamendo), CC BY 4.0, excerpt 0:40-1:30, faded. inputs/bad-side-lyrics.txt: the song's lyrics as normalised in JamendoLyrics (annotations MIT). Your own track and lyrics replace these files.
Prompt for your coding agent
# Decosa music video from your track: run it yourself (containers)
You are setting up Decosa's music-video studio on this machine: a CPU analyzer (beats, bars, sections, lyric timing),
a treatment from Qwen3.8-27B with checked citations, clips from Wan2.2-VACE-Fun-A14B through ComfyUI, and an ffmpeg edit
with captions, an AI label and a C2PA credential. Tracks, lyrics and renders stay here; nothing is sent to Decosa's
hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images or models.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/music-video-studio.zip (825 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py music-video-studio` (the api image carries the same bundle under /app/rehearsal/music-video-studio/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py music-video-studio --bundle music-video-studio.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "a request without the rights statement is refused", "a text file sent as audio is refused", "a steady beat near 112 BPM"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Licences first
All models are Apache-2.0: Qwen3.8-27B, Wan2.2-VACE-Fun-A14B with the Wan2.2-Lightning LoRAs, the Wan2.1 VAE and umt5
encoder, and wav2vec2-large-960h-lv60-self. librosa is ISC, FFmpeg LGPL/GPL, ComfyUI GPL-3.0. Only process tracks you
own or have the rights to use.
## Steps
1. Docker and the NVIDIA container toolkit: if `docker compose version` or
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi` fails, install them from the official
instructions. Plan on a 96 GB card (a 48 GB card is untested) plus about 20 GB for the text model or an existing
Qwen3.8-27B server, 8+ CPU cores and about 110 GB of disk.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the
`llm`, `comfyui`, `mvideo-analyze` and `api` services. For `api` set `DECOSA_LLM_ROUTE=direct`,
`DECOSA_STUDIO_WORKER=command`, `DECOSA_MVIDEO_ANALYZE_URL=http://mvideo-analyze:8488`; use named volumes; bind every
port to 127.0.0.1. The api image needs FFmpeg, fonts-dejavu-core and c2pa-python.
3. `docker compose pull && docker compose up -d`; wait for the health checks (the first start downloads the weights).
4. Smoke test: a token from `POST /demo/session {"vertical":"music-video-studio"}`; `POST /mvideo/analyze
{"sample":"rxbyn-bad-side","stream":false}` must be refused (403, no rights statement); with
`"rights":{"confirmed":true,"role":"open-licence"}` it must return about 112 BPM, 18 timed lines and scenes that cite
them; then `POST /mvideo/runs/{id}/render {"aspects":["9:16"]}`, poll `GET /mvideo/runs/{id}?watch=<watch>` to
`done`, and check `credential: "c2pa"`, the cut timing, and that `POST /record/verify` accepts `GET /mvideo/certificates/{id}`.
5. Report back: the signing key id (`GET /attest/signing-key`), the time per clip and the cut timing.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Runs with a smaller tierMusic video from your track on GeForce RTX 5090: use the Lite · timed captions and a cited treatment, no GPU for video tier
The standard tier does not fit: Needs about 62 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.
Lite · timed captions and a cited treatment, no GPU for video: what changesuses estimates
- Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Beat, bar and section detection: decosa-mvideo-analyze (services/mvideo). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Lyric timing: wav2vec2-large-960h-lv60-self. CPU. Runs on CPU (vram_gb 0 in stack.json).
- Treatment writer: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
- Sample and lyric clearance pre-check on the u...: decosa-api clearance module (decosa_api/verticals/clearance). CPU. Runs on CPU (vram_gb 0 in stack.json).
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Music video from your track, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Music video from your track on my hardware Fetch https://decosa.ai/prompts/music-video-studio-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=music-video-studio) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Lite · timed captions and a cited treatment, no GPU for video (lite). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Beat, bar and section detection: decosa-mvideo-analyze (services/mvideo), CPU - Lyric timing: wav2vec2-large-960h-lv60-self (facebook/wav2vec2-large-960h-lv60-self), CPU - Treatment writer: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. - Sample and lyric clearance pre-check on the u...: decosa-api clearance module (decosa_api/verticals/clearance), CPU GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/music-video-studio-assemble.md
Get an API key
- Call the music video from your track API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) shared with the text model, or a 48 GB card for the video model plus a remote text model; the analysis runs on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
A first music video from an artist's own track: cuts on the beat, the lyrics as timed captions, scenes written from the lyrics, labelled and credentialed.
An independent artist uploads a track they own, with its lyrics. A CPU analyzer finds the beat, bars and sections and times each lyric line to the audio; the sample-clearance pre-check looks for samples and lifted lyrics. An open text model writes a treatment whose scenes each cite the lines they show, code checks every citation and a receipted yes/no call checks each scene. An open video model renders a clip per scene, and the edit cuts on the bar lines with karaoke captions, in 16:9 and 9:16. Every file has a burned-in AI label and a C2PA credential, and a signed record holds the rights statement, the checks and the measured cut timing. It is a stylised first video at 480p upscaled, not a studio shoot.
- Deployment
- Hosted or self-host
- Regulatory
- Not legal advice; checked on 26 Sep 2026 unless stated. Rights: the artist confirms they own the track or have the rights to make and publish a video with it; the statement goes into the signed record and is not verified. The sample-clearance pre-check compares the upload with a small open catalogue only, and a clean result says nothing about commercial music (US courts disagree on short samples: Bridgeport v. Dimension Films, 6th Cir. 2005, against VMG Salsoul v. Ciccone, 9th Cir. 2016). Demo tracks are CC BY excerpts from JamendoLyrics, credited on screen, and one song made with MiniMax-Music3 through the music-gen-cleared tool. Disclosure: every file has a visible "AI-generated visuals / Made with AI" label on every frame and a C2PA credential that declares composite AI media. EU AI Act (Regulation (EU) 2024/1689) Art. 50(2) asks providers to mark generated video in a machine-readable way (from 2 Aug 2026; checked by the disclosure pre-flight on 25 Sep 2026; EUR-Lex link under Tools). New York General Business Law § 396-b requires a conspicuous disclosure of synthetic performers in advertisements (in force 9 Jun 2026; bill text linked under Tools): it applies if the video is used as an ad, and the console runs that rule on the file. YouTube asks creators to disclose realistic altered or synthetic content when they upload and labels it (help page read 26 Sep 2026, linked under Tools); TikTok and Meta have their own AI labels (not checked here): use each platform's toggle as well. Copyright: the US Copyright Office's Part 2 report (29 Jan 2025) says AI output is protected only where a human determined sufficient expressive elements; the record claims no copyright in the generated visuals. Scenes never depict real people, brands or readable text by prompt rule and screen.
Text description
An artist's track, lyrics, optional BPM and look, and a required rights statement go to a CPU analyzer: librosa finds beats, bars and sections, and wav2vec2 (Apache-2.0) aligns the lyrics to the audio. The sample-clearance pre-check compares the upload with a small open catalogue. Qwen3.8-27B (Apache-2.0) writes a treatment whose scenes cite their lyric lines; code checks the citations and a receipted yes/no call checks each scene. Wan2.2-VACE-Fun-A14B (Apache-2.0) renders a 5 s clip per scene per format on the studio GPU, one job at a time. FFmpeg cuts the clips on the bar lines at 30 fps, burns in karaoke captions, the AI label and the credit, and lays the artist's audio under it. The finished files are checked: cut timing measured from the pixels, the disclosure rules run on the file, a C2PA credential on each MP4, and a signed record. Text-model calls carry gateway receipts; renders carry receipts signed by the box. Self-hosted, everything stays on your machine.
At a glance
- Data retention
- The track is kept 24 hours after upload so it can be rendered (longer only while a render is queued or running), then deleted. The run (hashes, analysis, lyric timings, captions, plan, record) is kept 7 days; the rendered videos stay in the studio's media folder and open only through signed links that expire within hours, given to the run's owner or its watch link. Logs hold ids, counts and hashes, never lyrics.
- What leaves the box
- On the hosted route, the lyrics and plan go to the text model through our gateway and the render runs on Decosa's hosted service. Self-hosted, nothing leaves your machine.
- Rights
- You confirm you own the track or hold the rights; the statement is signed into the record, not verified. The clearance pre-check covers only a small open catalogue.
- Cost per video
- A fraction of a cent of model calls per analysis (measured). Rendering a video in one format takes several GPU-minutes, tens of cents at a GPU rental list price; both formats about twice that. Hosted demo renders are not billed.
- Typical time
- Seconds for the analysis and treatment of an excerpt; many minutes to render it in one format and about twice that in both, on a shared GPU (measured).
- Outputs
- MP4s in 16:9 (1280x720) and 9:16 (720x1280) at 30 fps with captions, an AI label on every frame and a C2PA credential; SRT, LRC and ASS captions; a signed record with the rights statement, checks, receipts and measured cut timing.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
- In the hosted demo
Lite
timed captions and a cited treatment, no GPU for video
Beat grid, lyric timing (SRT, LRC, ASS), the clearance pre-check and a checked treatment with an edit plan. Bring your own footage.
- Models
- decosa-mvideo-analyze (services/mvideo)
- wav2vec2-large-960h-lv60-self
- Qwen3.8-27B (NVFP4)
- decosa-api clearance module (decosa_api/verticals/clearance)
- Hardware
- CPU (8+ cores) plus the text model (about 20 GB of GPU, or a remote server)
- Quality evidence
- Lyric line starts within 0.3 s / 1 s of human timing (6 held-out CC BY-ND songs)81.8% / 87.7% (baseline 16.8% within 1 s)decosa-api docs/evals/music-video-studio.md, 2026-09-26
- Beat F-measure (±70 ms), 40 human-played grooves0.87 with the artist's BPM, 0.59 withoutdecosa-api docs/evals/music-video-studio.md, 2026-09-26; drums only
- Treatment citations valid / anchor words found (52 held-out scenes)100% / 96%decosa-api docs/evals/music-video-studio.md, 2026-09-26
- Latency
- measured: seconds per excerpt on the shared gateway
- Verification
- Proof: partial
- In the hosted demo
Standard
the hosted demo, clips on one 96 GB card
Everything in Lite plus a clip per scene on Wan2.2-VACE (Apache-2.0), cut on the bar lines in 16:9 and 9:16, labelled and credentialed.
- Models
- decosa-mvideo-analyze (services/mvideo)
- wav2vec2-large-960h-lv60-self
- Qwen3.8-27B (NVFP4)
- decosa-api clearance module (decosa_api/verticals/clearance)
- Wan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAs
- decosa-api mvideo module + FFmpeg + c2pa-python
- Hardware
- 1x RTX PRO 6000 Blackwell 96 GB (shared), CPU for the analyzer and the edit
- Quality evidence
- Cuts found in the rendered files / offset from the beat grid133 of 145 planned cuts found, 0 false; median 7-10 ms, max 16.3 ms (half a frame at 30 fps)decosa-api docs/evals/music-video-studio.md, 2026-09-26; 7 files from 5 renders; offsets against the detected beat grid
- GPU time per minute of videoabout 680 GPU-seconds for one format, 1,360-1,740 for bothdecosa-api docs/evals/music-video-studio.md, 2026-09-26; shared card
- Disclosure rules on the delivered files (tool 49)label read on 6 of 6 sampled frames and the C2PA marking valid in all 4 files checkeddecosa-api docs/evals/music-video-studio.md, 2026-09-26
- Typed grounding check: swapped (mismatched) scenes caught67% (the code checks on citations and anchor words do the rest)decosa-api docs/evals/music-video-studio.md, 2026-09-26
- Visual qualitynot measured yet
- Latency
- measured on our server: many minutes for a video in one format, about twice that for both, on a shared GPU
- Verification
- Proof: partial
Best
MiniMax H3 clips (licence pending), self-host
Everything in Standard with clips from MiniMax-H3 instead of Wan2.2. Self-host available where H3's current licence covers you; the pending licence would bring it to the hosted demo. Not wired in yet.
- Models
- decosa-mvideo-analyze (services/mvideo)
- wav2vec2-large-960h-lv60-self
- Qwen3.8-27B (NVFP4)
- decosa-api clearance module (decosa_api/verticals/clearance)
- decosa-api mvideo module + FFmpeg + c2pa-python
- MiniMax-H3 (licence pending)
- Hardware
- 1x RTX PRO 6000 96 GB plus about 115 GB of RAM for offload
- Quality evidence
- clip qualitynot measured yet
- Latency
- not measured yet
- Verification
- No proof yetSelf-host onlySelf-host: your box signs the render record. Not a community-provider model.
- Needs more compute
Wanted: the best setup
MiniMax H3 on two cards, no offload
A clip per scene adds up over a song. Both halves of MiniMax-H3 in BF16 on two 96 GB cards, so nothing is offloaded to RAM: faster renders (estimate, not measured). Self-host now where H3's current licence covers you; once the pending licence lands we add this compute for the hosted tier ourselves. Not a community-provider model.
- Models
- decosa-mvideo-analyze (services/mvideo)
- wav2vec2-large-960h-lv60-self
- Qwen3.8-27B (NVFP4)
- decosa-api clearance module (decosa_api/verticals/clearance)
- decosa-api mvideo module + FFmpeg + c2pa-python
- MiniMax-H3 (licence pending)
- Hardware
- 2x RTX PRO 6000 96 GB, no CPU offload (about 133 GB of BF16 weights; estimate)
- Quality evidence
- render time per clip against one card with offloadnot measured yet
- Latency
- not measured yet
- Verification
- No proof yetSelf-host onlySelf-host: your box signs the render record. Not a community-provider model.
Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Also runs on
- Sharper clips with LTX-2.3LTX-2.3 22B (distilled)self-host onlyLTX-2.3 is on the box but not wired in, and self-host only: its community licence has a competing-service clause, so we can't serve it hosted. Hardware: 1x 96 GB card.
We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Beat, bar and section detection (CPU): librosa beat tracker on a full-band plus low-band onset envelope, bar phase by a kick-and-snare heuristic (4/4), sections by checkerboard novelty on chroma and MFCC self-similaritydecosa-mvideo-analyze (services/mvideo) 0 GBNo proof yet | LiteStandardBestWanted | 0 GB | No proof yet | |
| ||||
Lyric timing (CPU): CTC forced alignment of the artist's own lyrics on the mix, with a repair pass for lines squeezed into too little time; English letterswav2vec2-large-960h-lv60-selffacebook/wav2vec2-large-960h-lv60-self on Hugging Face (opens in a new tab) 317M · 0 GBNo proof yet | LiteStandardBestWanted | 317M · 0 GB | No proof yet | |
| ||||
Treatment writer (scenes that cite the lyric lines they show) and the typed yes/no grounding check per scene; also labels near-duplicate lyric lines in the clearance pre-checkQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | LiteStandardBestWanted | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Sample and lyric clearance pre-check on the upload (tool 38, run in-process): audio landmarks and melody against a small open catalogue, lyric lines against a lyric setdecosa-api clearance module (decosa_api/verticals/clearance) 0 GBProof: partial | LiteStandardBestWanted | 0 GB | Proof: partial | |
| ||||
Clips: text-to-video, one 5 s clip per scene per formatWan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAsalibaba-pai/Wan2.2-VACE-Fun-A14B on Hugging Face (opens in a new tab) A14B (two 14B experts, high and low noise) (14B per step active)No proof yetIn the hosted demo | Standard | A14B (two 14B experts, high and low noise) (14B per step active) | No proof yetIn the hosted demo | |
| ||||
The edit and the marks (CPU): cuts on bar lines on a 30 fps grid, karaoke captions (ASS, libass), the AI label on a top bar, the credit, the artist's audio; cut timing measured back from the pixels; the disclosure pre-flight's rules (tool 49) on the file; C2PA credential per file; the signed recorddecosa-api mvideo module + FFmpeg + c2pa-python 0 GBProof: partial | StandardBestWanted | 0 GB | Proof: partial | |
| ||||
Alternate, self-host only: sharper clipsLTX-2.3 22B (distilled)Lightricks/LTX-2.3 on Hugging Face (opens in a new tab) 22BNo proof yetSelf-host only | Alternate | 22B | No proof yetSelf-host only | |
| ||||
A clip per scene with native audio, in place of Wan2.2MiniMax-H3 (licence pending)MiniMaxAI/MiniMax-H3 on Hugging Face (opens in a new tab) 33.1B (transformer) + 33.4B (text encoder) · about 133 GB (estimate)No proof yetSelf-host only | BestWanted | 33.1B (transformer) + 33.4B (text encoder) · about 133 GB (estimate) | No proof yetSelf-host only | |
| ||||
Tools, services and hardware
Tools
- JamendoLyrics MultiLang (opens in a new tab)Annotations MIT; audio CC per song (only CC BY excerpts are shown; CC BY-ND songs measured only)
Demo excerpts and the lyric-alignment eval (9 English songs without an NC clause, human word and line timings).
The beat-tracking eval: human drummers played to a click, synthesised with a small drum kit.
- FFmpeg with libass (opens in a new tab)LGPL-2.1+ / GPL for some builds
Cuts, captions, labels, audio, and measuring the cuts back from the pixels.
- c2pa-python (opens in a new tab)MIT OR Apache-2.0
The C2PA content credential on every MP4 (via the provenance kit).
- ComfyUI (opens in a new tab)GPL-3.0
Runs the VACE graph for the studio worker.
Primary source for the machine-readable marking duty (Art. 50(2)) the C2PA credential answers.
Primary source for the synthetic-performer disclosure in advertisements, run on each file by the disclosure pre-flight.
What YouTube asks creators to disclose at upload (read 26 Sep 2026).
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /mvideo/info, /mvideo/samples; POST /mvideo/analyze (SSE), /mvideo/runs/{id}/render; GET /mvideo/runs/{id} (token or watch key), /mvideo/runs/{id}/captions, /mvideo/certificates/{id}. Needs FFmpeg, fonts, c2pa-python (and Tesseract for the label check) in the image.
- decosa-mvideo-analyze:8488
CPU analyzer (services/mvideo/Dockerfile, built locally): POST /v1/analyze, GET /health. Keeps nothing.
- comfyui:8188
Wan2.2-VACE-Fun-A14B renders for the studio worker, one job at a time.
- vLLM:8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4, behind our gateway (hosted) or called directly (self-host).
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB, shared Fits
Measured on our server 2026-09-26: clips on GPU0 beside other services and a sibling's renders, 84-100 s per 5 s clip; the text model on GPU1; the analyzer on CPU.
- 1x 48 GB card
Not tested. ComfyUI offloads the idle expert, so the fp8 VACE graph should fit; the text model would need to be remote.
- CPU only Fits
The lite tier: analysis, lyric timing, treatment (with a remote text model), captions and the edit plan, no clips.
Latency per lane
- Analysis, clearance, treatment and checks for a 50-60 s excerpt15.0 s
Measuredmeasured on our server 2026-09-26: 13-32 s over the recorded, e2e and eval runs on the shared gateway
- Lyric alignment, 3-5 minute song24.0 s
Measuredmeasured on our server 2026-09-26: 18-33 s on 16 CPU threads
- One 5 s clip (Wan2.2-VACE, 4 steps)86.0 s
Measuredmeasured on our server 2026-09-26: medians 84-96 s in four renders on shared GPU0 (141 s when interleaved with another job)
- 55-60 s video in 9:16 (7 clips, edit, checks)645.0 s
Measuredmeasured on our server 2026-09-26: 628 s and 663 s
- 50 s video in 16:9 and 9:16 (12-14 clips)1400.0 s
Measuredmeasured on our server 2026-09-26: 1,359-1,449 s
Notes
- The visuals are 480p clips upscaled to 720p: a stylised visualiser, clearly below closed video models and the MiniMax H3 samples elsewhere on this site.
- The hosted demo caps the video at 60 s (a chorus or a teaser); self-hosted, raise DECOSA_MVIDEO_MAX_VIDEO_S.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa music-video studio on this machine
You are setting up a music-video maker for an artist's own tracks: a CPU analyzer finds the beat, bars and sections and
times the lyrics to the audio; an open text model writes a treatment whose scenes cite the lyric lines they show, and
checks each one; an open video model renders a clip per scene; ffmpeg cuts them on the bar lines with karaoke captions,
an AI label and the artist's audio; every file gets a C2PA credential and the run a signed record. Work step by step,
show me each command before you run anything with `sudo`, and stop to ask if a check fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/music-video-studio.zip (825 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py music-video-studio` (the api image carries the same bundle under /app/rehearsal/music-video-studio/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py music-video-studio --bundle music-video-studio.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "a request without the rights statement is refused", "a text file sent as audio is refused", "a steady beat near 112 BPM"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- All models here are Apache-2.0: Qwen3.8-27B (treatment and checks), Wan2.2-VACE-Fun-A14B with the Wan2.2-Lightning
4-step LoRAs (clips), the Wan2.1 VAE and umt5-xxl encoder, and `facebook/wav2vec2-large-960h-lv60-self` (lyric
alignment, CPU). Beat tracking is librosa (ISC). LTX-2.3 would be faster and sharper but is under the LTX-2 Community
License (free below USD 10M yearly revenue, conditions on distribution); it is not wired into this pipeline.
- Only upload tracks you own or have the rights to use in a video. The API refuses a request without the rights
statement and writes it (not verified) into the signed record.
- The AI label ("AI-generated visuals Made with AI", on a black bar at the top) is burned into every frame and each file carries a C2PA credential; keep both.
YouTube, TikTok and Meta ask creators to disclose realistic AI content when they upload: use their toggle as well.
- Bind every port to 127.0.0.1. Logs carry hashes and counts, never lyrics.
## 1. Check the machine
1. `nvidia-smi`: the video model runs in fp8 through ComfyUI, which offloads the expert it is not using; on our RTX PRO
6000 (96 GB) it shared the card with other services and took about 85 s per 5 s clip. A 48 GB card should work;
smaller cards are untested. The text model needs about 20 GB more (Qwen3.8-27B NVFP4), or a server you already run.
2. `docker --version`, `docker compose version`; if Docker or the NVIDIA container toolkit is missing, ask me, then
install them from the official repositories and check `docker run --rm --gpus all ubuntu nvidia-smi`.
3. Disk: about 110 GB (VACE experts 2x34.7 GB, umt5 11 GB, Qwen 20 GB, aligner 1.3 GB). CPU: 8+ cores for the analyzer.
## 2. Models and ComfyUI
Create `~/decosa/mvideo/models`, `pip install -U huggingface_hub`, then:
- `hf download alibaba-pai/Wan2.2-VACE-Fun-A14B --revision 4438caabbfdd7437bbb0283d86cb200d1b7223ba --local-dir models/vace-fun-a14b`
- `hf download lightx2v/Wan2.2-Lightning --local-dir models/lightning --include "Wan2.2-T2V-A14B-4steps-lora-250928/*"` (record the revision you got)
- `hf download Comfy-Org/Wan_2.1_ComfyUI_repackaged --revision 123acf1cc74bccbb9bfff8ac1ee72edc08c2341d --local-dir models/wan21 --include "split_files/text_encoders/umt5_xxl_fp16.safetensors" "split_files/vae/wan_2.1_vae.safetensors"`
Build `decosa-comfyui:local` as in the Decosa **characters** assemble prompt (ComfyUI at commit
`30bdda1ef13a3a34fce2cd2fec633f15d832122a`, torch with CUDA for your card) and expose these names through
`extra_model_paths.yaml`: `diffusion_models/vace-fun-a14b-high-noise.safetensors`, `vace-fun-a14b-low-noise.safetensors`,
`loras/wan22_lightning_t2v_4step_250928_high.safetensors`, `_low.safetensors`, `text_encoders/umt5_xxl_fp16.safetensors`,
`vae/Wan2.1_VAE.pth`. The graph the worker sends is `services/mvideo/workflows/vace-a14b-t2v-lightning.api.json`.
## 3. Images
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
`git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a release that contains
`decosa_api/verticals/mvideo/`, and `docker build -f docker/api/Dockerfile -t decosa-api:mvideo .`. The image already
carries FFmpeg (with libass), fonts-dejavu-core, Tesseract and c2pa-python, which the edit, the label check and the
credentials need.
- The analyzer, from the same checkout: `docker build -f services/mvideo/Dockerfile -t decosa-mvideo-analyze:local .`
(CPU torch 2.14.0, transformers 5.17.0, librosa 1.0.0; the aligner downloads on first start).
## 4. docker-compose.yml
Write this in `~/decosa/mvideo/`. `network_mode: host` lets the api reach ComfyUI and your model server on 127.0.0.1.
```yaml
services:
analyze:
image: decosa-mvideo-analyze:local
network_mode: host
environment: { MVIDEO_ANALYZE_HOST: 127.0.0.1, MVIDEO_ANALYZE_PORT: "8488", MVIDEO_ANALYZE_THREADS: "8" }
volumes: ["hf-cache:/hf"]
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8488/health', timeout=4)"], interval: 30s, retries: 20, start_period: 600s }
api:
image: decosa-api:mvideo
network_mode: host
environment:
DECOSA_HOST: 127.0.0.1
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://127.0.0.1:8114/v1 # your Qwen3.8-27B server (vLLM, served name qwen3.8-27b)
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_STUDIO_WORKER: command
DECOSA_STUDIO_WORKER_CMD: python /app/scripts/studio_worker.py
DECOSA_STUDIO_COMFY_URL: http://127.0.0.1:8188
DECOSA_STUDIO_WORK_DIR: /work
DECOSA_MVIDEO_ANALYZE_URL: http://127.0.0.1:8488
DECOSA_MVIDEO_MAX_VIDEO_S: "60" # raise for whole songs (about 85 s of GPU per scene per format)
DECOSA_MVIDEO_GPU_USD_PER_HOUR: "1.69" # what your GPU hour costs you, for the estimates
DECOSA_PROVENANCE_DIR: /provenance
DECOSA_CLEARANCE_LOAD: "0" # "1" with a sample-clearance catalogue (tool 38) to run the pre-check
volumes: ["decosa-data:/data", "decosa-work:/work", "decosa-provenance:/provenance"]
depends_on: { analyze: { condition: service_healthy } }
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/mvideo/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
decosa-data:
decosa-work:
decosa-provenance:
hf-cache:
```
Use named volumes, never host folders (the api runs as uid 10001). `docker compose up -d` and wait for both health
checks. Create the C2PA signing material once, then restart: `docker compose exec api python scripts/provenance_devcert.py &&
docker compose restart api`; `GET /provenance/status` must show `"c2pa": true`. It is a development CA: public validators
show the signature as valid and the issuer as untrusted. Back up the box's Ed25519 key: `docker compose cp api:/data/attest ./attest-backup`.
## 5. Smoke test
```bash
B=http://127.0.0.1:8445
T=$(curl -s -XPOST $B/demo/session -H 'content-type: application/json' -d '{"vertical":"music-video-studio"}' | jq -r .token)
curl -s -XPOST $B/mvideo/analyze -H "authorization: Bearer $T" -H 'content-type: application/json' \
-d '{"sample":"rxbyn-bad-side","stream":false}' | jq .error # refused: no rights statement (403)
curl -s -XPOST $B/mvideo/analyze -H "authorization: Bearer $T" -H 'content-type: application/json' \
-d '{"sample":"rxbyn-bad-side","stream":false,"rights":{"confirmed":true,"role":"open-licence"}}' > run.json
jq '{bpm: .analysis.tempo_bpm, lines: (.lyrics.lines|length), scenes: [.scenes[].verdict], cuts: .plan.cuts, est: .plan.estimate.render_minutes}' run.json
R=$(jq -r .run_id run.json); W=$(jq -r .watch run.json)
curl -s -XPOST $B/mvideo/runs/$R/render -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"aspects":["9:16"]}' | jq .status
until curl -s "$B/mvideo/runs/$R?watch=$W" | jq -e '.render.status=="done" or .render.status=="failed"' >/dev/null; do sleep 20; done
curl -s "$B/mvideo/runs/$R?watch=$W" | jq '.render | {status, measured, outputs: (.outputs|map_values({url, credential, cut: .cut_accuracy.median_abs_ms}))}'
curl -s $B/mvideo/certificates/$R | jq '{record: .}' | curl -s -XPOST $B/record/verify -H 'content-type: application/json' -d @- | jq .ok
```
Expect about 112 BPM, 18 lines, every scene `grounded` or `weak` with its lines, 15+ cuts; then a 9:16 file with
`credential: "c2pa"`, cuts within a few tens of ms of the beat, and a record that verifies (`true`). Report the times you
measure; ours were 20-35 s for the analysis on a shared gateway and about 85 s per clip on a shared GPU.
## 6. Your own tracks
Send `audio_b64` (WAV, FLAC, MP3, OGG or M4A, 20 MB), `lyrics` (English; `[Chorus]` tags help; LRC timestamps are used
as they are), `bpm` from your session (strongly recommended: without it the beat tracker can lock to half or double time),
and `start_s`/`end_s` for the part you want. Download `/mvideo/runs/{id}/captions?format=srt|lrc|ass` to reuse the timing.
## 7. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`. Contract: `API_CONTRACT.md`, section
"Music video from your track".What it does, in shortWho it's for, where it runs and the key results
A first music video for an independent artist's own track. It finds the beats, bars and sections, times the artist's lyrics to the audio, runs the sample-clearance pre-check, and has an open model write scenes that each cite the lines they illustrate; code and a receipted yes/no check test every citation. Clips render with an open video model, cut on the bar lines, with the lyrics as karaoke captions, in 16:9 and 9:16. Each file carries a C2PA credential and a burned-in AI label, and a signed record holds the rights statement, the checks and the measured cut timing. The visuals are 480p upscaled: a stylised first video, not a studio shoot.
In short
Last reviewed
- What it is
- A first music video from an artist's own track: cuts on the beat, the lyrics as timed captions, scenes written from the lyrics, labelled and credentialed.
- Who it's for
- Teams in music and film, tv and games.
- Where it runs
- Hosted (the artist's own or openly licensed tracks) or self-host
- Key numbers
- 81.8% / 87.7% Lyric lines within 0.3 s / 1 s of human timing (test split, n = 6)
- 0.868 Beat F-measure (±70 ms), artist's BPM given (test split, n = 40)
- 67.3% Typed check says no to a swapped (mismatched) scene (test split, n = 52)
- 15.1 s Median end-to-end run, hosted (QA sweep 2026-09-26)
How we tested itEnd-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 15 s · ~$0.002 per run · 7 receipts
Loading the nightly status…
Self-host: verified 26 Sep 2026 · Fresh clone of the branch into a clean directory, docker build of the api image and services/mvideo, compose with named volumes on host networking, pointed at the already-running local Qwen (direct route) and ComfyUI; then torn down.
Measured cost to run: about $0.19 per minute of video (hosted, 26 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
No rights statement: 403; the sample analysed in 19 s (112.35 BPM, 18 lines, 6 grounded scenes, 20 cuts, 7 attested receipts); a 9:16 render finished with a C2PA credential, the disclosure rules met and a record that verified. Found on the way: the containerised analyzer's beat grid sat about 160 ms later than the host's on the same file (different decoder), and a clip of flickering neon fooled the cut detector (fixed: isolated spikes only).
Known limits (6)
- Visuals are 480p clips upscaled to 720p: a stylised visualiser, clearly below closed video models and the H3 samples on this site.
- Lyric timing is English only; 82% of held-out lines start within 0.3 s, and when it slips whole passages slip.
- Without the artist's BPM the beat tracker got the tempo right on 43% of test grooves (100% with it); which beat is "one" is a heuristic.
- The hosted demo caps a video at 60 s and renders share one GPU: 3 renders per session, 3 per API key per day, about 11 minutes per minute of video per format.
- The rights statement is not verified; the clearance pre-check covers only a small open catalogue.
- C2PA credentials are signed by a development CA: valid signature, untrusted issuer in public validators.
Eval results, nightly checks and cost per run · held-out eval
Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
2 laws, rules and guidance pages cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels, what it is built from
- Models
- Qwen3.8-27B writes and checks the treatment; Wan2.2-VACE-Fun-A14B renders; wav2vec2 aligns the lyrics; librosa finds the beat
- Where
- Hosted (the artist's own or openly licensed tracks) or self-host
- Checks
- Receipt per model call; render receipt and C2PA credential per file; signed record with the rights statement and checks
- Industry
- Music · Film, TV and games
- Output
- Media · Signed record or verdict
- Data
- Confidential business data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Part of
- Decosa Studio: Music
- Runs in
- Decosa hosted · Self-host
- Built from
- Studio render · Typed judgment · Content credentials · Signed record
Every result carries a signed record of which model produced it, so you can check it later. How that works
Questions people ask
How good are the visuals?
They are 480p clips upscaled to 720p: a stylised visualiser, clearly below closed video models. It is a first video or a Shorts teaser, not a studio shoot.
How accurate are the beat cuts and lyric timing?
With the artist's BPM, beat F-measure was 0.868 on 40 test grooves (0.591 without it). 81.8% of held-out lyric lines started within 0.3 s of human timing (6 English songs); when it slips, whole passages slip.
Is the video labelled as AI?
Yes. Every frame carries a visible AI label and every file a C2PA credential declaring composite AI media. YouTube also asks creators to disclose realistic synthetic content at upload, so use each platform's toggle as well.
Who owns the rights?
You confirm you own the track or hold the rights; the statement is signed into the record but not verified. The record claims no copyright in the generated visuals, following the US Copyright Office's Part 2 report.
Does it check my track for samples?
It runs the sample-clearance pre-check, which compares the upload with a small open catalogue only; a clean result says nothing about commercial music.
What does it cost and how long does it take?
About 15 s for analysis and treatment of a 50-60 s excerpt, and about 11 minutes to render it in one format on a shared GPU. Rendering one format costs about $0.32 per minute of video at an assumed $1.69 per GPU-hour; hosted demo renders are not billed.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Music video from your track
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…