Skip to content
decosa

Make a music video for your track

A first music video for an independent artist's own track. It finds the beats, bars and sections, times the artist's lyrics to the audio, runs the sample-clearance pre-check, and has an open model write scenes that each cite the lines they illustrate; code and a receipted yes/no check test every citation. Clips render with an open video model, cut on the bar lines, with the lyrics as karaoke captions, in 16:9 and 9:16. Each file carries a C2PA credential and a burned-in AI label, and a signed record holds the rights statement, the checks and the measured cut timing. The visuals are 480p upscaled: a stylised first video, not a studio shoot.

For
Teams in music and film, tv and games.
  • Time per task15 stypical (median) on the sample
  • Cost per task~$0.19 per minute of videomeasured, at list price
  • Accuracy81.8% / 87.7%Lyric lines within 0.3 s / 1 s of human timing (held-out test)All results and caveats
  • Self-hostRun it on your own server: nothing is sent to us or anyone else.
  • HostedUse it on our servers; see where your data goes.

Your track

"Bad Side" by Rxbyn (Jamendo), CC BY; excerpt 0:40.0-1:30.0, faded in and out
Formats
RightsYour statement goes into the signed record. We do not verify it; the clearance pre-check only compares the track with a small open catalogue.

Confirm that you hold the rights to this track.

The visuals come from an open video model at 480p, upscaled: a stylised first video, not a studio shoot. Every file carries a C2PA credential and a burned-in "AI-generated visuals · Made with AI" bar. Your track is kept 24 hours so it can be rendered, then deleted. Renders share one GPU: about 20 minutes for 50 seconds of video in both formats.

Result

Live

Pick a sample or upload your own track, confirm the rights, and analyse. Nothing renders until you ask.

What it does, in shortWho it's for, where it runs and the key results

A first music video for an independent artist's own track. It finds the beats, bars and sections, times the artist's lyrics to the audio, runs the sample-clearance pre-check, and has an open model write scenes that each cite the lines they illustrate; code and a receipted yes/no check test every citation. Clips render with an open video model, cut on the bar lines, with the lyrics as karaoke captions, in 16:9 and 9:16. Each file carries a C2PA credential and a burned-in AI label, and a signed record holds the rights statement, the checks and the measured cut timing. The visuals are 480p upscaled: a stylised first video, not a studio shoot.

In short

Last reviewed

What it is
A first music video from an artist's own track: cuts on the beat, the lyrics as timed captions, scenes written from the lyrics, labelled and credentialed.
Who it's for
Teams in music and film, tv and games.
Where it runs
Hosted (the artist's own or openly licensed tracks) or self-host
Key numbers
  • 81.8% / 87.7% Lyric lines within 0.3 s / 1 s of human timing (test split, n = 6)
  • 0.868 Beat F-measure (±70 ms), artist's BPM given (test split, n = 40)
  • 67.3% Typed check says no to a swapped (mismatched) scene (test split, n = 52)
  • 15.1 s Median end-to-end run, hosted (QA sweep 2026-09-26)
All results, datasets and caveats
How we tested itEnd-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 15 s · ~$0.002 per run · 7 receipts

Loading the nightly status…

Self-host: verified 26 Sep 2026 · Fresh clone of the branch into a clean directory, docker build of the api image and services/mvideo, compose with named volumes on host networking, pointed at the already-running local Qwen (direct route) and ComfyUI; then torn down.

Measured cost to run: about $0.19 per minute of video (hosted, 26 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

No rights statement: 403; the sample analysed in 19 s (112.35 BPM, 18 lines, 6 grounded scenes, 20 cuts, 7 attested receipts); a 9:16 render finished with a C2PA credential, the disclosure rules met and a record that verified. Found on the way: the containerised analyzer's beat grid sat about 160 ms later than the host's on the same file (different decoder), and a clip of flickering neon fooled the cut detector (fixed: isolated spikes only).

Known limits (6)
  • Visuals are 480p clips upscaled to 720p: a stylised visualiser, clearly below closed video models and the H3 samples on this site.
  • Lyric timing is English only; 82% of held-out lines start within 0.3 s, and when it slips whole passages slip.
  • Without the artist's BPM the beat tracker got the tempo right on 43% of test grooves (100% with it); which beat is "one" is a heuristic.
  • The hosted demo caps a video at 60 s and renders share one GPU: 3 renders per session, 3 per API key per day, about 11 minutes per minute of video per format.
  • The rights statement is not verified; the clearance pre-check covers only a small open catalogue.
  • C2PA credentials are signed by a development CA: valid signature, untrusted issuer in public validators.

Eval results, nightly checks and cost per run · held-out eval

Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

2 laws, rules and guidance pages cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels, what it is built from
Models
Qwen3.8-27B writes and checks the treatment; Wan2.2-VACE-Fun-A14B renders; wav2vec2 aligns the lyrics; librosa finds the beat
Where
Hosted (the artist's own or openly licensed tracks) or self-host
Checks
Receipt per model call; render receipt and C2PA credential per file; signed record with the rights statement and checks
Output
Media · Signed record or verdict
Data
Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Every result carries a signed record of which model produced it, so you can check it later. How that works

Questions people ask

How good are the visuals?

They are 480p clips upscaled to 720p: a stylised visualiser, clearly below closed video models. It is a first video or a Shorts teaser, not a studio shoot.

How accurate are the beat cuts and lyric timing?

With the artist's BPM, beat F-measure was 0.868 on 40 test grooves (0.591 without it). 81.8% of held-out lyric lines started within 0.3 s of human timing (6 English songs); when it slips, whole passages slip.

Is the video labelled as AI?

Yes. Every frame carries a visible AI label and every file a C2PA credential declaring composite AI media. YouTube also asks creators to disclose realistic synthetic content at upload, so use each platform's toggle as well.

Who owns the rights?

You confirm you own the track or hold the rights; the statement is signed into the record but not verified. The record claims no copyright in the generated visuals, following the US Copyright Office's Part 2 report.

Does it check my track for samples?

It runs the sample-clearance pre-check, which compares the upload with a small open catalogue only; a clean result says nothing about commercial music.

What does it cost and how long does it take?

About 15 s for analysis and treatment of a 50-60 s excerpt, and about 11 minutes to render it in one format on a shared GPU. Rendering one format costs about $0.32 per minute of video at an assumed $1.69 per GPU-hour; hosted demo renders are not billed.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Music video from your track

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.