Skip to content
decosa

52 · Music · Film, TV and games · live

Music video from your track

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)

  • Lyric lines within 0.3 s / 1 s of human timing81.8% / 87.7%test splitn = 66 held-out songs; 4 of 6 songs have at least 80% of lines within 1 s. Baseline 16.8% within 1 s. Dev: 67.0% / 67.9%.
  • Beat F-measure (±70 ms), artist's BPM given0.868test splitn = 40Without BPM: 0.591. Drums-only grooves: an upper bound for full mixes.
  • Typed check says no to a swapped (mismatched) scene67.3%test splitn = 52The code checks (citations, anchor words) carry the grounding; the typed check lets a third through.
  • Anchor words found in the cited lines96.2%test splitn = 52Citations valid: 100%
  • Rendered cut offset from the beat grid (median / max)7.0-10.0 ms / 16.3 mssynthetic5 renders of one excerpt; against the detected grid, not a human one. Some cuts between similar shots are not detectable (13-21 of 20-21 found).
  • Render cost, one formatabout $0.32 per minute of videosyntheticAt an assumed $1.69/h GPU rental price; the hosted demo does not bill renders.

Dataset

JamendoLyrics MultiLang (9 English songs with human word/line timings: 3 dev, 6 test), Groove MIDI Dataset (40 test grooves; tuned on the validation split), and renders of one CC BY excerpt and one generated track.

Caveats

  • Visual quality is not scored; 480p upscaled clips are clearly softer than closed models.
  • Beat tracking measured on synthesised drums only, not real full mixes.
  • English only (the aligner is English-only).
  • Small sets: 6 test songs; the cut-detector threshold was set on the same render it was measured on.
  • The containerised analyzer placed one beat grid about 160 ms later than on the host.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
26 Sep 2026
Latency, this run
n/a
p50 over passed runs
15 s
Receipts
7
Model calls
n/a
Tokens
n/a
Cost per run
$0.002

Self-host verification

Verified on 26 Sep 2026: Fresh clone of the branch into a clean directory, docker build of the api image and services/mvideo, compose with named volumes on host networking, pointed at the already-running local Qwen (direct route) and ComfyUI; then torn down.

No rights statement: 403; the sample analysed in 19 s (112.35 BPM, 18 lines, 6 grounded scenes, 20 cuts, 7 attested receipts); a 9:16 render finished with a C2PA credential, the disclosure rules met and a record that verified. Found on the way: the containerised analyzer's beat grid sat about 160 ms later than the host's on the same file (different decoder), and a clip of flickering neon fooled the cut detector (fixed: isolated spikes only).

Rehearsal bundle: music-video-studio.zip (825 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Visuals are 480p clips upscaled to 720p: a stylised visualiser, clearly below closed video models and the H3 samples on this site.
  • Lyric timing is English only; 82% of held-out lines start within 0.3 s, and when it slips whole passages slip.
  • Without the artist's BPM the beat tracker got the tempo right on 43% of test grooves (100% with it); which beat is "one" is a heuristic.
  • The hosted demo caps a video at 60 s and renders share one GPU: 3 renders per session, 3 per API key per day, about 11 minutes per minute of video per format.
  • The rights statement is not verified; the clearance pre-check covers only a small open catalogue.
  • C2PA credentials are signed by a development CA: valid signature, untrusted issuer in public validators.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Beat, bar and section detection (CPU): librosa beat tracker on a full-band plus low-band onset envelope, bar phase by a kick-and-snare heuristic (4/4), sections by checkerboard novelty on chroma and MFCC self-similaritydecosa-mvideo-analyze (services/mvideo)AGPL-3.0-or-later (decosa-api)
  • Lyric timing (CPU): CTC forced alignment of the artist's own lyrics on the mix, with a repair pass for lines squeezed into too little time; English letterswav2vec2-large-960h-lv60-selfApache-2.0
  • Treatment writer (scenes that cite the lyric lines they show) and the typed yes/no grounding check per scene; also labels near-duplicate lyric lines in the clearance pre-checkQwen3.8-27B (NVFP4)Apache-2.0
  • Sample and lyric clearance pre-check on the upload (tool 38, run in-process): audio landmarks and melody against a small open catalogue, lyric lines against a lyric setdecosa-api clearance module (decosa_api/verticals/clearance)AGPL-3.0-or-later
  • Clips: text-to-video, one 5 s clip per scene per formatWan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAsApache-2.0
  • The edit and the marks (CPU): cuts on bar lines on a 30 fps grid, karaoke captions (ASS, libass), the AI label on a top bar, the credit, the artist's audio; cut timing measured back from the pixels; the disclosure pre-flight's rules (tool 49) on the file; C2PA credential per file; the signed recorddecosa-api mvideo module + FFmpeg + c2pa-pythonAGPL-3.0-or-later

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · timed captions and a cited treatment, no GPU for video (3)
  • Lyric line starts within 0.3 s / 1 s of human timing (6 held-out CC BY-ND songs): 81.8% / 87.7% (baseline 16.8% within 1 s)decosa-api docs/evals/music-video-studio.md, 2026-09-26
  • Beat F-measure (±70 ms), 40 human-played grooves: 0.87 with the artist's BPM, 0.59 withoutdecosa-api docs/evals/music-video-studio.md, 2026-09-26; drums only
  • Treatment citations valid / anchor words found (52 held-out scenes): 100% / 96%decosa-api docs/evals/music-video-studio.md, 2026-09-26
Standard · the hosted demo, clips on one 96 GB card (5)
  • Cuts found in the rendered files / offset from the beat grid: 133 of 145 planned cuts found, 0 false; median 7-10 ms, max 16.3 ms (half a frame at 30 fps)decosa-api docs/evals/music-video-studio.md, 2026-09-26; 7 files from 5 renders; offsets against the detected beat grid
  • GPU time per minute of video: about 680 GPU-seconds for one format, 1,360-1,740 for bothdecosa-api docs/evals/music-video-studio.md, 2026-09-26; shared card
  • Disclosure rules on the delivered files (tool 49): label read on 6 of 6 sampled frames and the C2PA marking valid in all 4 files checkeddecosa-api docs/evals/music-video-studio.md, 2026-09-26
  • Typed grounding check: swapped (mismatched) scenes caught: 67% (the code checks on citations and anchor words do the rest)decosa-api docs/evals/music-video-studio.md, 2026-09-26
  • Visual quality: not measured; 480p upscaled, clearly below closed models
Best · MiniMax H3 clips (licence pending), self-host (1)
  • clip quality: not measured yet
Wanted · MiniMax H3 on two cards, no offload (1)
  • render time per clip against one card with offload: not measured yet

How we measure · All tools