{"schema_version":"1","site":"https://decosa.ai","id":"music-video-studio","num":"52","name":"Music video from your track","tool_name":"Make a music video for your track","short":"Music video","blurb":"A first music video for an independent artist's own track. It finds the beats, bars and sections, times the artist's lyrics to the audio, runs the sample-clearance pre-check, and has an open model write scenes that each cite the lines they illustrate; code and a receipted yes/no check test every citation. Clips render with an open video model, cut on the bar lines, with the lyrics as karaoke captions, in 16:9 and 9:16. Each file carries a C2PA credential and a burned-in AI label, and a signed record holds the rights statement, the checks and the measured cut timing. The visuals are 480p upscaled: a stylised first video, not a studio shoot.","status":"live","labels":{"industry":["music","entertainment"],"job":["generate","review"],"input":["files","text"],"deploy":["hosted","selfhost"],"status":"live","output":["media","record"],"data":["confidential"],"hardware":"gpu-96","licence":"permissive"},"industries":["music","entertainment"],"runs_in":["hosted","selfhost"],"part_of":["studio-music"],"built_from":["studio-render","typed-judgment","content-credentials","signed-record"],"models":"Qwen3.8-27B writes and checks the treatment; Wan2.2-VACE-Fun-A14B renders; wav2vec2 aligns the lyrics; librosa finds the beat","where":"Hosted (the artist's own or openly licensed tracks) or self-host","hardware":"1× RTX PRO 6000 (96 GB) shared with the text model, or a 48 GB card for the video model plus a remote text model; the analysis runs on CPU","final_artifact":"Two MP4s (16:9 and 9:16) with captions, an AI label and a C2PA credential, SRT and LRC captions, and a signed record of the rights statement, checks and cut timing.","self_host_first":false,"verification":{"receipt_coverage":"partial","summary":"Receipt per model call; render receipt and C2PA credential per file; signed record with the rights statement and checks","manual_qa":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":15100,"p95_ms":null,"runs":null,"receipts_per_run":7,"cost_per_run_usd":0.0021},"selfhost":{"date":"2026-09-26","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of the api image and services/mvideo, compose with named volumes on host networking, pointed at the already-running local Qwen (direct route) and ComfyUI; then torn down.","notes":"No rights statement: 403; the sample analysed in 19 s (112.35 BPM, 18 lines, 6 grounded scenes, 20 cuts, 7 attested receipts); a 9:16 render finished with a C2PA credential, the disclosure rules met and a record that verified. Found on the way: the containerised analyzer's beat grid sat about 160 ms later than the host's on the same file (different decoder), and a clip of flickering neon fooled the cut detector (fixed: isolated spikes only)."},"known_limits":["Visuals are 480p clips upscaled to 720p: a stylised visualiser, clearly below closed video models and the H3 samples on this site.","Lyric timing is English only; 82% of held-out lines start within 0.3 s, and when it slips whole passages slip.","Without the artist's BPM the beat tracker got the tempo right on 43% of test grooves (100% with it); which beat is \"one\" is a heuristic.","The hosted demo caps a video at 60 s and renders share one GPU: 3 renders per session, 3 per API key per day, about 11 minutes per minute of video per format.","The rights statement is not verified; the clearance pre-check covers only a small open catalogue.","C2PA credentials are signed by a development CA: valid signature, untrusted issuer in public validators."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Lyric lines within 0.3 s / 1 s of human timing","value":"81.8% / 87.7%","unit":null,"n":6,"split":"test","note":"6 held-out songs; 4 of 6 songs have at least 80% of lines within 1 s. Baseline 16.8% within 1 s. Dev: 67.0% / 67.9%."},{"name":"Beat F-measure (±70 ms), artist's BPM given","value":"0.868","unit":null,"n":40,"split":"test","note":"Without BPM: 0.591. Drums-only grooves: an upper bound for full mixes."},{"name":"Typed check says no to a swapped (mismatched) scene","value":"67.3%","unit":null,"n":52,"split":"test","note":"The code checks (citations, anchor words) carry the grounding; the typed check lets a third through."},{"name":"Anchor words found in the cited lines","value":"96.2%","unit":null,"n":52,"split":"test","note":"Citations valid: 100%"},{"name":"Rendered cut offset from the beat grid (median / max)","value":"7.0-10.0 ms / 16.3 ms","unit":null,"n":null,"split":"synthetic","note":"5 renders of one excerpt; against the detected grid, not a human one. Some cuts between similar shots are not detectable (13-21 of 20-21 found)."},{"name":"Render cost, one format","value":"about $0.32 per minute of video","unit":null,"n":null,"split":"synthetic","note":"At an assumed $1.69/h GPU rental price; the hosted demo does not bill renders."}],"dataset":"JamendoLyrics MultiLang (9 English songs with human word/line timings: 3 dev, 6 test), Groove MIDI Dataset (40 test grooves; tuned on the validation split), and renders of one CC BY excerpt and one generated track.","held_out":true,"caveats":["Visual quality is not scored; 480p upscaled clips are clearly softer than closed models.","Beat tracking measured on synthesised drums only, not real full mixes.","English only (the aligner is English-only).","Small sets: 6 test songs; the cut-detector threshold was set on the same render it was measured on.","The containerised analyzer placed one beat grid about 160 ms later than on the host."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/music-video-studio"},"stack":{"summary":"An independent artist uploads a track they own, with its lyrics. A CPU analyzer finds the beat, bars and sections and times each lyric line to the audio; the sample-clearance pre-check looks for samples and lifted lyrics. An open text model writes a treatment whose scenes each cite the lines they show, code checks every citation and a receipted yes/no call checks each scene. An open video model renders a clip per scene, and the edit cuts on the bar lines with karaoke captions, in 16:9 and 9:16. Every file has a burned-in AI label and a C2PA credential, and a signed record holds the rights statement, the checks and the measured cut timing. It is a stylised first video at 480p upscaled, not a studio shoot.","tagline":"A first music video from an artist's own track: cuts on the beat, the lyrics as timed captions, scenes written from the lyrics, labelled and credentialed.","deployment":"hosted-or-self-host","regulatory_note":"Not legal advice; checked on 26 Sep 2026 unless stated. Rights: the artist confirms they own the track or have the rights to make and publish a video with it; the statement goes into the signed record and is not verified. The sample-clearance pre-check compares the upload with a small open catalogue only, and a clean result says nothing about commercial music (US courts disagree on short samples: Bridgeport v. Dimension Films, 6th Cir. 2005, against VMG Salsoul v. Ciccone, 9th Cir. 2016). Demo tracks are CC BY excerpts from JamendoLyrics, credited on screen, and one song made with MiniMax-Music3 through the music-gen-cleared tool. Disclosure: every file has a visible \"AI-generated visuals / Made with AI\" label on every frame and a C2PA credential that declares composite AI media. EU AI Act (Regulation (EU) 2024/1689) Art. 50(2) asks providers to mark generated video in a machine-readable way (from 2 Aug 2026; checked by the disclosure pre-flight on 25 Sep 2026; EUR-Lex link under Tools). New York General Business Law § 396-b requires a conspicuous disclosure of synthetic performers in advertisements (in force 9 Jun 2026; bill text linked under Tools): it applies if the video is used as an ad, and the console runs that rule on the file. YouTube asks creators to disclose realistic altered or synthetic content when they upload and labels it (help page read 26 Sep 2026, linked under Tools); TikTok and Meta have their own AI labels (not checked here): use each platform's toggle as well. Copyright: the US Copyright Office's Part 2 report (29 Jan 2025) says AI output is protected only where a human determined sufficient expressive elements; the record claims no copyright in the generated visuals. Scenes never depict real people, brands or readable text by prompt rule and screen.","components":[{"id":"analyzer","role":"Beat, bar and section detection (CPU): librosa beat tracker on a full-band plus low-band onset envelope, bar phase by a kick-and-snare heuristic (4/4), sections by checkerboard novelty on chroma and MFCC self-similarity","name":"decosa-mvideo-analyze (services/mvideo)","hf_repo":null,"license":"AGPL-3.0-or-later (decosa-api)","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12, librosa 1.0.0 (ISC), numpy, scipy; 8-16 threads","receipt_coverage":"none","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"aligner","role":"Lyric timing (CPU): CTC forced alignment of the artist's own lyrics on the mix, with a repair pass for lines squeezed into too little time; English letters","name":"wav2vec2-large-960h-lv60-self","hf_repo":"facebook/wav2vec2-large-960h-lv60-self","license":"Apache-2.0","params":"317M","quant":"FP32 on CPU","vram_gb":0,"memory_gb_estimate":null,"engine":"transformers 5.17.0, torch 2.14.0 (CPU); own Viterbi in numpy","receipt_coverage":"none","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"llm","role":"Treatment writer (scenes that cite the lyric lines they show) and the typed yes/no grounding check per scene; also labels near-duplicate lyric lines in the clearance pre-check","name":"Qwen3.8-27B (NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP NVFP4, GDN/attention FP8) + FP8 KV cache; MTP head, 3 draft tokens","vram_gb":20,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, thinking off; treatment at temperature 0.3 (one retry at 0), checks at 0","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["lite","standard"],"alternative_to":null},{"id":"clearance","role":"Sample and lyric clearance pre-check on the upload (tool 38, run in-process): audio landmarks and melody against a small open catalogue, lyric lines against a lyric set","name":"decosa-api clearance module (decosa_api/verticals/clearance)","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"numpy code; near-duplicate lyric lines labelled by the text model (receipted)","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"video","role":"Clips: text-to-video, one 5 s clip per scene per format","name":"Wan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAs","hf_repo":"alibaba-pai/Wan2.2-VACE-Fun-A14B","license":"Apache-2.0","params":"A14B (two 14B experts, high and low noise)","quant":"FP8 (e4m3fn) weights in ComfyUI; BF16 files on disk (34.7 GB per expert)","vram_gb":null,"memory_gb_estimate":null,"engine":"ComfyUI core WanVaceToVideo without a control video (text only), 832x480 or 480x832, 81 frames at 16 fps, 4 steps (2 per expert), cfg 1, euler/simple, shift 5, through the studio queue","receipt_coverage":"none","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"edit","role":"The edit and the marks (CPU): cuts on bar lines on a 30 fps grid, karaoke captions (ASS, libass), the AI label on a top bar, the credit, the artist's audio; cut timing measured back from the pixels; the disclosure pre-flight's rules (tool 49) on the file; C2PA credential per file; the signed record","name":"decosa-api mvideo module + FFmpeg + c2pa-python","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"FFmpeg (libx264, libass), c2pa-python, Tesseract for the label check","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["standard"],"alternative_to":null},{"id":"ltx","role":"Alternate, self-host only: sharper clips","name":"LTX-2.3 22B (distilled)","hf_repo":"Lightricks/LTX-2.3","license":"LTX-2 Community License","params":"22B","quant":"BF16","vram_gb":null,"memory_gb_estimate":null,"engine":"ComfyUI core LTX nodes","receipt_coverage":"none","in_hosted_demo":false,"tiers":["alternates"],"alternative_to":null},{"id":"h3","role":"A clip per scene with native audio, in place of Wan2.2","name":"MiniMax-H3 (licence pending)","hf_repo":"MiniMaxAI/MiniMax-H3","license":"MiniMax H3 Community License today (it excludes the US, EU, UK and Korea); a licence for our use is pending","params":"33.1B (transformer) + 33.4B (text encoder)","quant":"BF16 with CPU offload on one card; no offload on two","vram_gb":null,"memory_gb_estimate":133,"engine":"diffusers ModularPipeline (MiniMax-H3 blocks)","receipt_coverage":"none","in_hosted_demo":false,"tiers":["best"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · timed captions and a cited treatment, no GPU for video","summary":"Beat grid, lyric timing (SRT, LRC, ASS), the clearance pre-check and a checked treatment with an edit plan. Bring your own footage.","components":["analyzer","aligner","llm","clearance"],"hardware":"CPU (8+ cores) plus the text model (about 20 GB of GPU, or a remote server)","quality_evidence":[{"metric":"Lyric line starts within 0.3 s / 1 s of human timing (6 held-out CC BY-ND songs)","value":"81.8% / 87.7% (baseline 16.8% within 1 s)","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26"},{"metric":"Beat F-measure (±70 ms), 40 human-played grooves","value":"0.87 with the artist's BPM, 0.59 without","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26; drums only"},{"metric":"Treatment citations valid / anchor words found (52 held-out scenes)","value":"100% / 96%","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26"}],"latency_note":"measured: seconds per excerpt on the shared gateway","in_hosted_demo":true,"receipt_coverage":"partial","receipt_note":null,"hosting":null},{"id":"standard","label":"Standard · the hosted demo, clips on one 96 GB card","summary":"Everything in Lite plus a clip per scene on Wan2.2-VACE (Apache-2.0), cut on the bar lines in 16:9 and 9:16, labelled and credentialed.","components":["analyzer","aligner","llm","clearance","video","edit"],"hardware":"1x RTX PRO 6000 Blackwell 96 GB (shared), CPU for the analyzer and the edit","quality_evidence":[{"metric":"Cuts found in the rendered files / offset from the beat grid","value":"133 of 145 planned cuts found, 0 false; median 7-10 ms, max 16.3 ms (half a frame at 30 fps)","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26; 7 files from 5 renders; offsets against the detected beat grid"},{"metric":"GPU time per minute of video","value":"about 680 GPU-seconds for one format, 1,360-1,740 for both","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26; shared card"},{"metric":"Disclosure rules on the delivered files (tool 49)","value":"label read on 6 of 6 sampled frames and the C2PA marking valid in all 4 files checked","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26"},{"metric":"Typed grounding check: swapped (mismatched) scenes caught","value":"67% (the code checks on citations and anchor words do the rest)","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26"},{"metric":"Visual quality","value":"not measured; 480p upscaled, clearly below closed models","source":null}],"latency_note":"measured on our server: many minutes for a video in one format, about twice that for both, on a shared GPU","in_hosted_demo":true,"receipt_coverage":"partial","receipt_note":null,"hosting":null},{"id":"best","label":"Best · MiniMax H3 clips (licence pending), self-host","summary":"Everything in Standard with clips from MiniMax-H3 instead of Wan2.2. Self-host available where H3's current licence covers you; the pending licence would bring it to the hosted demo. Not wired in yet.","components":["analyzer","aligner","llm","clearance","h3","edit"],"hardware":"1x RTX PRO 6000 96 GB plus about 115 GB of RAM for offload","quality_evidence":[{"metric":"clip quality","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Self-host: your box signs the render record. Not a community-provider model.","hosting":null},{"id":"wanted","label":"Wanted · MiniMax H3 on two cards, no offload","summary":"A clip per scene adds up over a song. Both halves of MiniMax-H3 in BF16 on two 96 GB cards, so nothing is offloaded to RAM: faster renders (estimate, not measured). Self-host now where H3's current licence covers you; once the pending licence lands we add this compute for the hosted tier ourselves. Not a community-provider model.","components":["analyzer","aligner","llm","clearance","h3","edit"],"hardware":"2x RTX PRO 6000 96 GB, no CPU offload (about 133 GB of BF16 weights; estimate)","quality_evidence":[{"metric":"render time per clip against one card with offload","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Self-host: your box signs the render record. Not a community-provider model.","hosting":"own-hardware"}],"alternates":[{"id":"ltx","label":"Sharper clips with LTX-2.3","components":["ltx"],"hardware":"1x 96 GB card","use":"LTX-2.3 is on the box but not wired in, and self-host only: its community licence has a competing-service clause, so we can't serve it hosted.","status":"self-host only"}],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"GET /mvideo/info, /mvideo/samples; POST /mvideo/analyze (SSE), /mvideo/runs/{id}/render; GET /mvideo/runs/{id} (token or watch key), /mvideo/runs/{id}/captions, /mvideo/certificates/{id}. Needs FFmpeg, fonts, c2pa-python (and Tesseract for the label check) in the image."},{"name":"decosa-mvideo-analyze","port":8488,"image":null,"purpose":"CPU analyzer (services/mvideo/Dockerfile, built locally): POST /v1/analyze, GET /health. Keeps nothing."},{"name":"comfyui","port":8188,"image":null,"purpose":"Wan2.2-VACE-Fun-A14B renders for the studio worker, one job at a time."},{"name":"vLLM","port":8114,"image":"vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1","purpose":"Qwen3.8-27B NVFP4, behind our gateway (hosted) or called directly (self-host)."}],"tools":[{"name":"JamendoLyrics MultiLang","url":"https://github.com/f90/jamendolyrics","license":"Annotations MIT; audio CC per song (only CC BY excerpts are shown; CC BY-ND songs measured only)","purpose":"Demo excerpts and the lyric-alignment eval (9 English songs without an NC clause, human word and line timings)."},{"name":"Groove MIDI Dataset","url":"https://magenta.tensorflow.org/datasets/groove","license":"CC BY 4.0","purpose":"The beat-tracking eval: human drummers played to a click, synthesised with a small drum kit."},{"name":"FFmpeg with libass","url":"https://ffmpeg.org","license":"LGPL-2.1+ / GPL for some builds","purpose":"Cuts, captions, labels, audio, and measuring the cuts back from the pixels."},{"name":"c2pa-python","url":"https://github.com/contentauth/c2pa-python","license":"MIT OR Apache-2.0","purpose":"The C2PA content credential on every MP4 (via the provenance kit)."},{"name":"ComfyUI","url":"https://github.com/comfyanonymous/ComfyUI","license":"GPL-3.0","purpose":"Runs the VACE graph for the studio worker."},{"name":"EU AI Act, Regulation (EU) 2024/1689 (Art. 50)","url":"https://eur-lex.europa.eu/eli/reg/2024/1689/oj","license":"EU law (EUR-Lex)","purpose":"Primary source for the machine-readable marking duty (Art. 50(2)) the C2PA credential answers."},{"name":"New York S.8420-A / A.8887-B (GBL § 396-b)","url":"https://www.nysenate.gov/legislation/bills/2025/S8420/amendment/A","license":"State law","purpose":"Primary source for the synthetic-performer disclosure in advertisements, run on each file by the disclosure pre-flight."},{"name":"YouTube Help: disclosing altered or synthetic content","url":"https://support.google.com/youtube/answer/14328491","license":"Platform policy","purpose":"What YouTube asks creators to disclose at upload (read 26 Sep 2026)."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB, shared","fits":true,"notes":"Measured on our server 2026-09-26: clips on GPU0 beside other services and a sibling's renders, 84-100 s per 5 s clip; the text model on GPU1; the analyzer on CPU."},{"tier":"1x 48 GB card","fits":null,"notes":"Not tested. ComfyUI offloads the idle expert, so the fp8 VACE graph should fit; the text model would need to be remote."},{"tier":"CPU only","fits":true,"notes":"The lite tier: analysis, lyric timing, treatment (with a remote text model), captions and the edit plan, no clips."}],"latency":[{"lane":"Analysis, clearance, treatment and checks for a 50-60 s excerpt","typical_ms":15000,"source":"measured on our server 2026-09-26: 13-32 s over the recorded, e2e and eval runs on the shared gateway"},{"lane":"Lyric alignment, 3-5 minute song","typical_ms":24000,"source":"measured on our server 2026-09-26: 18-33 s on 16 CPU threads"},{"lane":"One 5 s clip (Wan2.2-VACE, 4 steps)","typical_ms":86000,"source":"measured on our server 2026-09-26: medians 84-96 s in four renders on shared GPU0 (141 s when interleaved with another job)"},{"lane":"55-60 s video in 9:16 (7 clips, edit, checks)","typical_ms":645000,"source":"measured on our server 2026-09-26: 628 s and 663 s"},{"lane":"50 s video in 16:9 and 9:16 (12-14 clips)","typical_ms":1400000,"source":"measured on our server 2026-09-26: 1,359-1,449 s"}],"benchmark":null,"notes":["The visuals are 480p clips upscaled to 720p: a stylised visualiser, clearly below closed video models and the MiniMax H3 samples elsewhere on this site.","The hosted demo caps the video at 60 s (a chorus or a teaser); self-hosted, raise DECOSA_MVIDEO_MAX_VIDEO_S."]},"buyer_facts":[{"label":"Data retention","value":"The track is kept 24 hours after upload so it can be rendered (longer only while a render is queued or running), then deleted. The run (hashes, analysis, lyric timings, captions, plan, record) is kept 7 days; the rendered videos stay in the studio's media folder and open only through signed links that expire within hours, given to the run's owner or its watch link. Logs hold ids, counts and hashes, never lyrics."},{"label":"What leaves the box","value":"On the hosted route, the lyrics and plan go to the text model through our gateway and the render runs on Decosa's hosted service. Self-hosted, nothing leaves your machine."},{"label":"Rights","value":"You confirm you own the track or hold the rights; the statement is signed into the record, not verified. The clearance pre-check covers only a small open catalogue."},{"label":"Cost per video","value":"A fraction of a cent of model calls per analysis (measured). Rendering a video in one format takes several GPU-minutes, tens of cents at a GPU rental list price; both formats about twice that. Hosted demo renders are not billed."},{"label":"Typical time","value":"Seconds for the analysis and treatment of an excerpt; many minutes to render it in one format and about twice that in both, on a shared GPU (measured)."},{"label":"Outputs","value":"MP4s in 16:9 (1280x720) and 9:16 (720x1280) at 30 fps with captions, an AI label on every frame and a C2PA credential; SRT, LRC and ASS captions; a signed record with the rights statement, checks, receipts and measured cut timing."}],"data_handling":{"page":"/data#music-video-studio","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"The track is kept 24 hours after upload so it can be rendered (longer only while a render is queued or running), then deleted. The run (hashes, analysis, lyric timings, captions, plan, record) is kept 7 days; the rendered videos stay in the studio's media folder and open only through signed links that expire within hours, given to the run's owner or its watch link. Logs hold ids, counts and hashes, never lyrics.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/apps/music-video-studio","input":"mvideo","lanes":[{"id":"analysis","title":"Beat, bars and sections","kind":"list"},{"id":"lyrics","title":"Lyrics on the audio","kind":"list"},{"id":"treatment","title":"Treatment","kind":"list"},{"id":"video","title":"Video","kind":"list"}],"samples":[{"n":1,"id":"rxbyn-bad-side","title":"Rxbyn bad side","deep_link":"/apps/music-video-studio?sample=1&autorun=0"},{"n":2,"id":"cortez-feel","title":"Cortez feel","deep_link":"/apps/music-video-studio?sample=2&autorun=0"},{"n":3,"id":"lanterns-music3","title":"Lanterns music3","deep_link":"/apps/music-video-studio?sample=3&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/music-video-studio-hosted.md","selfhost":"/prompts/music-video-studio-selfhost.md","assemble":"/prompts/music-video-studio-assemble.md","mac":null},"rehearsal":{"bundle":"/samples/music-video-studio.zip","bundle_url":"https://decosa.ai/samples/music-video-studio.zip","folder":"/samples/music-video-studio/","expected":"/samples/music-video-studio/expected.json","files":["/samples/music-video-studio/expected.json","/samples/music-video-studio/inputs/ATTRIBUTION.txt","/samples/music-video-studio/inputs/bad-side-excerpt.ogg","/samples/music-video-studio/inputs/bad-side-lyrics.txt"],"bytes":844310,"checks":["a request without the rights statement is refused","a text file sent as audio is refused","a steady beat near 112 BPM","all 18 sung lines are placed on the audio","the first line starts within 0.5 s of the human timing (0.78 s)","the ninth line starts within 0.5 s of the human timing (18.62 s)","every scene cites lines and passes the code checks","the edit plan cuts on the bar lines at least 10 times","the plan can be rendered","every model call is receipted and signed"],"licence":"inputs/bad-side-excerpt.ogg: \"Bad Side\" by Rxbyn (Jamendo), CC BY 4.0, excerpt 0:40-1:30, faded. inputs/bad-side-lyrics.txt: the song's lyrics as normalised in JamendoLyrics (annotations MIT). Your own track and lyrics replace these files.","about":"A 50 s excerpt of a CC BY song and its lyrics. A request without the rights statement must be refused; with it, the analysis must find a steady beat near 112 BPM, place all 18 sung lines with the first and ninth line within half a second of the dataset's human timing, write a treatment whose scenes each cite lines and pass the checks, and plan a renderable edit, with every model call receipted. A text file sent as audio must be refused. No render: a render takes about 20 minutes of GPU; run it from the console or with POST /mvideo/runs/{id}/render when you are ready.","run":{"containers":"docker compose exec api python scripts/rehearse.py music-video-studio","checkout":"python scripts/rehearse.py music-video-studio --bundle music-video-studio.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py music-video-studio"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=music-video-studio","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":91.6,"basis":"stack","unknown":[]},{"id":"best","gpu_gb":153.6,"basis":"stack","unknown":[]},{"id":"wanted","gpu_gb":153.6,"basis":"stack","unknown":[]},{"id":"alternate-ltx","gpu_gb":50,"basis":"stack","unknown":[]}],"mac":null},"links":{"page":"/apps/music-video-studio","json":"/use-cases/music-video-studio.json","metrics":"/metrics/music-video-studio","console":"/apps/music-video-studio","console_sample":"/apps/music-video-studio?sample=1&autorun=0","stack":"/apps/music-video-studio#stack","try_live":"/apps/music-video-studio#live","watch":"/apps/music-video-studio#watch","build":"/apps/music-video-studio#build","self_host":"/apps/music-video-studio#self-host","prompts":{"hosted":"/prompts/music-video-studio-hosted.md","selfhost":"/prompts/music-video-studio-selfhost.md","assemble":"/prompts/music-video-studio-assemble.md","mac":null}}}