04 · Creative and media · Personal and family · live
Decosa Studio
Eval results
Not held outRun 23 Sep 2026
Quality is not measured yet for this tool.
Dataset
No quality eval. One reproducibility check on our server: same-seed re-render is bit-identical for MiniMax-Music3 and Qwen-Image-2512, not identical for ACE-Step.
Caveats
- Output quality not measured yet on any tier.
- The same-seed check shows reproducibility, not quality.
Nightly smoke check
Loading the nightly status…
- Result
- partial
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- n/a
- Receipts
- 1
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- n/a
Self-host verification
Verified on 25 Sep 2026: Fresh clone of decosa-api; the api image plus ffmpeg running the repo's own studio worker against an already-running ComfyUI on the same box (no render or ComfyUI containers built, no new model loads).
Verified on 2026-09-25: the image builds, the service starts, and the prompt's MiniMax-Music3 smoke job (seed 5501) finished in 40 s against a local ComfyUI equivalent to the documented one; model-server startup and the image and video paths were not re-verified. Two runs of the same seed gave bit-identical decoded PCM. An api restart mid-render marked that job failed with a reason and ran the queued one (72 s). Fixed on the way: the api image left out services/ (every render failed), jobs were lost on restart.
Rehearsal bundle: studio.zip (942 KB, 11 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Renders share one GPU with the live audio demos: a job pauses while a live session runs.
- Each token or key can queue 3 jobs; hosted video needs an API key and fal credit.
- Receipts for renders are our signed statement of model, seed, weights and output hash; nobody re-renders to check them.
- Wan2.1 video takes about 35 minutes per 5 s clip on the shared GPU.
Models and licences
Standard tier. Licence posture: community licence (check the terms).
- Music generation (songs with vocals and lyrics)MiniMax-Music3MiniMax-Music3 Community License
- Image generationQwen-Image-2512Apache-2.0
- Hosted video (default): 5 s clips with audioMiniMax H3 Max (via fal)H3's own licence excludes the US; used here only through fal's hosted endpoints, listed as commercial use under fal's MiniMax partnership
- Text to speech (preset voices)Kokoro-82MApache-2.0
- Voice cloning (consent-gated)IndexTTS-2.5bilibili Model Use License (commercial use allowed below 100M MAU and RMB 1B yearly revenue; may not be used to improve other models)
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · drafts on one 24–32 GB card (1)
- output quality: not measured yet
Standard · the hosted demo, about 50 GB of one 96 GB card (2)
- output quality: not measured yet
- same-seed re-render: bit-identical for MiniMax-Music3 and Qwen-Image-2512; not identical for ACE-Stepmeasured on our server 2026-09-23
Best · self-host only, a whole 96 GB card (1)
- output quality: not measured yet
Wanted · MiniMax H3 on two cards, no offload (1)
- render time per clip against one card with offload: not measured yet