{"schema_version":"1","site":"https://decosa.ai","id":"audio-drama-studio","num":"51","name":"Audio drama and narrated story studio","tool_name":"Produce an audio drama","short":"Audio drama studio","blurb":"Paste a radio script or a prose chapter and review the parse: every speaker, line, direction and sound cue, before anything is voiced. Cast each role from consented house voices; every line passes the consent ledger right before it is spoken. Music comes from cleared cues with licence certificates, and sound from CC0 or public-domain recordings or code. The mix is mastered to a podcast or ACX-style audiobook spec, with chapters, captions, show notes with an AI disclosure, a C2PA credential, a signed record and sides for human actors.","status":"live","labels":{"industry":["entertainment","creative-media"],"job":["generate","attest"],"input":["text"],"deploy":["hosted","selfhost"],"status":"live","output":["media","record"],"data":["confidential"],"hardware":"gpu-96","licence":"permissive"},"industries":["entertainment","creative-media"],"runs_in":["hosted","selfhost"],"part_of":["studio-voice"],"built_from":["consent-gate","studio-render","content-credentials","signed-record"],"models":"Qwen3.8-27B (parse) · Kokoro-82M (voices, CPU) · ACE-Step 1.5 (music, via music-gen-cleared)","where":"Hosted with public-domain and original scripts; self-host for unpublished work","hardware":"Qwen3.8-27B (1× RTX 5090 32 GB or larger) for the parse; the voices, sound and mix run on CPU; ACE-Step only if you compose new music (a GPU with about 10 GB free)","final_artifact":"An MP3 episode with chapters (or ACX-style chapter files with credits), captions, a Podcasting 2.0 transcript and chapters, show notes, a sides pack for human actors, a C2PA credential and a signed record.","self_host_first":false,"verification":{"receipt_coverage":"partial","summary":"Receipt per model call; a signed consent decision per spoken line; C2PA credential linking every role's consent entry; signed production record","manual_qa":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":28000,"p95_ms":null,"runs":null,"receipts_per_run":1,"cost_per_run_usd":0.001},"selfhost":{"date":"2026-09-26","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of the api image and the docker/drama voice layer, compose with named volumes and host networking, pointed at the running local Qwen3.8-27B (direct route); rehearsal bundle; then torn down.","notes":"Builds took 24 s and 100 s (warm cache). The rehearsal passed 16 of 16 checks in 20.7 s, including an 18 s render on 8 CPU threads: parse, an out-of-project performer refused, the house cast allowed, loudness in spec, C2PA with consent links, and the signed record verified. No speaker model in the image, so the voice check said not run."},"known_limits":["Kokoro voices are clear but flat: deliveries change pace and level only. British voices are Kokoro's weakest.","Prose attribution is about 97% right on held-out stories: check the parse before rendering.","ACX does not accept AI narration; the audiobook mode meets the technical spec only.","C2PA credentials use a development certificate, so public validators show the issuer as untrusted.","Composing new music needs the studio GPU queue; the hosted demo uses the cue library."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Radio scripts: speaker right (code alone)","value":"100% (389/389)","unit":null,"n":389,"split":"test","note":"Spoken lines found: also 100% (389/389)"},{"name":"Sound cue mapped to the right library tag (model)","value":"99.0% (96/97)","unit":null,"n":97,"split":"test","note":"Keyword code alone: 92.8% (90/97)"},{"name":"Cue with nothing in the library left as \"no sound\"","value":"87.5% (7/8)","unit":null,"n":8,"split":"test","note":"Varies run to run: a second pass missed three such cues"},{"name":"Prose: speaker right","value":"96.8% (209/216)","unit":null,"n":216,"split":"test","note":"Dev: 94.2% (65/69)"},{"name":"Public-domain excerpts: speaker right","value":"100% (30/30)","unit":null,"n":30,"split":"test","note":"Adapted excerpts of The Red-Headed League and The Monkey's Paw"},{"name":"Word error rate of the finished episodes (machine transcription)","value":"0.6%-5.0%","unit":null,"n":4,"split":"synthetic","note":"4 sample episodes, 175-625 words each"},{"name":"Parse cost","value":"about $0.0007 per parse","unit":null,"n":62,"split":"test","note":"62 receipted calls on the test split, $0.044 at the gateway list price"}],"dataset":"Seeded synthetic radio scripts (8 dev, 30 test) and prose stories (8 dev, 30 test), plus hand-labelled public-domain excerpts (Poe as dev; Doyle and Jacobs adaptations as test); 4 rendered sample episodes for loudness and listening checks.","held_out":true,"caveats":["Mostly synthetic data from a generator written by the same author as the prompts; three prompt changes were made after looking at dev.","The two test excerpts were written from memory, so they are adaptations, not exact texts.","The eval's author cannot hear: acting, how effects land and music fit were not checked by ear. A human listen is needed before publishing.","Results vary run to run on the shared gateway even at temperature 0 (the \"no sound\" row)."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/audio-drama-studio"},"stack":{"summary":"For podcasters, indie authors and teachers. Paste a radio script or a prose chapter; the studio shows you every speaker, line, direction and sound cue to fix before anything is voiced. Each role gets a consented house voice, and the consent ledger checks every line right before it is spoken. Music comes from cleared cues with licence certificates, and sound from CC0 or public-domain recordings or code. The episode is mastered to a podcast or ACX-style spec, with chapters, captions, show notes with an AI disclosure, sides for human actors, a C2PA credential and a signed record.","tagline":"Your script, produced tonight, with every voice consented and every track cleared.","deployment":"hosted-or-self-host","regulatory_note":"Not legal advice. Checked 26 Sep 2026: ACX's Audio Submission Requirements (help.acx.com, page dated 15 Apr 2026) prohibit unauthorised text-to-speech or AI narration, so this studio's audiobook mode meets ACX's technical spec (RMS -23 to -18 dB, peak -3 dB, noise floor -60 dB, 1-5 s room tone, credits files) but does not make a title eligible for ACX. Apple Podcasts recommends -16 LKFS +/- 1 dB and true peak at most -1 dBFS (podcasters.apple.com/support/893-audio-requirements). EU AI Act Art. 50(2) (Regulation (EU) 2024/1689; applies from 2 Aug 2026) requires machine-readable marking of synthetic audio: the C2PA credential covers it, signed with a development certificate, so public validators show the issuer as untrusted. Voice-replica laws (Tennessee ELVIS Act, California AB 2602, New York S7676B, as summarised by the consent ledger, tool 47) are why every line goes through the ledger; this build uses stock synthetic voices only and does not clone. Platform AI-disclosure rules for podcasts (Apple, Spotify) were not verified.","components":[{"id":"llm","role":"Parse: voice hints, aliases and the sound and music cue mapping for radio scripts (one call); speaker attribution and sound suggestions for prose (one call per 36 quotations)","name":"Qwen3.8-27B (NVIDIA NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP) + FP8 (attention/GDN), FP8 KV cache","vram_gb":57,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, temperature 0, thinking off","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["standard","best"],"alternative_to":null},{"id":"tts","role":"Speaks each line with a stock voicepack after the consent ledger allows it; one process per episode on CPU","name":"Kokoro-82M","hf_repo":"hexgrad/Kokoro-82M","license":"Apache-2.0","params":"82M","quant":"FP32","vram_gb":0,"memory_gb_estimate":null,"engine":"kokoro 0.9.4 (misaki G2P) in its own CPU venv, 8 torch threads","receipt_coverage":"none","in_hosted_demo":true,"tiers":["all"],"alternative_to":null},{"id":"music","role":"Score: theme and sting cues rendered through the music-gen-cleared path (tool 37) with its prompt guard, similarity check and signed licence certificate; the hosted demo uses the library, composing new cues needs the studio GPU","name":"ACE-Step 1.5 turbo + 5Hz LM 1.7B","hf_repo":"ACE-Step/Ace-Step1.5","license":"MIT","params":"2.39B DiT + 1.85B LM","quant":"BF16","vram_gb":14.6,"memory_gb_estimate":null,"engine":"ACE-Step 1.5 Python inference (8 steps, ODE), through the studio queue","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["all"],"alternative_to":null},{"id":"mixer","role":"Timeline, sound library, ducking, compression, limiter, loudness to spec, chapters, captions, sides, C2PA and the signed record (CPU)","name":"decosa-api drama module (decosa_api/verticals/drama) + FFmpeg","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12 + numpy; FFmpeg 6.1 (acompressor, alimiter, ebur128, libmp3lame); c2pa-python for the credential","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["all"],"alternative_to":null},{"id":"verifier","role":"Consent ledger's speaker check (tool 47): does each role's rendered voice match the voice enrolled in its entry?","name":"ECAPA-TDNN speaker embeddings (ONNX export)","hf_repo":"speechbrain/spkrec-ecapa-voxceleb","license":"Apache-2.0","params":null,"quant":"FP32","vram_gb":0,"memory_gb_estimate":null,"engine":"ONNX Runtime on CPU (the consent extra); optional on self-host","receipt_coverage":"none","in_hosted_demo":true,"tiers":["all"],"alternative_to":null},{"id":"expressive-tts","role":"Alternate: a permissively licensed voice model with emotion control, driven by stock (not cloned) reference voices","name":"Chatterbox (exaggeration control)","hf_repo":"ResembleAI/chatterbox","license":"MIT","params":"0.5B","quant":null,"vram_gb":4,"memory_gb_estimate":null,"engine":"chatterbox-tts 0.1.4 (as in consented dubbing); not wired into this studio","receipt_coverage":"none","in_hosted_demo":false,"tiers":["alternates"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · CPU only, radio scripts","summary":"No language model: speaker lines and cues are read by code, cues mapped by keywords; prose chapters are not supported.","components":["tts","mixer"],"hardware":"Any 8-core CPU","quality_evidence":[{"metric":"Speakers right on held-out radio scripts (code alone)","value":"100% (389/389)","source":"decosa-api docs/evals/audio-drama-studio.md, 2026-09-26"},{"metric":"Cue mapped to the right library sound by keywords","value":"92.8% (90/97)","source":"decosa-api docs/evals/audio-drama-studio.md, 2026-09-26"}],"latency_note":"measured: renders as standard (seconds per finished minute); parse in milliseconds","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"No model calls, so no receipts; consent decisions, the render receipt and the record are still signed.","hosting":null},{"id":"standard","label":"Standard · the hosted demo","summary":"Qwen3.8-27B through our gateway for the parse, Kokoro on CPU, library music, full mix and credentials.","components":["llm","tts","music","mixer","verifier"],"hardware":"1x 96 GB card (or 32 GB for Qwen alone) plus an 8-core CPU","quality_evidence":[{"metric":"Prose speaker attribution (held-out)","value":"96.8% planted (209/216); 100% public domain (30/30)","source":"decosa-api docs/evals/audio-drama-studio.md, 2026-09-26"},{"metric":"Cue mapping (held-out)","value":"99.0% (96/97); music cues 26/26","source":"decosa-api docs/evals/audio-drama-studio.md, 2026-09-26"},{"metric":"Loudness spec met on delivered files","value":"all sample episodes and chapter files","source":"measured on our server 2026-09-26"},{"metric":"Word error rate heard back by ASR","value":"0.6-5.0% on 4 sample episodes","source":"decosa-api docs/evals/audio-drama-studio.md, 2026-09-26"}],"latency_note":"measured: parse in a few seconds, render in seconds per finished minute","in_hosted_demo":true,"receipt_coverage":"partial","receipt_note":null,"hosting":null},{"id":"best","label":"Best · compose new music per episode","summary":"Standard plus a theme and a sting composed for the episode from your brief, through the music-gen-cleared guard, similarity check and certificate.","components":["llm","tts","music","mixer","verifier"],"hardware":"Standard plus about 15 GB free on a GPU for ACE-Step","quality_evidence":[{"metric":"Music prompt guard on held-out prompts","value":"precision 100%, recall 98%","source":"decosa-api docs/evals/music-gen-cleared.md, 2026-09-25"},{"metric":"Composed cues in this studio","value":"not measured yet","source":null}],"latency_note":"measured for the library cues: under a minute per ACE-Step render on a shared GPU, one at a time","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":null,"hosting":null}],"alternates":[{"id":"expressive-voices","label":"Voices that can act","components":["expressive-tts"],"hardware":"About 4 GB of GPU beside the standard tier","use":"An MIT voice model with emotion control (Chatterbox, exaggeration setting) on stock reference voices.","status":"not served"}],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag> (publishing soon) plus the voice layer docker/drama/Dockerfile","purpose":"Parse, consent gate, Kokoro worker, mix, master, credential, record. Binds 127.0.0.1."},{"name":"decosa-llm","port":8114,"image":"${DECOSA_REGISTRY}/decosa-llm:0.1.0","purpose":"Qwen3.8-27B OpenAI endpoint (hosted: behind our gateway)."}],"tools":[{"name":"FFmpeg","url":"https://ffmpeg.org/","license":"LGPL-2.1+ (GPL builds vary)","purpose":"Decoding, compression, limiting, EBU R128 measurement, MP3 encoding with ID3 chapters."},{"name":"c2pa-python","url":"https://github.com/contentauth/c2pa-python","license":"MIT OR Apache-2.0","purpose":"Embeds the C2PA credential in each MP3."},{"name":"Kenney audio packs (RPG Audio, Impact Sounds)","url":"https://kenney.nl/assets/rpg-audio","license":"CC0-1.0","purpose":"Door, footstep, paper, blade and impact recordings."},{"name":"OpenGameArt and Wikimedia Commons recordings","url":"https://commons.wikimedia.org/","license":"CC0-1.0 or public domain (each file's page checked 26 Sep 2026)","purpose":"Rain, wind, fire, crowd, bells, clock, dog, horse, gunshot, telephone and more; the manifest keeps each source URL."},{"name":"Decosa consent ledger (tool 47)","url":"https://decosa.ai/apps/consent-ledger","license":"part of decosa-api","purpose":"A signed decision for every role and every line."},{"name":"Decosa music-gen-cleared (tool 37)","url":"https://decosa.ai/apps/music-gen-cleared","license":"part of decosa-api","purpose":"Prompt guard, similarity check and licence certificate for each music cue."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB + CPU","fits":true,"notes":"Hosted layout on our server: Qwen3.8-27B on GPU1 behind the gateway; voices, sound and mixing on CPU. Composing new music uses GPU0 through the studio queue."},{"tier":"CPU only (radio scripts)","fits":true,"notes":"Scripts in NAME: line format parse with code alone (speakers exact; cue mapping 92.8% by keywords); prose needs the model or a remote endpoint."}],"latency":[{"lane":"Parse (one model call for a script, one per 36 quotations for prose)","typical_ms":2700,"source":"measured on our server 2026-09-26: median of 62 held-out parses, gateway route; p90 4.8 s"},{"lane":"Render a 2.4-minute episode (20 lines, music, 9 cues)","typical_ms":19700,"source":"measured on our server 2026-09-26: Night Shift sample, 8 CPU threads"},{"lane":"Render a 3.6-minute episode (26 lines, music)","typical_ms":31000,"source":"measured on our server 2026-09-26: Holmes sample"},{"lane":"Consent refusal","typical_ms":10,"source":"measured on our server 2026-09-26: the 422 comes back before anything is voiced"}],"benchmark":null,"notes":["The words spoken are always the script's own: code splits the text, and the model only says who speaks and which sound a cue means.","Deliveries in the script (quietly, shouting) change pace and level only; (on phone) and (off) change the sound.","Files are kept 7 days on the hosted service; the signed record keeps hashes and ids, never script text."]},"buyer_facts":[{"label":"Consent","value":"Every role and every line is checked against the consent ledger before it is spoken, and each decision is signed; a revoked or out-of-scope voice stops the render. Only stock voices with entries can be cast; nothing is cloned."},{"label":"What you check","value":"The parse is shown before anything is voiced: speakers, lines, sound and music cues, all editable. Spoken words are always the script's own."},{"label":"Loudness","value":"Measured on the delivered file: podcast -16 LUFS +/- 1 and true peak at most -1 dBTP, or ACX RMS, peak, noise floor and room tone per chapter file. All sample episodes passed."},{"label":"Licences","value":"Voices Apache-2.0 (Kokoro-82M), music MIT (ACE-Step) with a signed certificate per cue, sounds CC0 or public domain with source links, or generated by code. The show notes list them."},{"label":"Cost per episode","value":"A fraction of a cent of model time for the parse (one receipted call for a script); voices and mixing run on CPU, a few seconds per finished minute (measured)."},{"label":"Data retention","value":"Episode files are deleted after 7 days; until then they open only through signed links that expire within hours, given to the token or key that made the episode. The server keeps ids, hashes, settings and measurements, and the signed record; never script text in logs."},{"label":"What leaves the box","value":"Hosted: the script, sent to Decosa's API; its parse calls go through the Decosa API. Self-host: nothing."},{"label":"Not for ACX","value":"ACX prohibits unauthorised AI narration (requirements dated 15 Apr 2026). The audiobook mode meets its technical spec only; use it for other platforms, classrooms and pitches."}],"hosted_now":{"needs":["music-generation","qwen3.8-27b"],"off":["music-generation"],"live_by_default":false,"live_status":"https://api.decosa.ai/status"},"data_handling":{"page":"/data#audio-drama-studio","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Episode files are deleted after 7 days; until then they open only through signed links that expire within hours, given to the token or key that made the episode. The server keeps ids, hashes, settings and measurements, and the signed record; never script text in logs.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/apps/audio-drama-studio","input":"drama","lanes":[{"id":"parse","title":"Parse to review","kind":"list"},{"id":"cast","title":"Cast and consent","kind":"list"},{"id":"episode","title":"Episode","kind":"list"},{"id":"deliverables","title":"Deliverables","kind":"json"}],"samples":[{"n":1,"id":"scandal-in-bohemia","title":"Scandal in bohemia","deep_link":"/apps/audio-drama-studio?sample=1&autorun=0"},{"n":2,"id":"cask-of-amontillado","title":"Cask of amontillado","deep_link":"/apps/audio-drama-studio?sample=2&autorun=0"},{"n":3,"id":"night-shift","title":"Night shift","deep_link":"/apps/audio-drama-studio?sample=3&autorun=0"},{"n":4,"id":"town-mouse","title":"Town mouse","deep_link":"/apps/audio-drama-studio?sample=4&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/audio-drama-studio-hosted.md","selfhost":"/prompts/audio-drama-studio-selfhost.md","assemble":"/prompts/audio-drama-studio-assemble.md","mac":null},"rehearsal":{"bundle":"/samples/audio-drama-studio.zip","bundle_url":"https://decosa.ai/samples/audio-drama-studio.zip","folder":"/samples/audio-drama-studio/","expected":"/samples/audio-drama-studio/expected.json","files":["/samples/audio-drama-studio/expected.json","/samples/audio-drama-studio/inputs/night-shift.txt"],"bytes":3134,"checks":["the script is read as a radio play","four roles: narrator, Rosa, Dispatch and Teo","twenty spoken lines","every sound cue maps to a library sound or a stop","the phone voice is marked as a phone effect","a fictional performer outside their consented project is refused","the suggested house cast is allowed for every role","the episode renders","every spoken line passed the consent gate (one signed decision per line)","integrated loudness within -16 LUFS +/- 1 dB","true peak at most -1 dBTP","the MP3 carries chapter markers","the file carries a C2PA credential with the consent link","captions, transcript, chapters, show notes and the sides pack are delivered","the signed production record verifies","every model call is receipted and signed"],"licence":"Original script written for Decosa (2026), free to reuse. Voices are stock Kokoro-82M voicepacks (Apache-2.0); music was rendered with ACE-Step 1.5 (MIT) through the music-gen-cleared path; sound effects are CC0 / public-domain recordings or generated by code. Part of decosa-api, which will be released under AGPL-3.0-or-later; until then the source is on request.","about":"An original short radio drama (Night Shift at Weather Station Nine, written for Decosa) is parsed into roles, lines and cues; one role is cast with a fictional performer from the consent-ledger demo whose consent covers a different project, which must be refused; the suggested house cast must pass; the episode renders to the podcast spec with library music, and the file must meet the loudness spec, carry a C2PA credential that links every role to its consent entry, and come with a signed record that verifies.","run":{"containers":"docker compose exec api python scripts/rehearse.py audio-drama-studio","checkout":"python scripts/rehearse.py audio-drama-studio --bundle audio-drama-studio.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py audio-drama-studio"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=audio-drama-studio","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":0,"basis":null,"unknown":[]},{"id":"standard","gpu_gb":72.2,"basis":"stack","unknown":[]},{"id":"best","gpu_gb":72.2,"basis":"stack","unknown":[]},{"id":"alternate-expressive-voices","gpu_gb":4,"basis":"stack","unknown":[]}],"mac":null},"links":{"page":"/apps/audio-drama-studio","json":"/use-cases/audio-drama-studio.json","metrics":"/metrics/audio-drama-studio","console":"/apps/audio-drama-studio","console_sample":"/apps/audio-drama-studio?sample=1&autorun=0","stack":"/apps/audio-drama-studio#stack","try_live":"/apps/audio-drama-studio#live","watch":"/apps/audio-drama-studio#watch","build":"/apps/audio-drama-studio#build","self_host":"/apps/audio-drama-studio#self-host","prompts":{"hosted":"/prompts/audio-drama-studio-hosted.md","selfhost":"/prompts/audio-drama-studio-selfhost.md","assemble":"/prompts/audio-drama-studio-assemble.md","mac":null}}}