Produce an audio drama
Paste a radio script or a prose chapter and review the parse: every speaker, line, direction and sound cue, before anything is voiced. Cast each role from consented house voices; every line passes the consent ledger right before it is spoken. Music comes from cleared cues with licence certificates, and sound from CC0 or public-domain recordings or code. The mix is mastered to a podcast or ACX-style audiobook spec, with chapters, captions, show notes with an AI disclosure, a C2PA credential, a signed record and sides for human actors.
- For
- Teams in film, tv and games and creative and media.
- Time per task28 stypical (median) on the sample
- Cost per task~$0.10 per 100 episodesmeasured, at list price
- Accuracy100% (389/389)Radio scripts: speaker right (code alone)All results and caveats
Review, cast, render
LiveParse the script. You will see every line, speaker and cue before anything is voiced, and can fix them.
Watch: scripts parsed, cast, gated and produced
Replay · not liveRecorded from real runs on the decosa-api pre-release server on 26 Sep 2026: Qwen3.8-27B through our gateway for the parse, Kokoro-82M on CPU for the voices, library music from music-gen-cleared. Public-domain and original scripts.
Watch the parse (four roles, eight sound cues, two music cues), then the render: 26 lines, each checked against the consent ledger, rain and fire beds, a knock, a sting and an off-mic visitor, mastered to -16 LUFS with a C2PA credential.
Get an API key
- Call the audio drama and narrated story studio API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on Qwen3.8-27B (1× RTX 5090 32 GB or larger) for the parse; the voices, sound and mix run on CPU; ACE-Step only if you compose new music (a GPU with about 10 GB free).
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- audio-drama-studio
Use the hosted API
# Decosa audio drama studio: use the hosted API
You are wiring Decosa's audio drama studio into this project. It takes a script (a radio play with `NAME: line`,
`[SFX: ...]` and `[MUSIC: ...]` cues, or a prose chapter) and returns a produced episode: every line voiced by a
consented stock voice, openly licensed or generated sound, cleared music, mastered to a podcast or ACX-style audiobook
spec, with chapters, captions, a transcript, show notes with an AI disclosure, sides for human actors, a C2PA credential
and a signed record. Use only what is listed below. If you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`. Health: `GET https://api.decosa.ai/healthz`.
- The hosted service speaks only stock voices with consent-ledger entries (`GET /drama/info` lists them). It never
clones a voice. Send unpublished work only if I say the hosted service is fine for it; otherwise use self-host.
- Always show the parse to a person before rendering: the model attributes quotations and maps cues, and it can be wrong.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page, in `DECOSA_API_KEY`, sent as
`Authorization: Bearer $DECOSA_API_KEY`. A key renders 20 episodes per UTC day.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "audio-drama-studio"}` returns `{"token", ...}`;
3 episodes per session. Over a limit you get HTTP 429 with `Retry-After`.
## Endpoints
- `POST /drama/parse` (token): `{"script": "...", "format"?: "auto"|"screenplay"|"prose"}` (30,000 characters at most) →
`{parse_id, parse: {roles, blocks, warnings, stats}, cast, receipts, usage}`. Blocks are `line` (role, text, delivery,
effect), `sfx` (sound tag, layer spot|bed|stop), `music` (cue theme|sting|bed|stop), `direction`, `scene`, `chapter`,
`pause`. Quotations the model could not attribute have role `unassigned`; fix them before rendering.
- `POST /drama/cast` (token): `{"parse", "cast"?: {role: identity_id}, "spec"?}` → one signed consent decision per role.
- `POST /drama/episodes` (token): `{"parse", "parse_id", "cast", "spec": "podcast"|"acx", "music": {"mode": "library",
"style": "victorian-mystery"} | {"mode": "none"}, "title"?, "author"?, "source_credit"?}` → 202 `{episode_id}`, or 422
`{category: "consent_refused", decisions}` when a voice's consent does not cover this use (nothing is rendered).
- `GET /drama/episodes/{id}` (token): poll every 3 s until `status` is `done`, `refused` or `failed`. `done` has `files`
(MP3 with C2PA and chapters; one per chapter for acx), `measured` loudness and `checks`, `documents` (captions.vtt,
transcript.json, chapters.json, transcript.txt, show-notes.md, sides.zip, record.json), `cast` with the consent entry
per role, `music` with licence certificates, `sounds` with licences, and `disclosure`.
- `GET /drama/records/{id}` (no token): the signed record; verify it at `POST /record/verify` with `{"record": ...}`.
- No token: `GET /drama/info`, `/drama/samples`, `/drama/voices/{voice}.mp3`, `/drama/music/{cue}.mp3`,
`/drama/music/{cue}/certificate`.
## Example (Python, `pip install httpx`)
```python
import httpx, os, time
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
script = open("script.txt").read()
p = httpx.post(f"{API}/drama/parse", headers=H, timeout=300, json={"script": script}).json()
# show p["parse"] to a person and let them fix speakers and cues here
r = httpx.post(f"{API}/drama/episodes", headers=H, timeout=60, json={"parse": p["parse"], "parse_id": p["parse_id"],
"cast": p["cast"], "spec": "podcast", "music": {"mode": "library", "style": "victorian-mystery"}})
if r.status_code == 422:
raise SystemExit([(d["role"], d["code"]) for d in r.json()["decisions"] if not d["allowed"]])
ep_id = r.json()["episode_id"]
while (ep := httpx.get(f"{API}/drama/episodes/{ep_id}", headers=H, timeout=60).json())["status"] in ("queued", "running"):
time.sleep(3)
assert ep["status"] == "done", ep.get("error")
f = ep["files"][0]
print(f["measured"], f["checks"], ep["disclosure"])
open("episode.mp3", "wb").write(httpx.get(API + f["url"]).content)
```
## Honest limits
- Kokoro-82M voices are clear but flat; delivery notes change pace and level only. Treat the result as a temp track.
- ACX does not accept AI narration; the acx mode meets its technical spec only.
- Files stay on the server for 7 days; download what you need.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa audio drama studio: run it yourself (containers)
You are setting up Decosa's audio drama studio on this machine: Qwen3.8-27B parses scripts, Kokoro-82M (Apache-2.0)
speaks each line on CPU with consented stock voices, and decosa-api mixes, masters, credentials and signs the episode.
Scripts and episodes stay here; nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Don't substitute other images or models.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/audio-drama-studio.zip (3 KB, 16 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py audio-drama-studio` (the api image carries the same bundle under /app/rehearsal/audio-drama-studio/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py audio-drama-studio --bundle audio-drama-studio.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the script is read as a radio play", "four roles: narrator, Rosa, Dispatch and Teo", "twenty spoken lines"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Rules first
- Every voice render goes through the consent ledger. Never bypass it. Enrol a real person only with their recorded
consent and a specific description of the use; this build does not clone voices.
- Keep the parse review step: a person checks speakers and cues before rendering.
## Steps
1. Docker (and, for the parse model, the NVIDIA container toolkit and a GPU with 32 GB or more): if
`docker compose version` or `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi` fails,
install them from the official instructions.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the
`llm` and `api` services. The api image needs the voice layer: `docker/drama/Dockerfile` in decosa-api (Kokoro in
its own CPU venv, model and voicepacks downloaded at build time). The tool's assemble prompt has the exact
commands. Use named volumes and bind every port to 127.0.0.1.
3. `docker compose run --rm api python scripts/provenance_devcert.py`, then `docker compose up -d`.
4. Smoke test: `docker compose exec api python scripts/rehearse.py audio-drama-studio --base-url http://127.0.0.1:8445`
must print `16/16 checks passed`.
5. Report back: the render time per finished minute and the loudness numbers from the rehearsal episode.
Off. Voice renders run on CPU and are not a network service; only the text model could be offered, and only if I say yes.
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMlite tierRuns with a smaller tier
The standard tier does not fit: Qwen3.8-27B (NVIDIA NVFP4) needs a GPU. The lite tier fits.
- GeForce RTX 4090lite tierRuns with a smaller tier
The standard tier does not fit: Needs about 34.6 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits.
- GeForce RTX 5090lite tierRuns with a smaller tier
The standard tier does not fit: Needs about 42.6 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits.
- 2x GeForce RTX 5090best tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. The best tier fits too.
- L40Slite tierRuns with a smaller tier
The standard tier does not fit: Needs about 48.2 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits.
- H100 80 GB (SXM)best tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too.
- RTX PRO 6000 Blackwell 96 GBbest tierRuns
The standard tier fits (72.2 of 96 GB). The best tier fits too.
- 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns
The standard tier fits (72.2 of 192 GB). The best tier fits too.
- Apple M3 Ultra (Mac Studio), 96 GBlite tierRuns with a smaller tier
The standard tier can't be checked: ACE-Step 1.5 turbo + 5Hz LM 1.7B has no mapped Apple Silicon build The lite tier fits.
- Apple M5 Max, 64 GBlite tierRuns with a smaller tier
The standard tier can't be checked: ACE-Step 1.5 turbo + 5Hz LM 1.7B has no mapped Apple Silicon build The lite tier fits.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"audio-drama-studio"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py audio-drama-studio
Download the mock-data bundle (3 KB, 16 checks)expected.json
An original short radio drama (Night Shift at Weather Station Nine, written for Decosa) is parsed into roles, lines and cues; one role is cast with a fictional performer from the consent-ledger demo whose consent covers a different project, which must be refused; the suggested house cast must pass; the episode renders to the podcast spec with library music, and the file must meet the loudness spec, carry a C2PA credential that links every role to its consent entry, and come with a signed record that verifies.
What the rehearsal checks
- the script is read as a radio play
- four roles: narrator, Rosa, Dispatch and Teo
- twenty spoken lines
- every sound cue maps to a library sound or a stop
- the phone voice is marked as a phone effect
- a fictional performer outside their consented project is refused
- the suggested house cast is allowed for every role
- the episode renders
- every spoken line passed the consent gate (one signed decision per line)
- integrated loudness within -16 LUFS +/- 1 dB
- true peak at most -1 dBTP
- the MP3 carries chapter markers
- the file carries a C2PA credential with the consent link
- captions, transcript, chapters, show notes and the sides pack are delivered
- the signed production record verifies
- every model call is receipted and signed
Licence: Original script written for Decosa (2026), free to reuse. Voices are stock Kokoro-82M voicepacks (Apache-2.0); music was rendered with ACE-Step 1.5 (MIT) through the music-gen-cleared path; sound effects are CC0 / public-domain recordings or generated by code. Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa audio drama studio: run it yourself (containers)
You are setting up Decosa's audio drama studio on this machine: Qwen3.8-27B parses scripts, Kokoro-82M (Apache-2.0)
speaks each line on CPU with consented stock voices, and decosa-api mixes, masters, credentials and signs the episode.
Scripts and episodes stay here; nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Don't substitute other images or models.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/audio-drama-studio.zip (3 KB, 16 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py audio-drama-studio` (the api image carries the same bundle under /app/rehearsal/audio-drama-studio/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py audio-drama-studio --bundle audio-drama-studio.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the script is read as a radio play", "four roles: narrator, Rosa, Dispatch and Teo", "twenty spoken lines"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Rules first
- Every voice render goes through the consent ledger. Never bypass it. Enrol a real person only with their recorded
consent and a specific description of the use; this build does not clone voices.
- Keep the parse review step: a person checks speakers and cues before rendering.
## Steps
1. Docker (and, for the parse model, the NVIDIA container toolkit and a GPU with 32 GB or more): if
`docker compose version` or `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi` fails,
install them from the official instructions.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the
`llm` and `api` services. The api image needs the voice layer: `docker/drama/Dockerfile` in decosa-api (Kokoro in
its own CPU venv, model and voicepacks downloaded at build time). The tool's assemble prompt has the exact
commands. Use named volumes and bind every port to 127.0.0.1.
3. `docker compose run --rm api python scripts/provenance_devcert.py`, then `docker compose up -d`.
4. Smoke test: `docker compose exec api python scripts/rehearse.py audio-drama-studio --base-url http://127.0.0.1:8445`
must print `16/16 checks passed`.
5. Report back: the render time per finished minute and the loudness numbers from the rehearsal episode.
Off. Voice renders run on CPU and are not a network service; only the text model could be offered, and only if I say yes.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Runs with a smaller tierAudio drama and narrated story studio on GeForce RTX 5090: use the Lite · CPU only, radio scripts tier
The standard tier does not fit: Needs about 42.6 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits.
Lite · CPU only, radio scripts: what changes
Nothing: it runs as listed in the stack.
Memory per component
- Speaks each line with a stock voicepack after...: Kokoro-82M. CPU. Runs on CPU (vram_gb 0 in stack.json).
- Timeline, sound library, ducking, compression...: decosa-api drama module (decosa_api/verticals/drama) + FFmpeg. CPU. Runs on CPU (vram_gb 0 in stack.json).
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Audio drama and narrated story studio, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Audio drama and narrated story studio on my hardware Fetch https://decosa.ai/prompts/audio-drama-studio-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=audio-drama-studio) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Lite · CPU only, radio scripts (lite). Fit check: runs, about 0 GB of 32 GB used. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Speaks each line with a stock voicepack after...: Kokoro-82M (hexgrad/Kokoro-82M), CPU - Timeline, sound library, ducking, compression...: decosa-api drama module (decosa_api/verticals/drama) + FFmpeg, CPU During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/audio-drama-studio-assemble.md
Get an API key
- Call the audio drama and narrated story studio API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on Qwen3.8-27B (1× RTX 5090 32 GB or larger) for the parse; the voices, sound and mix run on CPU; ACE-Step only if you compose new music (a GPU with about 10 GB free).
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Your script, produced tonight, with every voice consented and every track cleared.
For podcasters, indie authors and teachers. Paste a radio script or a prose chapter; the studio shows you every speaker, line, direction and sound cue to fix before anything is voiced. Each role gets a consented house voice, and the consent ledger checks every line right before it is spoken. Music comes from cleared cues with licence certificates, and sound from CC0 or public-domain recordings or code. The episode is mastered to a podcast or ACX-style spec, with chapters, captions, show notes with an AI disclosure, sides for human actors, a C2PA credential and a signed record.
- Deployment
- Hosted or self-host
- Regulatory
- Not legal advice. Checked 26 Sep 2026: ACX's Audio Submission Requirements (help.acx.com, page dated 15 Apr 2026) prohibit unauthorised text-to-speech or AI narration, so this studio's audiobook mode meets ACX's technical spec (RMS -23 to -18 dB, peak -3 dB, noise floor -60 dB, 1-5 s room tone, credits files) but does not make a title eligible for ACX. Apple Podcasts recommends -16 LKFS +/- 1 dB and true peak at most -1 dBFS (podcasters.apple.com/support/893-audio-requirements). EU AI Act Art. 50(2) (Regulation (EU) 2024/1689; applies from 2 Aug 2026) requires machine-readable marking of synthetic audio: the C2PA credential covers it, signed with a development certificate, so public validators show the issuer as untrusted. Voice-replica laws (Tennessee ELVIS Act, California AB 2602, New York S7676B, as summarised by the consent ledger, tool 47) are why every line goes through the ledger; this build uses stock synthetic voices only and does not clone. Platform AI-disclosure rules for podcasts (Apple, Spotify) were not verified.
Text description
A script enters decosa-api. Code splits it into lines, cues and quotations; Qwen3.8-27B through our gateway adds speakers and the cue mapping with a signed receipt. You review the parse and pick a voice per role. The consent ledger checks each role and every line right before Kokoro-82M speaks it on CPU. Sound comes from CC0 or public-domain recordings or code; music from cleared ACE-Step cues with licence certificates. The mixer ducks, compresses and masters to the podcast or ACX spec. A C2PA credential links every role's consent entry and a signed record lists every step. Outputs: MP3 with chapters, captions, transcript, show notes with an AI disclosure and sides for human actors.
At a glance
- Consent
- Every role and every line is checked against the consent ledger before it is spoken, and each decision is signed; a revoked or out-of-scope voice stops the render. Only stock voices with entries can be cast; nothing is cloned.
- What you check
- The parse is shown before anything is voiced: speakers, lines, sound and music cues, all editable. Spoken words are always the script's own.
- Loudness
- Measured on the delivered file: podcast -16 LUFS +/- 1 and true peak at most -1 dBTP, or ACX RMS, peak, noise floor and room tone per chapter file. All sample episodes passed.
- Licences
- Voices Apache-2.0 (Kokoro-82M), music MIT (ACE-Step) with a signed certificate per cue, sounds CC0 or public domain with source links, or generated by code. The show notes list them.
- Cost per episode
- A fraction of a cent of model time for the parse (one receipted call for a script); voices and mixing run on CPU, a few seconds per finished minute (measured).
- Data retention
- Episode files are deleted after 7 days; until then they open only through signed links that expire within hours, given to the token or key that made the episode. The server keeps ids, hashes, settings and measurements, and the signed record; never script text in logs.
- What leaves the box
- Hosted: the script, sent to Decosa's API; its parse calls go through the Decosa API. Self-host: nothing.
- Not for ACX
- ACX prohibits unauthorised AI narration (requirements dated 15 Apr 2026). The audiobook mode meets its technical spec only; use it for other platforms, classrooms and pitches.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
CPU only, radio scripts
No language model: speaker lines and cues are read by code, cues mapped by keywords; prose chapters are not supported.
- Models
- Kokoro-82M
- ACE-Step 1.5 turbo + 5Hz LM 1.7B
- decosa-api drama module (decosa_api/verticals/drama) + FFmpeg
- ECAPA-TDNN speaker embeddings (ONNX export)
- Hardware
- Any 8-core CPU
- Quality evidence
- Speakers right on held-out radio scripts (code alone)100% (389/389)decosa-api docs/evals/audio-drama-studio.md, 2026-09-26
- Cue mapped to the right library sound by keywords92.8% (90/97)decosa-api docs/evals/audio-drama-studio.md, 2026-09-26
- Latency
- measured: renders as standard (seconds per finished minute); parse in milliseconds
- Verification
- No proof yetSelf-host onlyNo model calls, so no receipts; consent decisions, the render receipt and the record are still signed.
- In the hosted demo
Standard
the hosted demo
Qwen3.8-27B through our gateway for the parse, Kokoro on CPU, library music, full mix and credentials.
- Models
- Qwen3.8-27B (NVIDIA NVFP4)
- Kokoro-82M
- ACE-Step 1.5 turbo + 5Hz LM 1.7B
- decosa-api drama module (decosa_api/verticals/drama) + FFmpeg
- ECAPA-TDNN speaker embeddings (ONNX export)
- Hardware
- 1x 96 GB card (or 32 GB for Qwen alone) plus an 8-core CPU
- Quality evidence
- Prose speaker attribution (held-out)96.8% planted (209/216); 100% public domain (30/30)decosa-api docs/evals/audio-drama-studio.md, 2026-09-26
- Cue mapping (held-out)99.0% (96/97); music cues 26/26decosa-api docs/evals/audio-drama-studio.md, 2026-09-26
- Loudness spec met on delivered filesall sample episodes and chapter filesmeasured on our server 2026-09-26
- Word error rate heard back by ASR0.6-5.0% on 4 sample episodesdecosa-api docs/evals/audio-drama-studio.md, 2026-09-26
- Latency
- measured: parse in a few seconds, render in seconds per finished minute
- Verification
- Proof: partial
Best
compose new music per episode
Standard plus a theme and a sting composed for the episode from your brief, through the music-gen-cleared guard, similarity check and certificate.
- Models
- Qwen3.8-27B (NVIDIA NVFP4)
- Kokoro-82M
- ACE-Step 1.5 turbo + 5Hz LM 1.7B
- decosa-api drama module (decosa_api/verticals/drama) + FFmpeg
- ECAPA-TDNN speaker embeddings (ONNX export)
- Hardware
- Standard plus about 15 GB free on a GPU for ACE-Step
- Quality evidence
- Music prompt guard on held-out promptsprecision 100%, recall 98%decosa-api docs/evals/music-gen-cleared.md, 2026-09-25
- Composed cues in this studionot measured yet
- Latency
- measured for the library cues: under a minute per ACE-Step render on a shared GPU, one at a time
- Verification
- Proof: partialSelf-host only
Also runs on
- Voices that can actChatterbox (exaggeration control)not servedAn MIT voice model with emotion control (Chatterbox, exaggeration setting) on stock reference voices. Hardware: About 4 GB of GPU beside the standard tier.
We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Parse: voice hints, aliases and the sound and music cue mapping for radio scripts (one call); speaker attribution and sound suggestions for prose (one call per 36 quotations)Qwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 57 GBProof: strongIn the hosted demo | StandardBest | 27.8B · 57 GB | Proof: strongIn the hosted demo | |
| ||||
Speaks each line with a stock voicepack after the consent ledger allows it; one process per episode on CPUKokoro-82Mhexgrad/Kokoro-82M on Hugging Face (opens in a new tab) 82M · 0 GBNo proof yetIn the hosted demo | LiteStandardBest | 82M · 0 GB | No proof yetIn the hosted demo | |
| ||||
Score: theme and sting cues rendered through the music-gen-cleared path (tool 37) with its prompt guard, similarity check and signed licence certificate; the hosted demo uses the library, composing new cues needs the studio GPUACE-Step 1.5 turbo + 5Hz LM 1.7BACE-Step/Ace-Step1.5 on Hugging Face (opens in a new tab) 2.39B DiT + 1.85B LM · 14.6 GBProof: partialIn the hosted demo | LiteStandardBest | 2.39B DiT + 1.85B LM · 14.6 GB | Proof: partialIn the hosted demo | |
| ||||
Timeline, sound library, ducking, compression, limiter, loudness to spec, chapters, captions, sides, C2PA and the signed record (CPU)decosa-api drama module (decosa_api/verticals/drama) + FFmpeg 0 GBProof: partialIn the hosted demo | LiteStandardBest | 0 GB | Proof: partialIn the hosted demo | |
| ||||
Consent ledger's speaker check (tool 47): does each role's rendered voice match the voice enrolled in its entry?ECAPA-TDNN speaker embeddings (ONNX export)speechbrain/spkrec-ecapa-voxceleb on Hugging Face (opens in a new tab) 0 GBNo proof yetIn the hosted demo | LiteStandardBest | 0 GB | No proof yetIn the hosted demo | |
| ||||
Alternate: a permissively licensed voice model with emotion control, driven by stock (not cloned) reference voicesChatterbox (exaggeration control)ResembleAI/chatterbox on Hugging Face (opens in a new tab) 0.5B · 4 GBNo proof yetSelf-host only | Alternate | 0.5B · 4 GB | No proof yetSelf-host only | |
| ||||
Tools, services and hardware
Tools
- FFmpeg (opens in a new tab)LGPL-2.1+ (GPL builds vary)
Decoding, compression, limiting, EBU R128 measurement, MP3 encoding with ID3 chapters.
- c2pa-python (opens in a new tab)MIT OR Apache-2.0
Embeds the C2PA credential in each MP3.
Door, footstep, paper, blade and impact recordings.
- OpenGameArt and Wikimedia Commons recordings (opens in a new tab)CC0-1.0 or public domain (each file's page checked 26 Sep 2026)
Rain, wind, fire, crowd, bells, clock, dog, horse, gunshot, telephone and more; the manifest keeps each source URL.
- Decosa consent ledger (tool 47) (opens in a new tab)part of decosa-api
A signed decision for every role and every line.
- Decosa music-gen-cleared (tool 37) (opens in a new tab)part of decosa-api
Prompt guard, similarity check and licence certificate for each music cue.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag> (publishing soon) plus the voice layer docker/drama/DockerfileParse, consent gate, Kokoro worker, mix, master, credential, record. Binds 127.0.0.1.
- decosa-llm:8114
${DECOSA_REGISTRY}/decosa-llm:0.1.0Qwen3.8-27B OpenAI endpoint (hosted: behind our gateway).
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB + CPU Fits
Hosted layout on our server: Qwen3.8-27B on GPU1 behind the gateway; voices, sound and mixing on CPU. Composing new music uses GPU0 through the studio queue.
- CPU only (radio scripts) Fits
Scripts in NAME: line format parse with code alone (speakers exact; cue mapping 92.8% by keywords); prose needs the model or a remote endpoint.
Latency per lane
- Parse (one model call for a script, one per 36 quotations for prose)2.7 s
Measuredmeasured on our server 2026-09-26: median of 62 held-out parses, gateway route; p90 4.8 s
- Render a 2.4-minute episode (20 lines, music, 9 cues)19.7 s
Measuredmeasured on our server 2026-09-26: Night Shift sample, 8 CPU threads
- Render a 3.6-minute episode (26 lines, music)31.0 s
Measuredmeasured on our server 2026-09-26: Holmes sample
- Consent refusal10 ms
Measuredmeasured on our server 2026-09-26: the 422 comes back before anything is voiced
Notes
- The words spoken are always the script's own: code splits the text, and the model only says who speaks and which sound a cue means.
- Deliveries in the script (quietly, shouting) change pace and level only; (on phone) and (off) change the sound.
- Files are kept 7 days on the hosted service; the signed record keeps hashes and ids, never script text.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa audio drama studio on this machine
You are setting up a studio that turns a script (a radio play or a prose chapter) into a produced episode: it parses the
script into roles, lines and cues for me to check, voices each line with a consented stock voice, adds openly licensed
sound and cleared music, masters to a podcast or ACX-style audiobook spec, and writes chapters, captions, show notes,
sides for human actors, a C2PA credential and a signed record. Work step by step, show me each command before you run
anything with `sudo`, and stop to ask if a check fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/audio-drama-studio.zip (3 KB, 16 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py audio-drama-studio` (the api image carries the same bundle under /app/rehearsal/audio-drama-studio/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py audio-drama-studio --bundle audio-drama-studio.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the script is read as a radio play", "four roles: narrator, Rosa, Dispatch and Teo", "twenty spoken lines"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- decosa-api (AGPL-3.0-or-later) runs the studio. Voices: Kokoro-82M (`hexgrad/Kokoro-82M`, Apache-2.0), stock voicepacks only,
on CPU; nothing is cloned. Parse model: Qwen3.8-27B (Apache-2.0) on one GPU, or any OpenAI-compatible endpoint I give
you. Library music was rendered with ACE-Step 1.5 (MIT) and ships with its licence certificates; sound effects are CC0
or public-domain recordings (sources in `decosa_api/verticals/drama/data/sfx/manifest.json`) or generated by code.
- Unpublished scripts are confidential: bind every port to 127.0.0.1 and keep the data volume private.
- Be honest about the output: these are temp tracks with synthetic voices. ACX does not accept AI narration (its
submission requirements, dated 15 Apr 2026, prohibit unauthorised TTS); the ACX mode checks the technical spec only.
## 1. Check the machine
1. `docker --version` and `docker compose version`. If Docker is missing, install it from Docker's official repository
after asking me.
2. Disk: about 7 GB free (the api image, the voice layer of about 4.5 GB, and room for episodes).
3. For the parse model: `nvidia-smi` shows a GPU with at least 32 GB and the NVIDIA container toolkit works
(`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`). Without a GPU, ask me for an
OpenAI-compatible endpoint; radio-play scripts still parse with code alone (prose needs the model).
## 2. The images
`${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
```bash
git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> # access required; a release with decosa_api/verticals/drama/
cd decosa-api
docker build -f docker/api/Dockerfile -t decosa-api:local .
docker build -f docker/drama/Dockerfile --build-arg BASE=decosa-api:local -t decosa-api:drama .
```
The second image adds Kokoro in its own CPU venv and downloads the model and the 16 voicepacks at build time, so the
container runs offline. In our run the two builds took 24 s and 100 s with a warm cache.
## 3. docker-compose.yml
Write this in `~/decosa/drama/`:
```yaml
services:
api:
image: decosa-api:drama
ports: ["127.0.0.1:8445:8445"]
environment:
DECOSA_HOST: 0.0.0.0
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_PROVENANCE_DIR: /provenance
DECOSA_PUBLIC_BASE_URL: http://127.0.0.1:8445 # written into the C2PA consent links; use your real host
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_DRAMA_TTS_THREADS: "8"
volumes: ["decosa-data:/data", "provenance:/provenance"]
depends_on: { llm: { condition: service_healthy } }
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/drama/info', timeout=4)"], interval: 30s, retries: 10 }
llm:
image: vllm/vllm-openai:v0.29.0
command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768"]
ports: ["127.0.0.1:8114:8000"]
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
volumes: ["hf-cache:/root/.cache/huggingface"]
healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 40 }
volumes:
decosa-data:
provenance:
hf-cache:
```
If a model server already runs on this machine, drop the `llm` service, use `network_mode: host` for `api` (remove
`ports`) and point `DECOSA_LLM_URL` at it with its served model name. The image runs as uid 10001: use named volumes.
## 4. Start
1. C2PA development certificate (once): `docker compose run --rm api python scripts/provenance_devcert.py`.
2. `docker compose up -d`, then `curl -s localhost:8445/drama/info | jq '{c2pa: .provenance.c2pa, ledger: .cast.consent_ledger, styles: [.music.styles[].id]}'`:
`c2pa: true`, `ledger: true`, and the music styles.
3. `curl -s localhost:8445/drama/info | jq '[.cast.voices[] | select(.state=="active")] | length'` is 16 (12 house
voices and 4 fictional demo performers, each with an active consent-ledger entry).
4. Optional voice check (does each role's output match its enrolled voice?): export the ECAPA speaker model as in the
consent-ledger assemble prompt and set `DECOSA_CONSENT_SPK_MODEL`. Without it the episode says "voice check not run".
## 5. Smoke test
1. Rehearse on the bundled mock script first: `docker compose exec api python scripts/rehearse.py audio-drama-studio --base-url http://127.0.0.1:8445`.
Expect `16/16 checks passed` (ours: 20.7 s, including an 18 s render on 8 CPU threads).
2. Or by hand: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"audio-drama-studio"}' | jq -r .token)`;
`POST /drama/parse {"script": "..."}`; check the parse; `POST /drama/episodes {"parse", "parse_id", "cast", "spec": "podcast", "music": {"mode": "library", "style": "victorian-mystery"}}`;
poll `GET /drama/episodes/{id}` until `done`; check `files[0].checks` (loudness), `files[0].c2pa.consent_link`,
and `POST /record/verify` with `GET /drama/records/{id}`.
3. Report back: the render time per finished minute, the loudness numbers and the signing key id (`GET /attest/signing-key`).
## 6. Use it from my project
Contract: `API_CONTRACT.md`, section "Audio drama and narrated story studio". Keep the parse review step in the UI: the
model attributes quotations and maps cues, and a person should check it before anything is voiced. To cast your own
performers, enrol them in the consent ledger (`POST /consent/entries`) with the uses and project; today they still speak
with a stock voice, because this build does not clone.
## 7. Optional: compose new music
Composing new cues per episode needs the studio render worker and ACE-Step on a GPU (see the music-gen-cleared assemble
prompt); without it, use the library styles.What it does, in shortWho it's for, where it runs and the key results
Paste a radio script or a prose chapter and review the parse: every speaker, line, direction and sound cue, before anything is voiced. Cast each role from consented house voices; every line passes the consent ledger right before it is spoken. Music comes from cleared cues with licence certificates, and sound from CC0 or public-domain recordings or code. The mix is mastered to a podcast or ACX-style audiobook spec, with chapters, captions, show notes with an AI disclosure, a C2PA credential, a signed record and sides for human actors.
In short
Last reviewed
- What it is
- Your script, produced tonight, with every voice consented and every track cleared.
- Who it's for
- Teams in film, tv and games and creative and media.
- Where it runs
- Hosted with public-domain and original scripts; self-host for unpublished work
- Key numbers
- 100% (389/389) Radio scripts: speaker right (code alone) (test split, n = 389)
- 99.0% (96/97) Sound cue mapped to the right library tag (model) (test split, n = 97)
- 87.5% (7/8) Cue with nothing in the library left as "no sound" (test split, n = 8)
- 28.0 s Median end-to-end run, hosted (QA sweep 2026-09-26)
How we tested itEnd-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 28 s · ~$0.001 per run · 1 receipt
Loading the nightly status…
Self-host: verified 26 Sep 2026 · Fresh clone of the branch into a clean directory, docker build of the api image and the docker/drama voice layer, compose with named volumes and host networking, pointed at the running local Qwen3.8-27B (direct route); rehearsal bundle; then torn down.
Measured cost to run: about $0.10 per 100 episodes (hosted, 26 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
Builds took 24 s and 100 s (warm cache). The rehearsal passed 16 of 16 checks in 20.7 s, including an 18 s render on 8 CPU threads: parse, an out-of-project performer refused, the house cast allowed, loudness in spec, C2PA with consent links, and the signed record verified. No speaker model in the image, so the voice check said not run.
Known limits (5)
- Kokoro voices are clear but flat: deliveries change pace and level only. British voices are Kokoro's weakest.
- Prose attribution is about 97% right on held-out stories: check the parse before rendering.
- ACX does not accept AI narration; the audiobook mode meets the technical spec only.
- C2PA credentials use a development certificate, so public validators show the issuer as untrusted.
- Composing new music needs the studio GPU queue; the hosted demo uses the cue library.
Eval results, nightly checks and cost per run · held-out eval
Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
1 law, rule and guidance page cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels, what it is built from
- Models
- Qwen3.8-27B (parse) · Kokoro-82M (voices, CPU) · ACE-Step 1.5 (music, via music-gen-cleared)
- Where
- Hosted with public-domain and original scripts; self-host for unpublished work
- Checks
- Receipt per model call; a signed consent decision per spoken line; C2PA credential linking every role's consent entry; signed production record
- Industry
- Film, TV and games · Creative and media
- Input
- Text and documents
- Output
- Media · Signed record or verdict
- Data
- Confidential business data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Part of
- Decosa Studio: Voice
- Runs in
- Decosa hosted · Self-host
- Built from
- Consent gate · Studio render · Content credentials · Signed record
Every result carries a signed record of which model produced it, so you can check it later. How that works
Questions people ask
Does ACX accept audiobooks made here?
No. ACX's requirements (dated 15 Apr 2026) prohibit unauthorised text-to-speech or AI narration. The audiobook mode meets ACX's technical spec only; use it for other platforms, classrooms and pitches.
Are the voices cloned from real people?
No. Only stock voices with a consent-ledger entry can be cast, and every role and line is checked against the ledger before it is spoken; a revoked or out-of-scope voice stops the render.
Will it change my script?
The spoken words are always the script's own: code splits the text, and the model only says who speaks and which sound a cue means. You see and can edit the parse before anything is voiced.
How accurate is the parse?
On held-out tests, speakers were right on 389 of 389 radio-script lines and 96.8% of prose lines, and sound cues mapped to the right library tag 99.0% of the time. The data is mostly synthetic, so check the parse before rendering.
What loudness does it deliver?
Measured on the delivered file: podcast -16 LUFS +/- 1 with true peak at most -1 dBTP, or ACX-style RMS, peak, noise floor and room tone per chapter file. All sample episodes passed.
How good are the voices?
Kokoro-82M voices are clear but flat: delivery notes change pace and level only, and British voices are its weakest. It is a produced temp track, and it exports sides for human actors.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Audio drama and narrated story studio
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…