Skip to content
decosa

Produce an audio drama

Paste a radio script or a prose chapter and review the parse: every speaker, line, direction and sound cue, before anything is voiced. Cast each role from consented house voices; every line passes the consent ledger right before it is spoken. Music comes from cleared cues with licence certificates, and sound from CC0 or public-domain recordings or code. The mix is mastered to a podcast or ACX-style audiobook spec, with chapters, captions, show notes with an AI disclosure, a C2PA credential, a signed record and sides for human actors.

For
Teams in film, tv and games and creative and media.
  • Time per task28 stypical (median) on the sample
  • Cost per task~$0.10 per 100 episodesmeasured, at list price
  • Accuracy100% (389/389)Radio scripts: speaker right (code alone)All results and caveats
  • Self-hostRun it on your own server: nothing is sent to us or anyone else.
  • HostedUse it on our servers; see where your data goes.

Script

Holmes, Watson and a masked visitor. A radio-play script with rain, fire, footsteps, a knock, a sting and an off-mic voice. Arthur Conan Doyle, "A Scandal in Bohemia" (1891), adapted as a radio scene; public domain (published 1891; the author died in 1930).

Voices are stock Kokoro-82M voices with consent-ledger entries; nothing is cloned. ACX does not accept AI narration: the audiobook mode meets its technical spec, not its eligibility rules. Each demo session renders 3 episodes; files are deleted after 7 days.

Review, cast, render

Live

Parse the script. You will see every line, speaker and cue before anything is voiced, and can fix them.

What it does, in shortWho it's for, where it runs and the key results

Paste a radio script or a prose chapter and review the parse: every speaker, line, direction and sound cue, before anything is voiced. Cast each role from consented house voices; every line passes the consent ledger right before it is spoken. Music comes from cleared cues with licence certificates, and sound from CC0 or public-domain recordings or code. The mix is mastered to a podcast or ACX-style audiobook spec, with chapters, captions, show notes with an AI disclosure, a C2PA credential, a signed record and sides for human actors.

In short

Last reviewed

What it is
Your script, produced tonight, with every voice consented and every track cleared.
Who it's for
Teams in film, tv and games and creative and media.
Where it runs
Hosted with public-domain and original scripts; self-host for unpublished work
Key numbers
  • 100% (389/389) Radio scripts: speaker right (code alone) (test split, n = 389)
  • 99.0% (96/97) Sound cue mapped to the right library tag (model) (test split, n = 97)
  • 87.5% (7/8) Cue with nothing in the library left as "no sound" (test split, n = 8)
  • 28.0 s Median end-to-end run, hosted (QA sweep 2026-09-26)
All results, datasets and caveats
How we tested itEnd-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 28 s · ~$0.001 per run · 1 receipt

Loading the nightly status…

Self-host: verified 26 Sep 2026 · Fresh clone of the branch into a clean directory, docker build of the api image and the docker/drama voice layer, compose with named volumes and host networking, pointed at the running local Qwen3.8-27B (direct route); rehearsal bundle; then torn down.

Measured cost to run: about $0.10 per 100 episodes (hosted, 26 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Builds took 24 s and 100 s (warm cache). The rehearsal passed 16 of 16 checks in 20.7 s, including an 18 s render on 8 CPU threads: parse, an out-of-project performer refused, the house cast allowed, loudness in spec, C2PA with consent links, and the signed record verified. No speaker model in the image, so the voice check said not run.

Known limits (5)
  • Kokoro voices are clear but flat: deliveries change pace and level only. British voices are Kokoro's weakest.
  • Prose attribution is about 97% right on held-out stories: check the parse before rendering.
  • ACX does not accept AI narration; the audiobook mode meets the technical spec only.
  • C2PA credentials use a development certificate, so public validators show the issuer as untrusted.
  • Composing new music needs the studio GPU queue; the hosted demo uses the cue library.

Eval results, nightly checks and cost per run · held-out eval

Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

1 law, rule and guidance page cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels, what it is built from
Models
Qwen3.8-27B (parse) · Kokoro-82M (voices, CPU) · ACE-Step 1.5 (music, via music-gen-cleared)
Where
Hosted with public-domain and original scripts; self-host for unpublished work
Checks
Receipt per model call; a signed consent decision per spoken line; C2PA credential linking every role's consent entry; signed production record
Output
Media · Signed record or verdict
Data
Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Every result carries a signed record of which model produced it, so you can check it later. How that works

Questions people ask

Does ACX accept audiobooks made here?

No. ACX's requirements (dated 15 Apr 2026) prohibit unauthorised text-to-speech or AI narration. The audiobook mode meets ACX's technical spec only; use it for other platforms, classrooms and pitches.

Are the voices cloned from real people?

No. Only stock voices with a consent-ledger entry can be cast, and every role and line is checked against the ledger before it is spoken; a revoked or out-of-scope voice stops the render.

Will it change my script?

The spoken words are always the script's own: code splits the text, and the model only says who speaks and which sound a cue means. You see and can edit the parse before anything is voiced.

How accurate is the parse?

On held-out tests, speakers were right on 389 of 389 radio-script lines and 96.8% of prose lines, and sound cues mapped to the right library tag 99.0% of the time. The data is mostly synthetic, so check the parse before rendering.

What loudness does it deliver?

Measured on the delivered file: podcast -16 LUFS +/- 1 with true peak at most -1 dBTP, or ACX-style RMS, peak, noise floor and room tone per chapter file. All sample episodes passed.

How good are the voices?

Kokoro-82M voices are clear but flat: delivery notes change pace and level only, and British voices are its weakest. It is a produced temp track, and it exports sides for human actors.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Audio drama and narrated story studio

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.