51 · Film, TV and games · Creative and media · live
Audio drama and narrated story studio
Eval results
Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)
- Radio scripts: speaker right (code alone)100% (389/389)test splitn = 389Spoken lines found: also 100% (389/389)
- Sound cue mapped to the right library tag (model)99.0% (96/97)test splitn = 97Keyword code alone: 92.8% (90/97)
- Cue with nothing in the library left as "no sound"87.5% (7/8)test splitn = 8Varies run to run: a second pass missed three such cues
- Prose: speaker right96.8% (209/216)test splitn = 216Dev: 94.2% (65/69)
- Public-domain excerpts: speaker right100% (30/30)test splitn = 30Adapted excerpts of The Red-Headed League and The Monkey's Paw
- Word error rate of the finished episodes (machine transcription)0.6%-5.0%syntheticn = 44 sample episodes, 175-625 words each
- Parse costabout $0.0007 per parsetest splitn = 6262 receipted calls on the test split, $0.044 at the gateway list price
Dataset
Seeded synthetic radio scripts (8 dev, 30 test) and prose stories (8 dev, 30 test), plus hand-labelled public-domain excerpts (Poe as dev; Doyle and Jacobs adaptations as test); 4 rendered sample episodes for loudness and listening checks.
Caveats
- Mostly synthetic data from a generator written by the same author as the prompts; three prompt changes were made after looking at dev.
- The two test excerpts were written from memory, so they are adaptations, not exact texts.
- The eval's author cannot hear: acting, how effects land and music fit were not checked by ear. A human listen is needed before publishing.
- Results vary run to run on the shared gateway even at temperature 0 (the "no sound" row).
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 26 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 28 s
- Receipts
- 1
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.001
Self-host verification
Verified on 26 Sep 2026: Fresh clone of the branch into a clean directory, docker build of the api image and the docker/drama voice layer, compose with named volumes and host networking, pointed at the running local Qwen3.8-27B (direct route); rehearsal bundle; then torn down.
Builds took 24 s and 100 s (warm cache). The rehearsal passed 16 of 16 checks in 20.7 s, including an 18 s render on 8 CPU threads: parse, an out-of-project performer refused, the house cast allowed, loudness in spec, C2PA with consent links, and the signed record verified. No speaker model in the image, so the voice check said not run.
Rehearsal bundle: audio-drama-studio.zip (3 KB, 16 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Kokoro voices are clear but flat: deliveries change pace and level only. British voices are Kokoro's weakest.
- Prose attribution is about 97% right on held-out stories: check the parse before rendering.
- ACX does not accept AI narration; the audiobook mode meets the technical spec only.
- C2PA credentials use a development certificate, so public validators show the issuer as untrusted.
- Composing new music needs the studio GPU queue; the hosted demo uses the cue library.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Parse: voice hints, aliases and the sound and music cue mapping for radio scripts (one call); speaker attribution and sound suggestions for prose (one call per 36 quotations)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
- Speaks each line with a stock voicepack after the consent ledger allows it; one process per episode on CPUKokoro-82MApache-2.0
- Score: theme and sting cues rendered through the music-gen-cleared path (tool 37) with its prompt guard, similarity check and signed licence certificate; the hosted demo uses the library, composing new cues needs the studio GPUACE-Step 1.5 turbo + 5Hz LM 1.7BMIT
- Timeline, sound library, ducking, compression, limiter, loudness to spec, chapters, captions, sides, C2PA and the signed record (CPU)decosa-api drama module (decosa_api/verticals/drama) + FFmpegAGPL-3.0-or-later
- Consent ledger's speaker check (tool 47): does each role's rendered voice match the voice enrolled in its entry?ECAPA-TDNN speaker embeddings (ONNX export)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · CPU only, radio scripts (2)
- Speakers right on held-out radio scripts (code alone): 100% (389/389)decosa-api docs/evals/audio-drama-studio.md, 2026-09-26
- Cue mapped to the right library sound by keywords: 92.8% (90/97)decosa-api docs/evals/audio-drama-studio.md, 2026-09-26
Standard · the hosted demo (4)
- Prose speaker attribution (held-out): 96.8% planted (209/216); 100% public domain (30/30)decosa-api docs/evals/audio-drama-studio.md, 2026-09-26
- Cue mapping (held-out): 99.0% (96/97); music cues 26/26decosa-api docs/evals/audio-drama-studio.md, 2026-09-26
- Loudness spec met on delivered files: all sample episodes and chapter filesmeasured on our server 2026-09-26
- Word error rate heard back by ASR: 0.6-5.0% on 4 sample episodesdecosa-api docs/evals/audio-drama-studio.md, 2026-09-26
Best · compose new music per episode (2)
- Music prompt guard on held-out prompts: precision 100%, recall 98%decosa-api docs/evals/music-gen-cleared.md, 2026-09-25
- Composed cues in this studio: not measured yet