Eval: audio drama and narrated story studio (51)
Run on our server, 26 Sep 2026, branch the pre-release branch. Scripts: scripts/drama_eval.py (parse),
scripts/drama_listen.py (machine listening), scripts/drama_record_demo.py (the sample episodes). Raw results:
docs/evals/audio-drama-studio/ (planted.json, real.json, results-*.json, summary-*.json, listen.json).
1. Parse accuracy
Data. Built by drama_eval.py build (deterministic, seeded), split by seed before anything was run:
- Radio scripts (38: 8 dev, 30 test): two to four characters, one to three scenes, in five layouts (
NAME: line, Title-case names,NAMEover its dialogue, mixed), with parentheticals ((off),(on phone),(quietly),V.O.), continuation lines, inline cues inside dialogue, cues in seven notations ([SFX: ],SFX:,(SOUND: ),FX:,[Sound: ],(SFX: ),[SOUND EFFECT: ]), music cues, directions, headings and metadata lines. Cue phrasings come from a bank of 54 (48 with a gold tag or tag set, 6 with nothing in the library, whose right answer is "no sound") plus four "stops" and seven music cues. - Prose (38: 8 dev, 30 test): two-person conversations with tags after, before and around the speech, pronoun tags, action beats and untagged alternation; about a third told in the first person; curly, straight and British quotes; sound events in the narration ("A knock came at the door.") with a gold tag.
- Public domain, labelled by hand: The Cask of Amontillado (Poe, 1846; 27 quotations; dev, because it was used while building) and adapted excerpts of The Red-Headed League (Doyle, 1891; 18) and The Monkey's Paw (Jacobs, 1902; 12) as test. The two test excerpts were written from memory, so they are adaptations, not exact texts.
Tuning on dev (disclosed). Three prompt changes were made after looking at dev results: first-person narrators'
quotations go to the id narrator; a note on action beats, pronoun tags and alternation; and a softer wording of the
alternation rule after the first version broke the Poe excerpt (20 of 27). Two generator bugs were fixed on dev (pronoun
tags before both speakers had been named, which even a reader cannot resolve; ?/! followed by a comma) and one parser
bug (a qualified name like MARCUS (angrily) over its dialogue was missed). The test split was run once, after that.
| Held-out test | Result |
|---|---|
| Radio scripts: spoken lines found (text exact after normalisation) | 100% (389/389) |
| Radio scripts: speaker right (code alone; the model is not asked) | 100% (389/389) |
| Radio scripts: cues found (sound and music) | 100% (128/128), no extra blocks |
| Sound cue mapped to the right library tag (model) | 99.0% (96/97); keyword code alone: 92.8% (90/97) |
| Cue with nothing in the library left as "no sound" | 87.5% (7/8) |
| Layer (background / one-off / stop) right | 98.9% (93/94) |
| Music cue type (theme, sting, bed, stop) | 100% (26/26) |
| Prose: quotations found (verbatim) | 100% (216/216) |
| Prose: speaker right | 96.8% (209/216), none left unassigned |
| Prose: planted sound events suggested | 63 of 63 found; all 61 suggestions matched a planted event |
| Public-domain excerpts: speaker right | 100% (30/30) |
Dev for comparison: scripts 100%, prose 94.2% (65/69), the Poe excerpt 100% (27/27) after the prompt change.
Misses. Prose: 7 quotations in 3 stories, all untagged turns after a run of same-speaker lines (the model restarts the
alternation from the wrong person); the review step is there for this. Scripts: the only tag miss was one of the eight cues with no
library sound, mapped to a near sound instead of "no sound": the model prefers a near sound to none. A second pass over the 30 test scripts (not the scored run) missed three such cues
(a zip to whoosh, a typewriter to paper, the kettle to pour), so the "no sound" row varies run to run on the shared
gateway even at temperature 0; the mapped-tag rows did not change.
Cost and time. One model call per radio script, one per ≈36 quotations for prose: median 2.7 s, p90 4.8 s per parse on the gateway (max 14.3 s under load). Test split: 62 receipted calls, 63,312 prompt and 16,594 completion tokens, $0.044 at the gateway's list price ($0.30 / $1.50 per million), so about $0.0007 per parse.
2. Rendering: loudness, time and cost
The recorded sample episodes (scripts/drama_record_demo.py, the site's Watch runs), rendered through the pre-release server
(gateway route for the parse, Kokoro-82M on 8 CPU threads, library music), measured on the delivered files (FFmpeg
ebur128 for loudness and true peak; ACX numbers from the decoded MP3):
| Episode | Spec | Audio | Render | Per finished minute | Measured |
|---|---|---|---|---|---|
| A Scandal in Bohemia (26 lines, 4 roles, 3 voices) | podcast | 215.5 s | 31.0 s | 8.6 s | -16.6 LUFS, TP -1.9 dBTP |
| The Cask of Amontillado (prose, 44 lines + 2 credit lines) | ACX | 3 files, 302.9 s | 37.1 s | 7.4 s | RMS -20.3 dB each, peak -3.9 to -4.1 dB, floor -70.9 to -71.7 dB, 192 kbps |
| Night Shift (original, 20 lines, 9 cues) | podcast | 144.5 s | 19.7 s | 8.2 s | -16.5 LUFS, TP -2.6 dBTP |
| The Town Mouse (reader's theatre) | podcast | 94.7 s | 13.3 s | 8.4 s | -16.6 LUFS, TP -2.3 dBTP |
Every file met its spec, here and in the earlier development renders. Render time: 7.4-8.6 s per finished minute on CPU (the voices are about 60% of it); a self-hosted container on the same box rendered the rehearsal script in 18 s. Cost: the parse's model call ($0.0005-0.001); voices, sound and mixing are CPU time we do not meter; library music adds nothing. Composing new music adds two ACE-Step renders on the shared GPU (22-31 s each when measured for the library).
Loudness chain: 2.5:1 compression above -24 dB, gain aimed at the target in up to four measured passes, a look-ahead
limiter, then MP3; the true peak is measured after encoding and the ceiling lowered if the encoder pushed it over.
(The first version used FFmpeg loudnorm, which fell back to dynamic mode on dialogue and delivered -18 LUFS; caught on
the first sample render and replaced.)
3. Listening check (honest version)
The author of this eval is a language model and cannot hear. What was done instead, on the delivered MP3s:
- Intelligibility: MOSS-Transcribe-Diarize (the pass-2 ASR on GPU0) transcribed each episode; word error rate against the script's spoken lines was 1.7% (Holmes, 518 words), 5.0% (Poe, 625 words: "Amontillado", "flambeaux", "roquelaire" and friends), 2.8% (Night Shift, 249) and 0.6% (Town Mouse, 175).
- Voices: the diarizer found as many speakers as cast voices in every episode (3, 2, 4, 3); the consent ledger's speaker check matched every role's output to its enrolled voice (scores 0.72-0.96, threshold 0.585).
- Clipping and dead air: no clipped samples. Longest stretch below -55 dBFS: 0.5 s (Holmes), 1.1 s (Town Mouse), 3.1 s (Night Shift: a written pause and a scene change), 4.0 s (Poe: the room tone at the end of one ACX file and the start of the next, which is what ACX asks for). An earlier render of the Town Mouse had 6.0 s of near-silence at the end, because the outro took the last seconds of the theme, which ACE-Step had faded out; the outro now uses the theme's opening.
- Not checked by ear: whether the acting is convincing (it is not expressive: Kokoro has no emotion control), whether the effects land naturally, and whether the music suits the scene. A human listen is still needed before anyone publishes an episode, and the site says so.
4. Library assets and their licences
- 12 house voices (Kokoro-82M stock voicepacks, Apache-2.0), each enrolled in the consent ledger for project
decosa-audio-drama(narration, character dialogue, audiobook; WW) from a consent clip in its own voice. - 36 sound recordings, CC0 or public domain, each file's page checked on 26 Sep 2026 (Kenney RPG Audio and Impact Sounds,
OpenGameArt entries, Wikimedia Commons files; source URLs in
data/sfx/manifest.json); every other tag is generated by code (synth.py). BBC Sound Effects were not used (their licence is non-commercial). - Music cues rendered with ACE-Step 1.5 (MIT) through music-gen-cleared, one at a time on the shared GPU0. Of the first
8 renders, 6 were "clear" and 2 stings were "review" (above the review threshold, below the flag); those two were
re-rendered with a new seed rather than shipped unreviewed (recorded in
cues.json,replaced). Both re-renders were clear (gothic 3.46, suspense 4.50; the suspense one failed once on a busy GPU and was queued again). All 8 shipped cues are "clear".
5. Expected properties of the sample run (for the rehearsal kit)
rehearsal/audio-drama-studio/expected.json, on the original Night Shift script:
- It parses as a radio play with 4 roles, 20 spoken lines and every sound cue mapped (none unmatched), and the phone
voice carries the
phoneeffect. - Casting Rosa with the fictional performer Arthur Penhale is refused with
project_not_in_scope. - The suggested house cast is allowed for every role, and the episode renders with exactly 20 signed line decisions.
- The MP3 is within -16 LUFS +/- 1 dB with true peak at most -1 dBTP, has at least 2 chapters, and carries a C2PA credential whose consent assertion links every role.
- Captions, transcript, chapters, show notes, the sides pack and the record are delivered, and the record verifies.
6. Bugs found while building (fixed, with tests in tests/test_drama.py)
loudnormdelivered -18 LUFS on dialogue (see section 2).- FFT filters and resampling on odd lengths were very slow (pocketfft on large primes: 40 s of a 75 s render); padded to smooth sizes, and beds are generated as 20 s loops.
- A reserved field name in the record's music entry failed every render that had music (test added).
- ACX chapter files skipped a number and their captions were offset when a chapter was too short to keep.
- Browser run: a script with no spoken lines parsed "successfully" with nothing to render (now a warning, and the console disables Render); a sample whose music style is missing on the server made the render fail with a 400 (the console now falls back to an installed style).