Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: audio drama and narrated story studio (51)

Run on our server, 26 Sep 2026, branch the pre-release branch. Scripts: scripts/drama_eval.py (parse), scripts/drama_listen.py (machine listening), scripts/drama_record_demo.py (the sample episodes). Raw results: docs/evals/audio-drama-studio/ (planted.json, real.json, results-*.json, summary-*.json, listen.json).

1. Parse accuracy

Data. Built by drama_eval.py build (deterministic, seeded), split by seed before anything was run:

  • Radio scripts (38: 8 dev, 30 test): two to four characters, one to three scenes, in five layouts (NAME: line, Title-case names, NAME over its dialogue, mixed), with parentheticals ((off), (on phone), (quietly), V.O.), continuation lines, inline cues inside dialogue, cues in seven notations ([SFX: ], SFX:, (SOUND: ), FX:, [Sound: ], (SFX: ), [SOUND EFFECT: ]), music cues, directions, headings and metadata lines. Cue phrasings come from a bank of 54 (48 with a gold tag or tag set, 6 with nothing in the library, whose right answer is "no sound") plus four "stops" and seven music cues.
  • Prose (38: 8 dev, 30 test): two-person conversations with tags after, before and around the speech, pronoun tags, action beats and untagged alternation; about a third told in the first person; curly, straight and British quotes; sound events in the narration ("A knock came at the door.") with a gold tag.
  • Public domain, labelled by hand: The Cask of Amontillado (Poe, 1846; 27 quotations; dev, because it was used while building) and adapted excerpts of The Red-Headed League (Doyle, 1891; 18) and The Monkey's Paw (Jacobs, 1902; 12) as test. The two test excerpts were written from memory, so they are adaptations, not exact texts.

Tuning on dev (disclosed). Three prompt changes were made after looking at dev results: first-person narrators' quotations go to the id narrator; a note on action beats, pronoun tags and alternation; and a softer wording of the alternation rule after the first version broke the Poe excerpt (20 of 27). Two generator bugs were fixed on dev (pronoun tags before both speakers had been named, which even a reader cannot resolve; ?/! followed by a comma) and one parser bug (a qualified name like MARCUS (angrily) over its dialogue was missed). The test split was run once, after that.

Held-out test Result
Radio scripts: spoken lines found (text exact after normalisation) 100% (389/389)
Radio scripts: speaker right (code alone; the model is not asked) 100% (389/389)
Radio scripts: cues found (sound and music) 100% (128/128), no extra blocks
Sound cue mapped to the right library tag (model) 99.0% (96/97); keyword code alone: 92.8% (90/97)
Cue with nothing in the library left as "no sound" 87.5% (7/8)
Layer (background / one-off / stop) right 98.9% (93/94)
Music cue type (theme, sting, bed, stop) 100% (26/26)
Prose: quotations found (verbatim) 100% (216/216)
Prose: speaker right 96.8% (209/216), none left unassigned
Prose: planted sound events suggested 63 of 63 found; all 61 suggestions matched a planted event
Public-domain excerpts: speaker right 100% (30/30)

Dev for comparison: scripts 100%, prose 94.2% (65/69), the Poe excerpt 100% (27/27) after the prompt change.

Misses. Prose: 7 quotations in 3 stories, all untagged turns after a run of same-speaker lines (the model restarts the alternation from the wrong person); the review step is there for this. Scripts: the only tag miss was one of the eight cues with no library sound, mapped to a near sound instead of "no sound": the model prefers a near sound to none. A second pass over the 30 test scripts (not the scored run) missed three such cues (a zip to whoosh, a typewriter to paper, the kettle to pour), so the "no sound" row varies run to run on the shared gateway even at temperature 0; the mapped-tag rows did not change.

Cost and time. One model call per radio script, one per ≈36 quotations for prose: median 2.7 s, p90 4.8 s per parse on the gateway (max 14.3 s under load). Test split: 62 receipted calls, 63,312 prompt and 16,594 completion tokens, $0.044 at the gateway's list price ($0.30 / $1.50 per million), so about $0.0007 per parse.

2. Rendering: loudness, time and cost

The recorded sample episodes (scripts/drama_record_demo.py, the site's Watch runs), rendered through the pre-release server (gateway route for the parse, Kokoro-82M on 8 CPU threads, library music), measured on the delivered files (FFmpeg ebur128 for loudness and true peak; ACX numbers from the decoded MP3):

Episode Spec Audio Render Per finished minute Measured
A Scandal in Bohemia (26 lines, 4 roles, 3 voices) podcast 215.5 s 31.0 s 8.6 s -16.6 LUFS, TP -1.9 dBTP
The Cask of Amontillado (prose, 44 lines + 2 credit lines) ACX 3 files, 302.9 s 37.1 s 7.4 s RMS -20.3 dB each, peak -3.9 to -4.1 dB, floor -70.9 to -71.7 dB, 192 kbps
Night Shift (original, 20 lines, 9 cues) podcast 144.5 s 19.7 s 8.2 s -16.5 LUFS, TP -2.6 dBTP
The Town Mouse (reader's theatre) podcast 94.7 s 13.3 s 8.4 s -16.6 LUFS, TP -2.3 dBTP

Every file met its spec, here and in the earlier development renders. Render time: 7.4-8.6 s per finished minute on CPU (the voices are about 60% of it); a self-hosted container on the same box rendered the rehearsal script in 18 s. Cost: the parse's model call ($0.0005-0.001); voices, sound and mixing are CPU time we do not meter; library music adds nothing. Composing new music adds two ACE-Step renders on the shared GPU (22-31 s each when measured for the library).

Loudness chain: 2.5:1 compression above -24 dB, gain aimed at the target in up to four measured passes, a look-ahead limiter, then MP3; the true peak is measured after encoding and the ceiling lowered if the encoder pushed it over. (The first version used FFmpeg loudnorm, which fell back to dynamic mode on dialogue and delivered -18 LUFS; caught on the first sample render and replaced.)

3. Listening check (honest version)

The author of this eval is a language model and cannot hear. What was done instead, on the delivered MP3s:

  • Intelligibility: MOSS-Transcribe-Diarize (the pass-2 ASR on GPU0) transcribed each episode; word error rate against the script's spoken lines was 1.7% (Holmes, 518 words), 5.0% (Poe, 625 words: "Amontillado", "flambeaux", "roquelaire" and friends), 2.8% (Night Shift, 249) and 0.6% (Town Mouse, 175).
  • Voices: the diarizer found as many speakers as cast voices in every episode (3, 2, 4, 3); the consent ledger's speaker check matched every role's output to its enrolled voice (scores 0.72-0.96, threshold 0.585).
  • Clipping and dead air: no clipped samples. Longest stretch below -55 dBFS: 0.5 s (Holmes), 1.1 s (Town Mouse), 3.1 s (Night Shift: a written pause and a scene change), 4.0 s (Poe: the room tone at the end of one ACX file and the start of the next, which is what ACX asks for). An earlier render of the Town Mouse had 6.0 s of near-silence at the end, because the outro took the last seconds of the theme, which ACE-Step had faded out; the outro now uses the theme's opening.
  • Not checked by ear: whether the acting is convincing (it is not expressive: Kokoro has no emotion control), whether the effects land naturally, and whether the music suits the scene. A human listen is still needed before anyone publishes an episode, and the site says so.

4. Library assets and their licences

  • 12 house voices (Kokoro-82M stock voicepacks, Apache-2.0), each enrolled in the consent ledger for project decosa-audio-drama (narration, character dialogue, audiobook; WW) from a consent clip in its own voice.
  • 36 sound recordings, CC0 or public domain, each file's page checked on 26 Sep 2026 (Kenney RPG Audio and Impact Sounds, OpenGameArt entries, Wikimedia Commons files; source URLs in data/sfx/manifest.json); every other tag is generated by code (synth.py). BBC Sound Effects were not used (their licence is non-commercial).
  • Music cues rendered with ACE-Step 1.5 (MIT) through music-gen-cleared, one at a time on the shared GPU0. Of the first 8 renders, 6 were "clear" and 2 stings were "review" (above the review threshold, below the flag); those two were re-rendered with a new seed rather than shipped unreviewed (recorded in cues.json, replaced). Both re-renders were clear (gothic 3.46, suspense 4.50; the suspense one failed once on a busy GPU and was queued again). All 8 shipped cues are "clear".

5. Expected properties of the sample run (for the rehearsal kit)

rehearsal/audio-drama-studio/expected.json, on the original Night Shift script:

  1. It parses as a radio play with 4 roles, 20 spoken lines and every sound cue mapped (none unmatched), and the phone voice carries the phone effect.
  2. Casting Rosa with the fictional performer Arthur Penhale is refused with project_not_in_scope.
  3. The suggested house cast is allowed for every role, and the episode renders with exactly 20 signed line decisions.
  4. The MP3 is within -16 LUFS +/- 1 dB with true peak at most -1 dBTP, has at least 2 chapters, and carries a C2PA credential whose consent assertion links every role.
  5. Captions, transcript, chapters, show notes, the sides pack and the record are delivered, and the record verifies.

6. Bugs found while building (fixed, with tests in tests/test_drama.py)

  • loudnorm delivered -18 LUFS on dialogue (see section 2).
  • FFT filters and resampling on odd lengths were very slow (pocketfft on large primes: 40 s of a 75 s render); padded to smooth sizes, and beds are generated as 20 s loops.
  • A reserved field name in the record's music entry failed every render that had music (test added).
  • ACX chapter files skipped a number and their captions were offset when a chapter was too short to keep.
  • Browser run: a script with no spoken lines parsed "successfully" with nothing to render (now a warning, and the console disables Render); a sample whose music style is missing on the server made the render fail with a 400 (the console now falls back to an installed style).