Eval: Find the themes in your interviews, privately (use case 174)
29 Sep 2026, build-interview-themes-opus. Code: decosa_api/verticals/themes. Harness scripts are in the wiki scratch
folder an internal measurement script pre-release branch (eval_asr.py, eval_prro.py, eval_own.py, build_sample.py).
LLM: Qwen3.8-27B (direct route to the same weights for the eval runs; the hosted sample run went through the receipted
gateway). No Anthropic API. Hand coding by a blind Claude Code sub-agent (Opus 5.5) playing a careful researcher.
Data and licences
| Set | What | Licence | Used for |
|---|---|---|---|
| NASA "Houston We Have a Podcast" | 14 episodes, host interviews one guest; mp3 + NASA's own speaker-labelled transcripts (nasa.gov) | US government work, public domain in the US (17 U.S.C. 105; NASA media guidelines) | attribution error (2 dev + 12 test episodes, first 21 min each), the hosted sample study, cost and time |
| Same, degraded | first 10 min of 4 test episodes pushed through a video-call chain (300-3,400 Hz, 8 kHz, pink noise ~20 dB SNR, Opus 16 kbit/s) | as above | attribution under call-quality audio |
| Knowledge Exchange PRRO interview coding | 1,000 interview passages coded by the project team in NVivo, exported with the report structure (6 sections > 28 subsections > 102 codes); no transcripts or audio | CC BY 4.0, Zenodo 10.5281/zenodo.5512420 | code agreement with human coders, the own-model lens, the blind hand-coding baseline |
1. Speaker attribution (the headline)
The complaint we measure: an AI tool credits the interviewer's words to a participant. Method: our transcripts are aligned word by word to NASA's published transcript (unique 3-gram anchors, then difflib); the host is the interviewer. Attribution error = aligned words whose role differs. "Other" voices (the podcast's intro clip) count as not-participant.
| Condition (12 test interviews, 4.2 h, 41,624 aligned words) | Error | Interviewer words credited to a participant | Participant words credited to the interviewer | Worst interview |
|---|---|---|---|---|
| Diarizer in 10-min pieces, roles per piece (no voice linking) | 4.0% | 1,269 | 392 | 31% |
| Product: pieces + voices linked across pieces (ECAPA, CPU) + roles once per interview | 0.05% (22 words) | 4 | 18 | 0.28% |
| Product + voice check (moves long segments whose voice matches the other speaker) | 0.05% | 4 | 18 | 0.28% |
| Degraded video-call copy (4 x 10 min, 6,285 aligned words; one piece each, so no linking needed) | 0.03% (2 words) | 1 | 1 | 0.12% |
- The diarizer itself (MOSS-Transcribe-Diarize 0.9B) separates two studio voices almost perfectly. Nearly all the error in the unlinked condition comes from speaker ids that change meaning between pieces: a 15-minute limit forces long interviews into pieces, and naive stitching swaps who is who. Voice linking removes it.
- The voice check never fired (0 segments moved in all sets). It is a safety net with no measured benefit here.
- Coverage: about 39% of NASA's words were aligned (we transcribed the first 21 minutes of 35-70 minute episodes); the error rate is over aligned words only. ASR word accuracy was not scored.
- Tuning: thresholds (LINK_SIM 0.50, MOVE_MARGIN 0.12) were set on the 2 dev episodes. After the first test build we saw extra speaker ids on the test interviews (the podcast's launch-countdown clip, and one voice split in two within a piece) and added the same-piece merge (0.65), the small-voice attach (under 5% of words, cosine >= 0.30) and the "other" role (a voice under 2% of the words and 60 words); the dev set was re-run first, then test. So the test set influenced the design of those two rules, not their thresholds.
- Limits: clean podcast audio, professional hosts, one guest; the degraded copy is simulated, not a real call; no crosstalk-heavy or same-gender similar-voice set was scored. Expect more errors there.
2. Code agreement with human coders (PRRO, 1,000 passages)
Our segment coder (study.code_units, the product prompt, batches of 10) applied the published codebook. Primary code
= its first code ("none" when it gave none). Human coding is one code per passage (8 passages carried two).
| Codebook given to the coder | Codes | Cohen's kappa, all 1,000 | kappa, test split (750) | Accuracy |
|---|---|---|---|---|
| Code names only | 6 sections | 0.215 | 0.226 | 37.0% |
| Code names only | 28 subsections | 0.291 | 0.294 | 32.0% |
| Names + the published hierarchy as definitions, passages in export order | 6 | 0.354 | 0.348 | 49.9% |
| same | 28 | 0.567 | 0.552 | 59.1% |
| Names + hierarchy, passages shuffled (honest) | 6 | 0.355 | 0.357 | 49.4% |
| same | 28 | 0.516 | 0.489 | 54.3% |
- The export lists passages grouped by code, so a batch of 10 consecutive passages leaked its neighbours' group; the shuffled rows are the honest numbers (the leak was worth ~0.05 kappa at 28 codes).
- The definitions variant was chosen after seeing the names-only result on the same passages; treat the test-split column as not held out. The 6-section codes are broad report sections, which is why they score lower than the 28.
- Quotes: 1,033 of 1,042 coder quotes at 28 codes were word for word in their passage (99.1%); the 9 that were not are kept as code assignments without a quote and are never shown.
Blind second coder (120 passages from the test split; a blind sub-agent saw only the codebook and passages, coded by hand, no scripts): kappa with the published coding 0.624; our coder with the published coding 0.511 on the same 120; blind coder with our coder 0.706; Fleiss' kappa across the three 0.613. The two independent coders agree with each other more than either agrees with the published coding (which was one team's report-driven coding). Caveat: passage ids followed the export order; the blind coder noticed the grouping and says it coded on content.
3. Quote fidelity
- Hosted sample run (12 interviews): 15 of 15 theme quotes word for word in the transcript and in a participant's turn (the check is code, not a model).
- Planted test on 100 real sample quotes: clean 100/100 pass; one word changed 100/100 caught; one word dropped 100/100 caught; 10-word stretches of interviewer turns 100/100 refused; wrong speaker credited 100/100 caught.
- Attribution of a quote is only as good as the transcript's roles: 4 interviewer words in 41,624 were credited to a participant on the test set (section 1).
4. Your own model (CPU)
bge-small-en-v1.5 (MIT, ONNX) sentence embeddings + one logistic regression per code, thresholds by cross-validation on the reviewed passages only. PRRO, test passages never used in training:
| Codes | Reviewed passages | Own model vs human (kappa) | Our LLM coder vs human (same passages) | Own model vs LLM coder |
|---|---|---|---|---|
| 28 | 100 | 0.156 | 0.489 | 0.146 |
| 28 | 250 | 0.284 | 0.489 | 0.288 |
| 28 | 250 + 250 LLM-coded | 0.277 | 0.453 | 0.309 |
| 6 | 250 | 0.343 | 0.357 | 0.331 |
| 6 | 250 + 250 LLM-coded | 0.333 | 0.362 | 0.411 |
With a small codebook and ~250 reviewed passages (about 2.5 interview hours: the sample averaged 95 passages per hour) the lab's own CPU model matches the LLM coder's agreement with the human coding; with 28 fine codes (~9 examples each) it does not. Training and coding 750 passages takes ~80 s on CPU. On the hosted sample (demo review = the coder's codes on 3 interviews accepted as they were) it agreed with the LLM coder at kappa 0.24 on the other 9 interviews. Honest reading: a useful second coder for small codebooks after a few reviewed interviews, not yet a replacement.
5. Cost and time
Hosted sample study (12 interviews, 4.2 h of audio), 29 Sep 2026:
- Transcription: 1,095 GPU-s of diarizer time (10-minute pieces, one job at a time on a shared RTX PRO 6000; ~14x real
time) + 87 s of CPU for voice embeddings. Live
POST /themes/transcribeon one 21-minute sample: 99 s end to end. - Codebook proposal 42.8 s ($0.0257); coding 397 passages (278 got a code) and drafting 5 themes 74.5 s ($0.0691); own model 5.2 s CPU. (The recorded run on the page, made on production on 29 Sep 2026.)
- At list prices: ~$0.10 per interview hour (content/pricing: LLM $0.023 + 261 GPU-s; range $0.06-0.19).
- 30 one-hour interviews (extrapolated, linear in audio): transcription ~2.2 GPU-hours on one shared card, codebook ~1 min, coding and themes ~6 min: about 2.3 hours of machine time and about $2.90 (range $1.80-5.80). Page 85's estimate was $1-5.
6. Hours saved against hand coding
The blind researcher estimated ~4 hours of hand coding per interview hour even with an existing codebook (its 120 passages: ~4 h), i.e. ~120 hours for 30 one-hour interviews, near the low end of the 4-8 h methods rule of thumb. The tool does the transcription, first-pass codebook, coding, counts and quote checks in ~2.3 machine hours. The researcher's review time (editing the codebook, reading the coded passages and quotes, reviewing 3-4 interviews for the own model) was not measured; if it takes 15-25 hours, the saving is ~95-105 hours per 30-interview study. Transcription time (often outsourced at a per-minute price) is extra saving, not counted.
7. Rehearsal: checkable properties of the sample run
/themes/samples/hwhap-12has 12 interviews; every interview has exactly one interviewer voice./themes/analyzeon two sample interviews with the sample's approved codebook codes at least 10 passages and returns at least one theme.- Every theme quote returned is
ok(word for word, in a participant's turn). - No passage (unit) comes from an interviewer or "other" turn.
- The signed record verifies at
POST /record/verify, and a changed record fails. - The hosted instance refuses a pasted transcript (403) and a study that is not a sample (400).
8. Not measured
Theme quality against a human thematic analysis (no public set with published themes and full transcripts); ASR word error rate; long crosstalk; non-English interviews (the pipeline is English-first; MOSS and Qwen are multilingual but untested here); researcher review time.