Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: Find the themes in your interviews, privately (use case 174)

29 Sep 2026, build-interview-themes-opus. Code: decosa_api/verticals/themes. Harness scripts are in the wiki scratch folder an internal measurement script pre-release branch (eval_asr.py, eval_prro.py, eval_own.py, build_sample.py). LLM: Qwen3.8-27B (direct route to the same weights for the eval runs; the hosted sample run went through the receipted gateway). No Anthropic API. Hand coding by a blind Claude Code sub-agent (Opus 5.5) playing a careful researcher.

Data and licences

Set What Licence Used for
NASA "Houston We Have a Podcast" 14 episodes, host interviews one guest; mp3 + NASA's own speaker-labelled transcripts (nasa.gov) US government work, public domain in the US (17 U.S.C. 105; NASA media guidelines) attribution error (2 dev + 12 test episodes, first 21 min each), the hosted sample study, cost and time
Same, degraded first 10 min of 4 test episodes pushed through a video-call chain (300-3,400 Hz, 8 kHz, pink noise ~20 dB SNR, Opus 16 kbit/s) as above attribution under call-quality audio
Knowledge Exchange PRRO interview coding 1,000 interview passages coded by the project team in NVivo, exported with the report structure (6 sections > 28 subsections > 102 codes); no transcripts or audio CC BY 4.0, Zenodo 10.5281/zenodo.5512420 code agreement with human coders, the own-model lens, the blind hand-coding baseline

1. Speaker attribution (the headline)

The complaint we measure: an AI tool credits the interviewer's words to a participant. Method: our transcripts are aligned word by word to NASA's published transcript (unique 3-gram anchors, then difflib); the host is the interviewer. Attribution error = aligned words whose role differs. "Other" voices (the podcast's intro clip) count as not-participant.

Condition (12 test interviews, 4.2 h, 41,624 aligned words) Error Interviewer words credited to a participant Participant words credited to the interviewer Worst interview
Diarizer in 10-min pieces, roles per piece (no voice linking) 4.0% 1,269 392 31%
Product: pieces + voices linked across pieces (ECAPA, CPU) + roles once per interview 0.05% (22 words) 4 18 0.28%
Product + voice check (moves long segments whose voice matches the other speaker) 0.05% 4 18 0.28%
Degraded video-call copy (4 x 10 min, 6,285 aligned words; one piece each, so no linking needed) 0.03% (2 words) 1 1 0.12%
  • The diarizer itself (MOSS-Transcribe-Diarize 0.9B) separates two studio voices almost perfectly. Nearly all the error in the unlinked condition comes from speaker ids that change meaning between pieces: a 15-minute limit forces long interviews into pieces, and naive stitching swaps who is who. Voice linking removes it.
  • The voice check never fired (0 segments moved in all sets). It is a safety net with no measured benefit here.
  • Coverage: about 39% of NASA's words were aligned (we transcribed the first 21 minutes of 35-70 minute episodes); the error rate is over aligned words only. ASR word accuracy was not scored.
  • Tuning: thresholds (LINK_SIM 0.50, MOVE_MARGIN 0.12) were set on the 2 dev episodes. After the first test build we saw extra speaker ids on the test interviews (the podcast's launch-countdown clip, and one voice split in two within a piece) and added the same-piece merge (0.65), the small-voice attach (under 5% of words, cosine >= 0.30) and the "other" role (a voice under 2% of the words and 60 words); the dev set was re-run first, then test. So the test set influenced the design of those two rules, not their thresholds.
  • Limits: clean podcast audio, professional hosts, one guest; the degraded copy is simulated, not a real call; no crosstalk-heavy or same-gender similar-voice set was scored. Expect more errors there.

2. Code agreement with human coders (PRRO, 1,000 passages)

Our segment coder (study.code_units, the product prompt, batches of 10) applied the published codebook. Primary code = its first code ("none" when it gave none). Human coding is one code per passage (8 passages carried two).

Codebook given to the coder Codes Cohen's kappa, all 1,000 kappa, test split (750) Accuracy
Code names only 6 sections 0.215 0.226 37.0%
Code names only 28 subsections 0.291 0.294 32.0%
Names + the published hierarchy as definitions, passages in export order 6 0.354 0.348 49.9%
same 28 0.567 0.552 59.1%
Names + hierarchy, passages shuffled (honest) 6 0.355 0.357 49.4%
same 28 0.516 0.489 54.3%
  • The export lists passages grouped by code, so a batch of 10 consecutive passages leaked its neighbours' group; the shuffled rows are the honest numbers (the leak was worth ~0.05 kappa at 28 codes).
  • The definitions variant was chosen after seeing the names-only result on the same passages; treat the test-split column as not held out. The 6-section codes are broad report sections, which is why they score lower than the 28.
  • Quotes: 1,033 of 1,042 coder quotes at 28 codes were word for word in their passage (99.1%); the 9 that were not are kept as code assignments without a quote and are never shown.

Blind second coder (120 passages from the test split; a blind sub-agent saw only the codebook and passages, coded by hand, no scripts): kappa with the published coding 0.624; our coder with the published coding 0.511 on the same 120; blind coder with our coder 0.706; Fleiss' kappa across the three 0.613. The two independent coders agree with each other more than either agrees with the published coding (which was one team's report-driven coding). Caveat: passage ids followed the export order; the blind coder noticed the grouping and says it coded on content.

3. Quote fidelity

  • Hosted sample run (12 interviews): 15 of 15 theme quotes word for word in the transcript and in a participant's turn (the check is code, not a model).
  • Planted test on 100 real sample quotes: clean 100/100 pass; one word changed 100/100 caught; one word dropped 100/100 caught; 10-word stretches of interviewer turns 100/100 refused; wrong speaker credited 100/100 caught.
  • Attribution of a quote is only as good as the transcript's roles: 4 interviewer words in 41,624 were credited to a participant on the test set (section 1).

4. Your own model (CPU)

bge-small-en-v1.5 (MIT, ONNX) sentence embeddings + one logistic regression per code, thresholds by cross-validation on the reviewed passages only. PRRO, test passages never used in training:

Codes Reviewed passages Own model vs human (kappa) Our LLM coder vs human (same passages) Own model vs LLM coder
28 100 0.156 0.489 0.146
28 250 0.284 0.489 0.288
28 250 + 250 LLM-coded 0.277 0.453 0.309
6 250 0.343 0.357 0.331
6 250 + 250 LLM-coded 0.333 0.362 0.411

With a small codebook and ~250 reviewed passages (about 2.5 interview hours: the sample averaged 95 passages per hour) the lab's own CPU model matches the LLM coder's agreement with the human coding; with 28 fine codes (~9 examples each) it does not. Training and coding 750 passages takes ~80 s on CPU. On the hosted sample (demo review = the coder's codes on 3 interviews accepted as they were) it agreed with the LLM coder at kappa 0.24 on the other 9 interviews. Honest reading: a useful second coder for small codebooks after a few reviewed interviews, not yet a replacement.

5. Cost and time

Hosted sample study (12 interviews, 4.2 h of audio), 29 Sep 2026:

  • Transcription: 1,095 GPU-s of diarizer time (10-minute pieces, one job at a time on a shared RTX PRO 6000; ~14x real time) + 87 s of CPU for voice embeddings. Live POST /themes/transcribe on one 21-minute sample: 99 s end to end.
  • Codebook proposal 42.8 s ($0.0257); coding 397 passages (278 got a code) and drafting 5 themes 74.5 s ($0.0691); own model 5.2 s CPU. (The recorded run on the page, made on production on 29 Sep 2026.)
  • At list prices: ~$0.10 per interview hour (content/pricing: LLM $0.023 + 261 GPU-s; range $0.06-0.19).
  • 30 one-hour interviews (extrapolated, linear in audio): transcription ~2.2 GPU-hours on one shared card, codebook ~1 min, coding and themes ~6 min: about 2.3 hours of machine time and about $2.90 (range $1.80-5.80). Page 85's estimate was $1-5.

6. Hours saved against hand coding

The blind researcher estimated ~4 hours of hand coding per interview hour even with an existing codebook (its 120 passages: ~4 h), i.e. ~120 hours for 30 one-hour interviews, near the low end of the 4-8 h methods rule of thumb. The tool does the transcription, first-pass codebook, coding, counts and quote checks in ~2.3 machine hours. The researcher's review time (editing the codebook, reading the coded passages and quotes, reviewing 3-4 interviews for the own model) was not measured; if it takes 15-25 hours, the saving is ~95-105 hours per 30-interview study. Transcription time (often outsourced at a per-minute price) is extra saving, not counted.

7. Rehearsal: checkable properties of the sample run

  1. /themes/samples/hwhap-12 has 12 interviews; every interview has exactly one interviewer voice.
  2. /themes/analyze on two sample interviews with the sample's approved codebook codes at least 10 passages and returns at least one theme.
  3. Every theme quote returned is ok (word for word, in a participant's turn).
  4. No passage (unit) comes from an interviewer or "other" turn.
  5. The signed record verifies at POST /record/verify, and a changed record fails.
  6. The hosted instance refuses a pasted transcript (403) and a study that is not a sample (400).

8. Not measured

Theme quality against a human thematic analysis (no public set with published themes and full transcripts); ASR word error rate; long crosstalk; non-English interviews (the pipeline is English-first; MOSS and Qwen are multilingual but untested here); researcher review time.