174 · Science and research · preview
Interview themes
Eval results
Not held outRun 29 Sep 2026Eval write-up (decosa-api, access required)
- Words credited to the wrong speaker (12 public-domain interviews)0.05% (22 of 41,624)test splitn = 41,624Interviewer words put in a participant's mouth: 4. Without linking voices across 10-minute pieces: 4.0% (1,269 interviewer words credited to participants; worst interview 31%).
- Agreement with published human coding (Cohen's κ, 28 codes, test split)0.49test splitAll 1,000 passages: 0.52. With code names only as definitions: 0.29.
- Agreement with published human coding (Cohen's κ, 28 codes, all passages)0.52dev (tuned on)n = 1,000
- Our coder vs a blind second coder (Cohen's κ)0.71dev (tuned on)n = 120The blind second coder is a model playing a careful researcher, coding by hand.
- Blind second coder vs the published coding, for comparison (Cohen's κ)0.62dev (tuned on)n = 120
- Three coders: published, blind, ours (Fleiss' κ)0.61dev (tuned on)n = 120
- Quotes passing the word-for-word and speaker check15 of 15test splitn = 15Recorded sample run of 12 interviews. Planted test on 100 real quotes: a changed word, a dropped word, interviewer words and a wrong speaker were each caught 100 of 100.
Dataset
Speaker attribution: 12 episodes of NASA's Houston We Have a Podcast (US government work, public domain), first 21 minutes each (4.2 h), scored word by word against NASA's published transcripts. Coding: Knowledge Exchange PRRO interview coding (Zenodo 10.5281/zenodo.5512420, CC BY 4.0), 1,000 human-coded passages against a 28-code hierarchy, shuffled; a 120-passage sample coded blind by a second coder.
Caveats
- Attribution was measured on clean studio recordings with one guest; a simulated video-call copy of 4 interviews scored 0.03%, but real noisy calls, crosstalk and similar voices were not tested.
- Two speaker-linking rules were designed after seeing extra speaker ids on the test interviews (thresholds were set on 2 dev episodes).
- Coding agreement is one published dataset in one field; the definitions variant was chosen after the names-only run on the same passages, so the kappa is not held out.
- The blind second coder is a model playing a careful researcher, not a person.
- Agreement drops to 0.29 when codes have names only: definitions matter.
- Theme quality against a human thematic analysis was not measured.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 30 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 69 s
- Receipts
- 62
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.099
Self-host verification
partial on 29 Sep 2026: the branch's API run directly on a GPU server with DECOSA_THEMES_SAMPLES_ONLY=0 against the local model servers (not a fresh compose)
Pasted transcripts, codebook, coding, themes, own model and every export worked, and the rehearsal bundle passed 8 of 8; the docker compose in the assemble prompt was not run end to end.
Rehearsal bundle: interview-themes.zip (32 KB, 8 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted numbers are the whole sample task measured on production (codebook proposal, approval, then coding and themes for all 12 sample interviews, through the production API), run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
- Transcription runs about 14x faster than real time on one shared GPU: about 4 to 5 minutes per interview hour, one recording at a time.
- The hosted demo takes the public-domain sample interviews only.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Recordings in: speech recognition with speaker turns, one pass per 6 to 10 minute pieceMOSS-Transcribe-Diarize 0.9BApache-2.0
- Voice check: links each piece's speakers into one voice per person for the whole recording, and moves segments whose voice matches the other speaker (marked in the transcript)ECAPA-TDNN speaker embeddings (ONNX export)Apache-2.0
- Proposes the codebook, applies the approved codebook to every passage of participant talk (with the words that justify each code), groups codes into themes and picks candidate quotes; participant counts and the quote check are codeQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
- Your own model: sentence embeddings for the small classifier trained on your reviewed codes (one logistic-regression head per code)bge-small-en-v1.5 (ONNX)MIT
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 48 GB card (1)
- agreement with human coding: not measured yet
Standard · the hosted demo, one 96 GB card (3)
- Words credited to the wrong speaker, 12 public-domain interviews (41,624 words): 0.05% (4.0% without voice linking)decosa-api docs/evals/interview-themes.md, measured 2026-09-29
- Cohen's kappa with published human coding, 28 codes, 1,000 passages (test split): 0.52 (0.49)decosa-api interview-themes eval, PRRO coding (Zenodo 10.5281/zenodo.5512420, CC BY 4.0), measured 2026-09-29
- a blind second coder vs the published coding, 120 passages (for comparison): 0.62same eval, 2026-09-29
Wanted · a much larger coder on your own hardware (1)
- agreement with human coding: not measured yet