Skip to content
decosa

174 · Science and research · preview

Interview themes

Open the toolJSON

Eval results

Not held outRun 29 Sep 2026Eval write-up (decosa-api, access required)

  • Words credited to the wrong speaker (12 public-domain interviews)0.05% (22 of 41,624)test splitn = 41,624Interviewer words put in a participant's mouth: 4. Without linking voices across 10-minute pieces: 4.0% (1,269 interviewer words credited to participants; worst interview 31%).
  • Agreement with published human coding (Cohen's κ, 28 codes, test split)0.49test splitAll 1,000 passages: 0.52. With code names only as definitions: 0.29.
  • Agreement with published human coding (Cohen's κ, 28 codes, all passages)0.52dev (tuned on)n = 1,000
  • Our coder vs a blind second coder (Cohen's κ)0.71dev (tuned on)n = 120The blind second coder is a model playing a careful researcher, coding by hand.
  • Blind second coder vs the published coding, for comparison (Cohen's κ)0.62dev (tuned on)n = 120
  • Three coders: published, blind, ours (Fleiss' κ)0.61dev (tuned on)n = 120
  • Quotes passing the word-for-word and speaker check15 of 15test splitn = 15Recorded sample run of 12 interviews. Planted test on 100 real quotes: a changed word, a dropped word, interviewer words and a wrong speaker were each caught 100 of 100.

Dataset

Speaker attribution: 12 episodes of NASA's Houston We Have a Podcast (US government work, public domain), first 21 minutes each (4.2 h), scored word by word against NASA's published transcripts. Coding: Knowledge Exchange PRRO interview coding (Zenodo 10.5281/zenodo.5512420, CC BY 4.0), 1,000 human-coded passages against a 28-code hierarchy, shuffled; a 120-passage sample coded blind by a second coder.

Caveats

  • Attribution was measured on clean studio recordings with one guest; a simulated video-call copy of 4 interviews scored 0.03%, but real noisy calls, crosstalk and similar voices were not tested.
  • Two speaker-linking rules were designed after seeing extra speaker ids on the test interviews (thresholds were set on 2 dev episodes).
  • Coding agreement is one published dataset in one field; the definitions variant was chosen after the names-only run on the same passages, so the kappa is not held out.
  • The blind second coder is a model playing a careful researcher, not a person.
  • Agreement drops to 0.29 when codes have names only: definitions matter.
  • Theme quality against a human thematic analysis was not measured.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
30 Sep 2026
Latency, this run
n/a
p50 over passed runs
69 s
Receipts
62
Model calls
n/a
Tokens
n/a
Cost per run
$0.099

Self-host verification

partial on 29 Sep 2026: the branch's API run directly on a GPU server with DECOSA_THEMES_SAMPLES_ONLY=0 against the local model servers (not a fresh compose)

Pasted transcripts, codebook, coding, themes, own model and every export worked, and the rehearsal bundle passed 8 of 8; the docker compose in the assemble prompt was not run end to end.

Rehearsal bundle: interview-themes.zip (32 KB, 8 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted numbers are the whole sample task measured on production (codebook proposal, approval, then coding and themes for all 12 sample interviews, through the production API), run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
  • Transcription runs about 14x faster than real time on one shared GPU: about 4 to 5 minutes per interview hour, one recording at a time.
  • The hosted demo takes the public-domain sample interviews only.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Recordings in: speech recognition with speaker turns, one pass per 6 to 10 minute pieceMOSS-Transcribe-Diarize 0.9BApache-2.0
  • Voice check: links each piece's speakers into one voice per person for the whole recording, and moves segments whose voice matches the other speaker (marked in the transcript)ECAPA-TDNN speaker embeddings (ONNX export)Apache-2.0
  • Proposes the codebook, applies the approved codebook to every passage of participant talk (with the words that justify each code), groups codes into themes and picks candidate quotes; participant counts and the quote check are codeQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
  • Your own model: sentence embeddings for the small classifier trained on your reviewed codes (one logistic-regression head per code)bge-small-en-v1.5 (ONNX)MIT

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 48 GB card (1)
  • agreement with human coding: not measured yet
Standard · the hosted demo, one 96 GB card (3)
  • Words credited to the wrong speaker, 12 public-domain interviews (41,624 words): 0.05% (4.0% without voice linking)decosa-api docs/evals/interview-themes.md, measured 2026-09-29
  • Cohen's kappa with published human coding, 28 codes, 1,000 passages (test split): 0.52 (0.49)decosa-api interview-themes eval, PRRO coding (Zenodo 10.5281/zenodo.5512420, CC BY 4.0), measured 2026-09-29
  • a blind second coder vs the published coding, 120 passages (for comparison): 0.62same eval, 2026-09-29
Wanted · a much larger coder on your own hardware (1)
  • agreement with human coding: not measured yet

How we measure · All tools