{"schema_version":"1","site":"https://decosa.ai","id":"interview-themes","num":"174","name":"Interview themes","tool_name":"Find the themes in your interviews","short":"Interview themes","blurb":"For PhD students, academic researchers, UX researchers and program evaluators with 10 to 60 recorded interviews. Recordings come in as transcripts with every speaker kept apart; a voice check moves segments credited to the wrong person and marks them, and interviewer turns are never coded or quoted. An open model proposes a codebook of 8 to 12 codes that you rename, merge, split and approve; then every passage of participant talk is coded against it. Themes come with participant counts worked out in code and three quotes each, every quote checked word for word against the transcript and the person it is credited to, and playable from its timestamp. Export a REFI-QDA project, CSVs, a methods paragraph and a signed record of every model call; after you review a few interviews, a small model of your own codes the rest your way.","status":"preview","labels":{"industry":["science-research"],"job":["generate","review"],"input":["voice","text"],"deploy":["hosted","selfhost"],"status":"preview","output":["data","record"],"data":["confidential"],"hardware":"gpu-96","licence":"permissive"},"industries":["science-research"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["diarize","signed-record"],"models":"MOSS-Transcribe-Diarize 0.9B (speech and speakers) · ECAPA-TDNN voice check (CPU) · Qwen3.8-27B (codebook, coding, themes) · bge-small-en-v1.5 (your own model, CPU)","where":"Self-host, or a confidential enclave on request, for interviews under an ethics or IRB protocol; the hosted demo takes the public-domain sample interviews only","hardware":"1x RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the voice check, the quote check and your own model run on CPU","final_artifact":"An approved codebook with its edit history, coded passages, themes with participant counts and checked, playable quotes, a REFI-QDA project, CSVs, a methods paragraph and a signed record.","self_host_first":true,"verification":{"receipt_coverage":"full","summary":"Every quote re-checked word for word against the transcript and against its speaker; participant counts computed in code; interviewer turns excluded; signed record of every model call","manual_qa":{"hosted":{"date":"2026-09-30","result":"pass","p50_ms":69449,"p95_ms":75043,"runs":5,"receipts_per_run":62,"cost_per_run_usd":0.098598},"selfhost":{"date":"2026-09-29","result":"partial","method":"the branch's API run directly on a GPU server with DECOSA_THEMES_SAMPLES_ONLY=0 against the local model servers (not a fresh compose)","notes":"Pasted transcripts, codebook, coding, themes, own model and every export worked, and the rehearsal bundle passed 8 of 8; the docker compose in the assemble prompt was not run end to end."},"known_limits":["Hosted numbers are the whole sample task measured on production (codebook proposal, approval, then coding and themes for all 12 sample interviews, through the production API), run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.","Transcription runs about 14x faster than real time on one shared GPU: about 4 to 5 minutes per interview hour, one recording at a time.","The hosted demo takes the public-domain sample interviews only."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Words credited to the wrong speaker (12 public-domain interviews)","value":"0.05% (22 of 41,624)","unit":null,"n":41624,"split":"test","note":"Interviewer words put in a participant's mouth: 4. Without linking voices across 10-minute pieces: 4.0% (1,269 interviewer words credited to participants; worst interview 31%)."},{"name":"Agreement with published human coding (Cohen's κ, 28 codes, test split)","value":"0.49","unit":null,"n":null,"split":"test","note":"All 1,000 passages: 0.52. With code names only as definitions: 0.29."},{"name":"Agreement with published human coding (Cohen's κ, 28 codes, all passages)","value":"0.52","unit":null,"n":1000,"split":"dev","note":null},{"name":"Our coder vs a blind second coder (Cohen's κ)","value":"0.71","unit":null,"n":120,"split":"dev","note":"The blind second coder is a model playing a careful researcher, coding by hand."},{"name":"Blind second coder vs the published coding, for comparison (Cohen's κ)","value":"0.62","unit":null,"n":120,"split":"dev","note":null},{"name":"Three coders: published, blind, ours (Fleiss' κ)","value":"0.61","unit":null,"n":120,"split":"dev","note":null},{"name":"Quotes passing the word-for-word and speaker check","value":"15 of 15","unit":null,"n":15,"split":"test","note":"Recorded sample run of 12 interviews. Planted test on 100 real quotes: a changed word, a dropped word, interviewer words and a wrong speaker were each caught 100 of 100."}],"dataset":"Speaker attribution: 12 episodes of NASA's Houston We Have a Podcast (US government work, public domain), first 21 minutes each (4.2 h), scored word by word against NASA's published transcripts. Coding: Knowledge Exchange PRRO interview coding (Zenodo 10.5281/zenodo.5512420, CC BY 4.0), 1,000 human-coded passages against a 28-code hierarchy, shuffled; a 120-passage sample coded blind by a second coder.","held_out":false,"caveats":["Attribution was measured on clean studio recordings with one guest; a simulated video-call copy of 4 interviews scored 0.03%, but real noisy calls, crosstalk and similar voices were not tested.","Two speaker-linking rules were designed after seeing extra speaker ids on the test interviews (thresholds were set on 2 dev episodes).","Coding agreement is one published dataset in one field; the definitions variant was chosen after the names-only run on the same passages, so the kappa is not held out.","The blind second coder is a model playing a careful researcher, not a person.","Agreement drops to 0.29 when codes have names only: definitions matter.","Theme quality against a human thematic analysis was not measured."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/interview-themes"},"stack":{"summary":"For researchers with 10 to 60 recorded interviews under an ethics protocol. Transcripts keep every speaker apart and never code the interviewer; an open model proposes 8 to 12 codes that you rename, merge, split and approve; every passage is then coded against your codebook. Themes come with participant counts computed in code and three quotes each, checked word for word against the transcript and the speaker and playable at their timestamp. Exports go to REFI-QDA, CSV and a methods paragraph, with a signed record of every model call.","tagline":"Recorded interviews to a codebook you approve, every passage coded and every quote checked word for word, on open models you can run yourself.","deployment":"hosted-or-self-host","regulatory_note":"Not a replacement for ethics or IRB approval, and it does not decide findings: codes and themes are proposals the researcher reviews. Interviews under a protocol belong on your own hardware or in a confidential enclave (on request); the hosted demo takes public-domain sample interviews only.","components":[{"id":"diarizer","role":"Recordings in: speech recognition with speaker turns, one pass per 6 to 10 minute piece","name":"MOSS-Transcribe-Diarize 0.9B","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize","license":"Apache-2.0","params":"0.9B","quant":"BF16","vram_gb":null,"memory_gb_estimate":null,"engine":"transformers (trust_remote_code) + moss_transcribe_diarize package; hosted: decosa-diarize service, 127.0.0.1:8092","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["standard","lite","wanted"],"alternative_to":null},{"id":"voice","role":"Voice check: links each piece's speakers into one voice per person for the whole recording, and moves segments whose voice matches the other speaker (marked in the transcript)","name":"ECAPA-TDNN speaker embeddings (ONNX export)","hf_repo":"speechbrain/spkrec-ecapa-voxceleb","license":"Apache-2.0","params":null,"quant":"FP32","vram_gb":0,"memory_gb_estimate":null,"engine":"ONNX Runtime on CPU","receipt_coverage":"none","in_hosted_demo":true,"tiers":["all"],"alternative_to":null},{"id":"llm","role":"Proposes the codebook, applies the approved codebook to every passage of participant talk (with the words that justify each code), groups codes into themes and picks candidate quotes; participant counts and the quote check are code","name":"Qwen3.8-27B (NVIDIA NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP) + FP8 (attention/GDN), FP8 KV cache, MTP speculative decoding k=3","vram_gb":57,"memory_gb_estimate":null,"engine":"vLLM 0.29.0","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"embedder","role":"Your own model: sentence embeddings for the small classifier trained on your reviewed codes (one logistic-regression head per code)","name":"bge-small-en-v1.5 (ONNX)","hf_repo":"BAAI/bge-small-en-v1.5","license":"MIT","params":"33M","quant":"FP32","vram_gb":0,"memory_gb_estimate":null,"engine":"ONNX Runtime on CPU","receipt_coverage":"none","in_hosted_demo":true,"tiers":["all"],"alternative_to":null},{"id":"llm-lite","role":"Lite tier: the same codebook, coding and themes on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","hf_repo":"google/gemma-4-26B-A4B-it","license":"Apache-2.0 (model card also links the Gemma 4 licence page)","params":"25.2B","quant":"BF16 weights; FP8 at load time (vLLM --quantization fp8) to fit a 48 GB card","vram_gb":null,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 (Gemma4ForConditionalGeneration is in its model registry)","receipt_coverage":"none","in_hosted_demo":false,"tiers":["lite"],"alternative_to":null},{"id":"llm-best","role":"Wanted: a much larger coder for long studies and finer codebooks","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","hf_repo":"nvidia/DeepSeek-V4-Flash-NVFP4","license":"MIT","params":"284B","quant":"NVFP4 experts + FP8 (about 159-176 GB of weights)","vram_gb":192,"memory_gb_estimate":null,"engine":"vLLM B12X community build, TP2 on 2x RTX PRO 6000, MTP draft fixed by the kit's patches","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · one 48 GB card","summary":"The same codebook, coding and themes on a smaller mixture-of-experts model. Faster and cheaper; agreement on this task unknown.","components":["llm-lite","diarizer","voice","embedder"],"hardware":"1x L40S or RTX 6000 Ada 48 GB (not measured)","quality_evidence":[{"metric":"agreement with human coding","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"Direct route: calls are attested by the box's key; no gateway receipts.","hosting":null},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","summary":"Qwen3.8-27B proposes the codebook, codes every passage and groups themes; the diarizer and voice check make the transcripts; counts, the quote check and your own model are code on CPU.","components":["diarizer","voice","llm","embedder"],"hardware":"1x RTX PRO 6000 Blackwell 96 GB","quality_evidence":[{"metric":"Words credited to the wrong speaker, 12 public-domain interviews (41,624 words)","value":"0.05% (4.0% without voice linking)","source":"decosa-api docs/evals/interview-themes.md, measured 2026-09-29"},{"metric":"Cohen's kappa with published human coding, 28 codes, 1,000 passages (test split)","value":"0.52 (0.49)","source":"decosa-api interview-themes eval, PRRO coding (Zenodo 10.5281/zenodo.5512420, CC BY 4.0), measured 2026-09-29"},{"metric":"a blind second coder vs the published coding, 120 passages (for comparison)","value":"0.62","source":"same eval, 2026-09-29"}],"latency_note":"measured on the dev run: seconds each for the codebook and coding of two interviews; a few minutes to code a thousand passages (direct route)","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"On the gateway route every language-model call gets a gateway-signed receipt, listed in the signed record; the speech and voice steps are signed by the instance key (attestation).","hosting":null},{"id":"wanted","label":"Wanted · a much larger coder on your own hardware","summary":"DeepSeek-V4-Flash codes the study; it may come closer to a careful researcher on fine codebooks. Interviews stay on your own hardware, never on community providers. Not served yet.","components":["llm-best","diarizer","voice","embedder"],"hardware":"Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash (estimate)","quality_evidence":[{"metric":"agreement with human coding","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"On your own hardware its calls are attested by the box's key only; no gateway receipts.","hosting":"own-hardware"}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:0.1.0","purpose":"Transcript cleanup, the voice check, coding units, the quote check, participant counts, exports, your own model (CPU) and signing; the HTTP API (/themes/*). Binds 127.0.0.1 by default."},{"name":"decosa-llm","port":8000,"image":"${DECOSA_REGISTRY}/decosa-llm:0.1.0","purpose":"vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network."},{"name":"decosa-diarize","port":8092,"image":"${DECOSA_REGISTRY}/decosa-diarize:0.1.0","purpose":"MOSS-Transcribe-Diarize for recordings. Not needed when you bring transcripts."}],"tools":[{"name":"REFI-QDA Project and Codebook exchange formats","url":"https://www.qdasoftware.org/","license":"Open standard","purpose":"The .qdpx project and .qdc codebook exports, validated against the published XML schemas, which most QDA programs import."},{"name":"decosa record (vertical 07) and POST /record/verify","url":"https://decosa.ai/apps/record","license":"AGPL-3.0-or-later (decosa-api)","purpose":"The hash chain and the Ed25519-signed record of every model call; anyone can re-check it offline."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured: the hosted demo's Qwen3.8-27B and the diarizer run on our server's cards."},{"tier":"1x L40S / RTX 6000 Ada 48 GB","fits":null,"notes":"Not measured. Gemma 4 26B A4B (lite) for the text steps; transcribe elsewhere or bring transcripts."},{"tier":"CPU only","fits":true,"notes":"With transcripts you already have and a model endpoint elsewhere: the voice check, quote check, counts, exports, signing and your own model need no GPU."}],"latency":[{"lane":"codebook proposal, 2 interviews (161 passages)","typical_ms":18000,"source":"measured on our server 2026-09-29, direct route, one run (dev)"},{"lane":"coding and themes, 2 interviews (161 passages, 12 quotes checked)","typical_ms":18300,"source":"measured on our server 2026-09-29, direct route, one run (dev)"},{"lane":"coding 1,000 passages (28 codes)","typical_ms":140000,"source":"measured on our server 2026-09-29, PRRO eval, direct route"}],"benchmark":null,"notes":[]},"buyer_facts":[{"label":"What you get","value":"Speaker-separated transcripts with interviewer turns excluded; an approved codebook with its edit history; every passage coded; themes with participant counts and three checked quotes each, playable at their timestamp; a REFI-QDA project (.qdpx) and codebook (.qdc), which most QDA programs import; CSVs; a methods paragraph; a signed record."},{"label":"What it never does","value":"It never decides your findings (codes and themes are proposals you approve), never quotes a line it credits to the interviewer, never edits a quote, and never trains a model on what you send."},{"label":"Data retention","value":"Nothing kept on the server: audio and text live in memory for each request, and logs carry counts and timings only. The study lives in your browser until you export it; the recording is dropped after transcription and only its hash is kept in the record."},{"label":"Where it runs","value":"Self-host on your own server or laptop for interviews under an ethics or IRB protocol, or a confidential enclave on request. The hosted demo takes the public-domain sample interviews only."},{"label":"Your own model","value":"After you review a few interviews, a small classifier (MIT-licensed sentence embeddings plus one logistic-regression head per code, trained on CPU in seconds) codes the rest your way. It is returned to you and not kept."},{"label":"Cost","value":"Cents per interview hour at list price (model time plus GPU-minutes of transcription), measured on a sample of interviews and scaled."}],"data_handling":{"page":"/data#interview-themes","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Nothing kept on the server: audio and text live in memory for each request, and logs carry counts and timings only. The study lives in your browser until you export it; the recording is dropped after transcription and only its hash is kept in the record.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/tools/research/interview-themes","input":"themes","lanes":[{"id":"interviews","title":"Interviews","kind":"list"},{"id":"codebook","title":"Codebook","kind":"list"},{"id":"themes","title":"Themes","kind":"list"},{"id":"model","title":"Your own model","kind":"json"},{"id":"export","title":"Export and record","kind":"json"}],"samples":[],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/interview-themes-hosted.md","selfhost":"/prompts/interview-themes-selfhost.md","assemble":"/prompts/interview-themes-assemble.md","mac":null},"rehearsal":{"bundle":"/samples/interview-themes.zip","bundle_url":"https://decosa.ai/samples/interview-themes.zip","folder":"/samples/interview-themes/","expected":"/samples/interview-themes/expected.json","files":["/samples/interview-themes/expected.json","/samples/interview-themes/inputs/codebook-draft.json","/samples/interview-themes/inputs/codebook.json","/samples/interview-themes/inputs/interviews.json"],"bytes":32632,"checks":["at least 40 passages of participant talk were coded","no passage comes from the interviewer or another voice","at least three themes","every quote of the first theme is word for word in a participant's turn","a draft codebook is refused","the signed record verifies","a changed record fails","every model call has a signed receipt"],"licence":"Interview audio and NASA's transcripts are US government works (public domain in the US, 17 U.S.C. 105). The transcripts here are Decosa's machine transcription of that audio; the codebook was proposed by Decosa's tool and renamed by us. Part of decosa-api, AGPL-3.0-or-later.","about":"The first two interviews of the hosted sample study (NASA's Houston We Have a Podcast, first 21 minutes of episodes 407 and 409, transcribed with speakers separated) and the study's approved 10-code codebook. The coder must code participant talk only, every theme quote must be word for word in a participant's turn, a draft codebook must be refused, and the signed record must verify and fail when changed.","run":{"containers":"docker compose exec api python scripts/rehearse.py interview-themes","checkout":"python scripts/rehearse.py interview-themes --bundle interview-themes.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py interview-themes"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=interview-themes","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":44,"basis":"estimate","unknown":[]},{"id":"standard","gpu_gb":61.6,"basis":"estimate","unknown":[]},{"id":"wanted","gpu_gb":196,"basis":"estimate","unknown":[]}],"mac":null},"links":{"page":"/tools/research/interview-themes","json":"/use-cases/interview-themes.json","metrics":"/metrics/interview-themes","console":"/tools/research/interview-themes","console_sample":"/tools/research/interview-themes?sample=1&autorun=0","stack":"/tools/research/interview-themes#stack","try_live":"/tools/research/interview-themes","watch":"/tools/research/interview-themes","build":"/tools/research/interview-themes#build","self_host":"/tools/research/interview-themes#self-host","prompts":{"hosted":"/prompts/interview-themes-hosted.md","selfhost":"/prompts/interview-themes-selfhost.md","assemble":"/prompts/interview-themes-assemble.md","mac":null}}}