Skip to content
decosa

Make and do

Find the themes in your interviews

A codebook you approve, every passage coded, and themes with participant counts and word-for-word quotes that play from the recording, on your own hardware.

69 s typical (median) on the sample; slowest 1 in 20: 75 s~$0.099 per interview hour in model time, measured, at list price0.05% (22 of 41,624) of words credited to the wrong speaker (interviewer vs participant), 12 public-domain interviewsDetails, API and self-host

Time and cost: coding and themes for the 12-interview sample. Transcription adds about 4–5 minutes and about $0.08 of GPU time per interview hour.

Mode

Nothing is saved on our servers; this page holds your study, and a copy stays in this browser on this device.

The file stays on your device; opening one never uploads it.

Loading the study…

How it works

  1. Recordings are transcribed with each speaker kept apart; a voice check moves segments credited to the wrong person and marks them. In the self-hosted version you can also paste transcripts you already have (Name: lines, WebVTT or SRT); the hosted demo takes the sample interviews only.
  2. You confirm who is the interviewer. Interviewer turns are never coded or quoted.
  3. An open model proposes 8 to 12 codes with definitions from the participants' talk. You rename, merge, split or delete them, then approve.
  4. Every passage of participant talk is coded against your approved codebook. Themes come with participant counts worked out in code and three quotes each; you rename, merge or reject them, untick passages and pick other checked quotes.
  5. Every quote is checked word for word against the transcript and against the person it is credited to, and plays from its timestamp.
  6. Export a REFI-QDA project, CSVs, the methods text and a signed record of every model call.

What it never does

  • It doesn't decide your findings: codes and themes are proposals you edit and approve, and the methods text says so.
  • It never quotes a line it credits to the interviewer, and never edits a quote.
  • It doesn't keep your study on a server: the page holds it, a copy stays in your browser on this device, and you can save it to a file and reopen it. The server keeps nothing after each request.
  • It doesn't redact names: a name someone says in an interview stays in its quote and in the exports. Check quotes before they go in a deck or a paper.
  • It doesn't train any model on what you send. The small model trained on your reviewed codes is handed to you, not kept.
  • It doesn't replace the ethics approval your study needs.

Today, and with this

Hand coding runs at about four hours per hour of interview: 30 one-hour interviews are about 120 hours before the first theme. Here a first codebook takes minutes, the coding of a thousand passages about two minutes, and you spend your time on the judgment calls.

Interviews under an ethics or IRB protocol belong on your own hardware or in a confidential enclave, on request. The hosted demo takes the public-domain sample interviews only.

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the voice check, the quote check and your own model run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the interview themes API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • The hosted API takes the public-domain sample interviews only, and refuses anything else. Your own interviews run on your own hardware (self-host) or in the confidential enclave.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
interview-themes

Use the hosted API

# Decosa interview themes: use the hosted API

You are wiring Decosa interview themes into this project. It returns:
- speaker-separated transcripts with the interviewer's turns kept apart (never coded or quoted);
- a proposed codebook of 8 to 12 codes that the researcher edits and approves;
- every passage of participant talk coded against the approved codebook;
- themes with participant counts and three quotes each, every quote checked word for word against the transcript and its speaker;
- REFI-QDA and CSV exports, and a signed record of every model call.

Every model call has a signed receipt. Use only what is listed below; if you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes the public-domain sample interviews only** (it refuses other interviews with HTTP 400 or 403). Interviews under an ethics or IRB protocol belong on a self-hosted box (see the self-host prompt) or a confidential enclave on request. Say so wherever this is wired in.
- It proposes codes and themes; the researcher decides them. Nothing is analysed until the codebook's `status` is `"approved"`.
- The server keeps nothing: every call sends the interviews it needs, and your client holds the study.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "interview-themes"}` returns `{"token", "expires_at", "budget"}`. Demo sessions are limited per IP per hour; over a limit you get HTTP 429 with `Retry-After`. A token runs one request at a time (409 otherwise); 402 means its budget is used up.

## Endpoints
The POSTs below stream server-sent events (`data: {...}` lines): `progress`, `receipt`, `result` (the payload), `error` (`message`), `budget`, `done`.
- `GET /themes/info`: what it does and doesn't, limits, and the data-flow statement per mode (self-host, confidential, hosted).
- `GET /themes/samples`, `GET /themes/samples/{id}`: the sample studies (interviews with timed turns, and a recorded run). `GET /themes/audio/{id}/{interview}`: the recording (mp3, range requests).
- `POST /themes/transcript` `{sample, interview}` and `POST /themes/transcribe?sample=&interview=`: a sample interview's transcript, or transcribed live (minutes per hour of audio).
- `POST /themes/codebook` `{interviews, question, n_codes?: [8, 12]}` → `codebook` (status `proposed`).
- `POST /themes/codebook/split` `{code, passages, hint?}` → two narrower codes.
- `POST /themes/analyze` `{interviews, codebook (status "approved", with its edit history), question}` → `units`, `coded`, `themes` (with `participants`, `passages`, `quotes[]` each with `ok` and `reasons`), `quotes {total, ok}`, `record`, `usage`.
- `POST /themes/train` `{interviews, codebook, reviewed: {passage id: [code ids]}, coded}` → your own small model (`model_b64`), its `card`, its codes for the rest, and its agreement with the model coder.
- `POST /themes/export` `{interviews, codebook, coded, themes, records, methods, format}` → a file: `zip`, `project.qdpx`, `codebook.qdc`, `codebook.csv`, `segments.csv`, `themes.md`, `methods.md` or `record.json`.
- `POST /record/verify` (no auth): re-check a signed record.

Show quotes with their participant id and timestamp, and show a failed check's `reasons` instead of hiding the quote. Never present the themes as findings; they are proposals for the researcher.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa interview themes: run it yourself (containers)

You are setting up Decosa interview themes on this machine, so the interviews never leave it. It returns:
- speaker-separated transcripts with the interviewer's turns kept apart (never coded or quoted);
- a proposed codebook of 8 to 12 codes that I edit and approve;
- every passage of participant talk coded against the approved codebook;
- themes with participant counts and three quotes each, every quote checked word for word against the transcript and its speaker;
- REFI-QDA and CSV exports, a methods paragraph and a signed record of every model call.

Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with "not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/interview-themes.zip (32 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py interview-themes` (the api image carries the same bundle under /app/rehearsal/interview-themes/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py interview-themes --bundle interview-themes.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least 40 passages of participant talk were coded", "no passage comes from the interviewer or another voice", "at least three themes"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin (docs.docker.com/engine/install), and the NVIDIA container toolkit; check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service; on `api` set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and `DECOSA_THEMES_SAMPLES_ONLY=0` (so it takes my own interviews); keep its data on a named volume; bind every port to 127.0.0.1. For recordings, also run the diarizer and set `DECOSA_DIARIZE_URL` (the assemble prompt has the details); transcripts need no diarizer.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/themes/info` answers with `"samples_only": false`; `GET /attest/signing-key` shows this box's public key.
5. Smoke test: get a token with `POST /demo/session {"vertical":"interview-themes"}`, then `POST /themes/transcript` with `{"text": "Interviewer: What is a normal day like?\nMaria: We start at seven with a check-in, then patient calls all morning.\nInterviewer: What helps?\nMaria: The routine helps, and talking with colleagues at lunch keeps me going."}` (a label counts once it appears twice or reads as a role, like "Interviewer"). Expect a `result` event whose interview has four turns, alternating `interviewer` and `participant` (Maria); every receipt has `"status": "attested"` and the record verifies with `POST /record/verify`.

Finish with a summary and these reminders:
- Interview audio and transcripts stay on this machine; the model route stays local (`direct`) and the API keeps nothing after each request.
- It proposes codes and themes; I decide them. It does not replace my ethics approval.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    MOSS-Transcribe-Diarize 0.9B needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (61.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (61.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Replace MOSS-Transcribe-Diarize 0.9B with MOSS-Transcribe-Diarize, MLX 8-bit. MLX build for Apple Silicon. (Memory is an estimate.)

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits with changes: Replace MOSS-Transcribe-Diarize 0.9B with MOSS-Transcribe-Diarize, MLX 8-bit. MLX build for Apple Silicon. (Memory is an estimate.)

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"interview-themes"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py interview-themes

Download the mock-data bundle (32 KB, 8 checks)expected.json

The first two interviews of the hosted sample study (NASA's Houston We Have a Podcast, first 21 minutes of episodes 407 and 409, transcribed with speakers separated) and the study's approved 10-code codebook. The coder must code participant talk only, every theme quote must be word for word in a participant's turn, a draft codebook must be refused, and the signed record must verify and fail when changed.

What the rehearsal checks
  • at least 40 passages of participant talk were coded
  • no passage comes from the interviewer or another voice
  • at least three themes
  • every quote of the first theme is word for word in a participant's turn
  • a draft codebook is refused
  • the signed record verifies
  • a changed record fails
  • every model call has a signed receipt

Licence: Interview audio and NASA's transcripts are US government works (public domain in the US, 17 U.S.C. 105). The transcripts here are Decosa's machine transcription of that audio; the codebook was proposed by Decosa's tool and renamed by us. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa interview themes: run it yourself (containers)

You are setting up Decosa interview themes on this machine, so the interviews never leave it. It returns:
- speaker-separated transcripts with the interviewer's turns kept apart (never coded or quoted);
- a proposed codebook of 8 to 12 codes that I edit and approve;
- every passage of participant talk coded against the approved codebook;
- themes with participant counts and three quotes each, every quote checked word for word against the transcript and its speaker;
- REFI-QDA and CSV exports, a methods paragraph and a signed record of every model call.

Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with "not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/interview-themes.zip (32 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py interview-themes` (the api image carries the same bundle under /app/rehearsal/interview-themes/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py interview-themes --bundle interview-themes.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least 40 passages of participant talk were coded", "no passage comes from the interviewer or another voice", "at least three themes"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin (docs.docker.com/engine/install), and the NVIDIA container toolkit; check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service; on `api` set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and `DECOSA_THEMES_SAMPLES_ONLY=0` (so it takes my own interviews); keep its data on a named volume; bind every port to 127.0.0.1. For recordings, also run the diarizer and set `DECOSA_DIARIZE_URL` (the assemble prompt has the details); transcripts need no diarizer.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/themes/info` answers with `"samples_only": false`; `GET /attest/signing-key` shows this box's public key.
5. Smoke test: get a token with `POST /demo/session {"vertical":"interview-themes"}`, then `POST /themes/transcript` with `{"text": "Interviewer: What is a normal day like?\nMaria: We start at seven with a check-in, then patient calls all morning.\nInterviewer: What helps?\nMaria: The routine helps, and talking with colleagues at lunch keeps me going."}` (a label counts once it appears twice or reads as a role, like "Interviewer"). Expect a `result` event whose interview has four turns, alternating `interviewer` and `participant` (Maria); every receipt has `"status": "attested"` and the record verifies with `POST /record/verify`.

Finish with a summary and these reminders:
- Interview audio and transcripts stay on this machine; the model route stays local (`direct`) and the API keeps nothing after each request.
- It proposes codes and themes; I decide them. It does not replace my ethics approval.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsInterview themes on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

Standard · the hosted demo, one 96 GB card: what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Recordings in: MOSS-Transcribe-Diarize 0.9B. ~4 GB, weights 1.8 GB (estimate). MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate.
  • Voice check: ECAPA-TDNN speaker embeddings (ONNX export). CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Proposes the codebook, applies the approved c...: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)
  • Your own model: bge-small-en-v1.5 (ONNX). CPU. Runs on CPU (vram_gb 0 in stack.json).

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Interview themes, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Interview themes on my hardware

Fetch https://decosa.ai/prompts/interview-themes-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=interview-themes)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 32 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Recordings in: MOSS-Transcribe-Diarize 0.9B (OpenMOSS-Team/MOSS-Transcribe-Diarize), 4 GB
- Voice check: ECAPA-TDNN speaker embeddings (ONNX export) (speechbrain/spkrec-ecapa-voxceleb), CPU
- Proposes the codebook, applies the approved c...: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- Your own model: bge-small-en-v1.5 (ONNX) (BAAI/bge-small-en-v1.5), CPU

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%), MOSS-Transcribe-Diarize 0.9B ~4 GB (13%); about 0 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/interview-themes-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 30 Sep 2026 · measured 30 Sep 2026: · p50 69 s · p95 75 s (5 runs) · ~$0.099 per run · 62 receipts

Loading the nightly status…

Self-host: partial on 29 Sep 2026 · the branch's API run directly on a GPU server with DECOSA_THEMES_SAMPLES_ONLY=0 against the local model servers (not a fresh compose)

Measured cost to run: about $0.099 per interview hour (hosted, 30 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Pasted transcripts, codebook, coding, themes, own model and every export worked, and the rehearsal bundle passed 8 of 8; the docker compose in the assemble prompt was not run end to end.

Known limits (3)
  • Hosted numbers are the whole sample task measured on production (codebook proposal, approval, then coding and themes for all 12 sample interviews, through the production API), run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
  • Transcription runs about 14x faster than real time on one shared GPU: about 4 to 5 minutes per interview hour, one recording at a time.
  • The hosted demo takes the public-domain sample interviews only.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the voice check, the quote check and your own model run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the interview themes API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • The hosted API takes the public-domain sample interviews only, and refuses anything else. Your own interviews run on your own hardware (self-host) or in the confidential enclave.
The open stack

Recorded interviews to a codebook you approve, every passage coded and every quote checked word for word, on open models you can run yourself.

For researchers with 10 to 60 recorded interviews under an ethics protocol. Transcripts keep every speaker apart and never code the interviewer; an open model proposes 8 to 12 codes that you rename, merge, split and approve; every passage is then coded against your codebook. Themes come with participant counts computed in code and three quotes each, checked word for word against the transcript and the speaker and playable at their timestamp. Exports go to REFI-QDA, CSV and a methods paragraph, with a signed record of every model call.

Deployment
Hosted or self-host
Regulatory
Not a replacement for ethics or IRB approval, and it does not decide findings: codes and themes are proposals the researcher reviews. Interviews under a protocol belong on your own hardware or in a confidential enclave (on request); the hosted demo takes public-domain sample interviews only.
Architecture
Text description

Recordings (or transcripts) go in. The diarizer transcribes each piece with speaker turns and a CPU voice check links voices across pieces and moves misattributed segments. The researcher confirms who is the interviewer. Qwen3.8-27B proposes a codebook, the researcher edits and approves it, then the model codes every passage of participant talk. Code counts participants and checks every quote word for word against the transcript and the speaker. Out come themes with playable quotes, REFI-QDA and CSV exports, a methods paragraph, an optional small model trained on the researcher's codes on CPU, and a signed record of every model call. Self-hosted, everything stays on the researcher's machine.

Architecture

At a glance

What you get
Speaker-separated transcripts with interviewer turns excluded; an approved codebook with its edit history; every passage coded; themes with participant counts and three checked quotes each, playable at their timestamp; a REFI-QDA project (.qdpx) and codebook (.qdc), which most QDA programs import; CSVs; a methods paragraph; a signed record.
What it never does
It never decides your findings (codes and themes are proposals you approve), never quotes a line it credits to the interviewer, never edits a quote, and never trains a model on what you send.
Data retention
Nothing kept on the server: audio and text live in memory for each request, and logs carry counts and timings only. The study lives in your browser until you export it; the recording is dropped after transcription and only its hash is kept in the record.
Where it runs
Self-host on your own server or laptop for interviews under an ethics or IRB protocol, or a confidential enclave on request. The hosted demo takes the public-domain sample interviews only.
Your own model
After you review a few interviews, a small classifier (MIT-licensed sentence embeddings plus one logistic-regression head per code, trained on CPU in seconds) codes the rest your way. It is returned to you and not kept.
Cost
Cents per interview hour at list price (model time plus GPU-minutes of transcription), measured on a sample of interviews and scaled.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 48 GB card

    The same codebook, coding and themes on a smaller mixture-of-experts model. Faster and cheaper; agreement on this task unknown.

    Models
    • MOSS-Transcribe-Diarize 0.9B
    • ECAPA-TDNN speaker embeddings (ONNX export)
    • bge-small-en-v1.5 (ONNX)
    • Gemma 4 26B A4B (instruction-tuned)
    Hardware
    1x L40S or RTX 6000 Ada 48 GB (not measured)
    Quality evidence
    • agreement with human codingnot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
  • In the hosted demo

    Standard

    the hosted demo, one 96 GB card

    Qwen3.8-27B proposes the codebook, codes every passage and groups themes; the diarizer and voice check make the transcripts; counts, the quote check and your own model are code on CPU.

    Models
    • MOSS-Transcribe-Diarize 0.9B
    • ECAPA-TDNN speaker embeddings (ONNX export)
    • Qwen3.8-27B (NVIDIA NVFP4)
    • bge-small-en-v1.5 (ONNX)
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • Words credited to the wrong speaker, 12 public-domain interviews (41,624 words)0.05% (4.0% without voice linking)decosa-api docs/evals/interview-themes.md, measured 2026-09-29
    • Cohen's kappa with published human coding, 28 codes, 1,000 passages (test split)0.52 (0.49)decosa-api interview-themes eval, PRRO coding (Zenodo 10.5281/zenodo.5512420, CC BY 4.0), measured 2026-09-29
    • a blind second coder vs the published coding, 120 passages (for comparison)0.62same eval, 2026-09-29
    Latency
    measured on the dev run: seconds each for the codebook and coding of two interviews; a few minutes to code a thousand passages (direct route)
    Verification
    Proof: strongOn the gateway route every language-model call gets a gateway-signed receipt, listed in the signed record; the speech and voice steps are signed by the instance key (attestation).
  • Needs more compute

    Wanted: the best setup

    a much larger coder on your own hardware

    DeepSeek-V4-Flash codes the study; it may come closer to a careful researcher on fine codebooks. Interviews stay on your own hardware, never on community providers. Not served yet.

    Models
    • MOSS-Transcribe-Diarize 0.9B
    • ECAPA-TDNN speaker embeddings (ONNX export)
    • bge-small-en-v1.5 (ONNX)
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    Hardware
    Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash (estimate)
    Quality evidence
    • agreement with human codingnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only; no gateway receipts.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Recordings in: speech recognition with speaker turns, one pass per 6 to 10 minute pieceMOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab)
0.9BProof: partialIn the hosted demo
Voice check: links each piece's speakers into one voice per person for the whole recording, and moves segments whose voice matches the other speaker (marked in the transcript)ECAPA-TDNN speaker embeddings (ONNX export)speechbrain/spkrec-ecapa-voxceleb on Hugging Face (opens in a new tab)
0 GBNo proof yetIn the hosted demo
Proposes the codebook, applies the approved codebook to every passage of participant talk (with the words that justify each code), groups codes into themes and picks candidate quotes; participant counts and the quote check are codeQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
Your own model: sentence embeddings for the small classifier trained on your reviewed codes (one logistic-regression head per code)bge-small-en-v1.5 (ONNX)BAAI/bge-small-en-v1.5 on Hugging Face (opens in a new tab)
33M · 0 GBNo proof yetIn the hosted demo
Lite tier: the same codebook, coding and themes on a 48 GB cardGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
25.2B (3.8B active)No proof yetSelf-host only
Wanted: a much larger coder for long studies and finer codebooksDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active) · 192 GBNo proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    Transcript cleanup, the voice check, coding units, the quote check, participant counts, exports, your own model (CPU) and signing; the HTTP API (/themes/*). Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.

  • decosa-diarize:8092
    ${DECOSA_REGISTRY}/decosa-diarize:0.1.0

    MOSS-Transcribe-Diarize for recordings. Not needed when you bring transcripts.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured: the hosted demo's Qwen3.8-27B and the diarizer run on our server's cards.

  • 1x L40S / RTX 6000 Ada 48 GB

    Not measured. Gemma 4 26B A4B (lite) for the text steps; transcribe elsewhere or bring transcripts.

  • CPU only Fits

    With transcripts you already have and a model endpoint elsewhere: the voice check, quote check, counts, exports, signing and your own model need no GPU.

Latency per lane

  • codebook proposal, 2 interviews (161 passages)18.0 s

    Measuredmeasured on our server 2026-09-29, direct route, one run (dev)

  • coding and themes, 2 interviews (161 passages, 12 quotes checked)18.3 s

    Measuredmeasured on our server 2026-09-29, direct route, one run (dev)

  • coding 1,000 passages (28 codes)140.0 s

    Measuredmeasured on our server 2026-09-29, PRRO eval, direct route

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

interview-themes/assemble-prompt.md168 lines
# Assemble Decosa interview themes on this machine

You are setting up Decosa interview themes on this Linux machine for a researcher (or a lab) with recorded interviews under an ethics or IRB protocol. It returns:
- speaker-separated transcripts, with interviewer turns kept apart and never coded or quoted;
- a proposed codebook of 8 to 12 codes that I rename, merge, split and approve before anything is coded;
- every passage of participant talk coded against the approved codebook;
- themes with participant counts (computed in code) and three quotes each, every quote checked word for word against the transcript and the person it is credited to;
- a REFI-QDA project (.qdpx) and codebook (.qdc), CSVs, a methods paragraph and a signed record of every model call;
- optionally, a small model trained on my own reviewed codes that codes the rest (CPU).

Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- Interview audio and transcripts stay on this machine: the model route is local (`direct`), `DECOSA_THEMES_SAMPLES_ONLY=0` lets the API take my own files, and the API keeps nothing after each request. The study lives in my browser or my own client until I export it.
- It proposes codes and themes; I decide them. It does not replace my ethics approval.

Repeat both points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/interview-themes.zip (32 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py interview-themes` (the api image carries the same bundle under /app/rehearsal/interview-themes/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py interview-themes --bundle interview-themes.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least 40 passages of participant talk were coded", "no passage comes from the interviewer or another voice", "at least three themes"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | voice check `speechbrain/spkrec-ecapa-voxceleb` (Apache-2.0, CPU); your own model `BAAI/bge-small-en-v1.5` (MIT, CPU) | `127.0.0.1:8445` |
| `diarize` (only for recordings) | built from the decosa-api source (`services/diarize`) | `OpenMOSS-Team/MOSS-Transcribe-Diarize` @ `704aa4a9c304e8520be88901e0d1960158ef5b15`, Apache-2.0 | 8092 on the docker bridge |

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
   - Hopper (H100/H200) or a 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main`; on 48 GB also set `LLM_MAX_LEN=32768`. Not measured.
   - No suitable GPU: I can still bring transcripts and point `DECOSA_LLM_URL` at a model server my institution approves. Ask me before doing that.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories, and show me the commands first.
3. Confirm about 70 GB of free disk.

## 2. Get the images

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`.
If a pull fails, build from source once the `decosa-api` source is published: in the clone, `docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`, and `docker compose build llm` from its compose file.

## 3. Write the compose file

Create `~/decosa-themes/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.80
DECOSA_SIGNER_NAME="<who signs these records, e.g. the lab's name>"
```

Create `~/decosa-themes/docker-compose.yml` with exactly these services:

```yaml
name: decosa-themes
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }
    extra_hosts: ["host.docker.internal:host-gateway"]
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"                # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_THEMES_SAMPLES_ONLY: "0"           # take my own recordings and transcripts
      DECOSA_THEMES_EMBED_DIR: /data/models/bge-small-en-v1.5
      DECOSA_DIARIZE_URL: ""                    # step 4 sets this for recordings
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "2000000"       # a 40-interview study codes a few thousand passages
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

Use the named volume `decosa-data` exactly as written; a host bind mount owned by root makes the API fail on `/data/keys.sqlite`.
For "your own model", put the ONNX export of `BAAI/bge-small-en-v1.5` (`model.onnx` and `tokenizer.json`) in the volume at `/data/models/bge-small-en-v1.5`; without it the API falls back to hashed word features and says so on the model card.
Run `docker compose up -d`, then poll `docker compose ps` until both services are healthy (5-10 minutes the first time). `curl -s localhost:8445/healthz` should show `"llm": true`.

## 4. Recordings (diarize service)

Skip this if I only bring transcripts (plain text with `Name:` labels, or WebVTT/SRT).
The `diarize` service (MOSS-Transcribe-Diarize 0.9B, Apache-2.0) has no published image yet. Run `services/diarize` from the decosa-api source on this box (install with the `uv` commands at the top of `services/diarize/requirements.txt`, then `DIARIZE_HOST=172.17.0.1 DIARIZE_PORT=8092 DIARIZE_DEVICE=cuda:0 .venv-diarize/bin/python services/diarize/server.py`; `172.17.0.1` is the docker bridge address from `ip -4 addr show docker0`, reachable from containers and not from the LAN). Set `DECOSA_DIARIZE_URL: http://host.docker.internal:8092` on `api` and `docker compose up -d api`. The 0.9B model needs a few GB beside the LLM; if it doesn't fit, lower `LLM_GPU_UTIL` to 0.70. That fit is an estimate, not measured.

## 5. Smoke test

```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"interview-themes"}' | jq -r .token)
curl -s $API/themes/info | jq '{samples_only, limits}'
printf 'Interviewer: What is a normal day like for you?\nMaria: We start at seven with a check-in, then I spend most of the morning on patient calls, which I find draining but meaningful.\nInterviewer: What helps?\nMaria: Honestly the routine helps, and talking with my colleagues at lunch keeps me going when the calls are hard.\n' > /tmp/iv1.txt
jq -Rs '{text: ., title: "Test interview"}' /tmp/iv1.txt | curl -sN $API/themes/transcript -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @- \
  | sed -n 's/^data: //p' | jq -c 'select(.type=="result") | .interview' > /tmp/iv1.json
jq '{turns: (.turns | length), roles: [.turns[].role] | unique}' /tmp/iv1.json
```

Pass if `samples_only` is `false`, the interview has 4 turns, and the roles are `interviewer` and `participant` (Maria). Then propose a codebook (it needs at least a few participant passages, so this tiny interview may be refused with a clear message; that is a pass too):

```bash
jq '{interviews: [.], question: "What helps staff cope with hard calls?", n_codes: [3, 5]}' /tmp/iv1.json \
  | curl -sN $API/themes/codebook -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @- | sed -n 's/^data: //p' | jq -c '{type, message, codes: (.codebook.codes // [] | map(.name))}'
```

Then verify a signed record from any `result` event: `jq '{record}' result.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'` must say `ok: true`.

## 6. The signing key

1. `curl -s localhost:8445/attest/signing-key` shows the public key. Show me the `pubkey`, and tell me to back up the `decosa-data` volume (it holds `/data/attest/ed25519.pem`).
2. Receipts from this box say `status: "attested"`: signed by its own key. They show nothing was changed after signing and which key signed; they do not prove the transcript or the codes are right.

## 7. Point the app at the local API

The console at decosa.ai runs in my browser; my study never leaves it except to go to the API I choose. To use this box, run the site from source with `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` (`npm ci && npm run build && npm start`), or call the API from my own scripts with the same bearer token: `POST /themes/transcribe` (recording as the body), `/themes/transcript`, `/themes/codebook`, `/themes/codebook/split`, `/themes/analyze`, `/themes/train` and `/themes/export`.

## 8. Optional: provide compute to the network

Off by default, and never for interview data: this box runs the tool for me only. If I later want to offer spare GPU time to the Decosa network for public workloads, that is a separate setup on a separate machine; see decosa.ai/provide.

## Final summary

Tell me: which services are healthy, the signing key's `pubkey`, where the compose file is, whether recordings are enabled (step 4), and repeat the two reminders from the top.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A thematic analysis tool for interview studies: speaker-separated transcripts with the interviewer left out, a codebook of 8 to 12 codes you edit and approve, every passage coded against it, and themes with participant counts and quotes checked word for word and playable at their timestamp.
Who it's for
PhD students, academic researchers, UX researchers and program evaluators with 10 to 60 recorded interviews.
Where it runs
Self-host, or a confidential enclave on request, for interviews under an ethics or IRB protocol; the hosted demo takes the public-domain sample interviews only
Key numbers

On 1,000 passages of published human-coded interviews (28 codes) the coder agreed with the published coding at Cohen's κ 0.52 (0.49 on the test split); a blind second coder reached 0.62 on a 120-passage sample. One dataset, one domain.

  • 0.05% (22 of 41,624) Words credited to the wrong speaker (12 public-domain interviews) (test split, n = 41624)
  • 0.49 Agreement with published human coding (Cohen's κ, 28 codes, test split) (test split)
  • 15 of 15 Quotes passing the word-for-word and speaker check (test split, n = 15)
  • 69.4 s Median end-to-end run, hosted (QA sweep 2026-09-30)
All results, datasets and caveats
Models
MOSS-Transcribe-Diarize 0.9B (speech and speakers) · ECAPA-TDNN voice check (CPU) · Qwen3.8-27B (codebook, coding, themes) · bge-small-en-v1.5 (your own model, CPU)
Where
Self-host, or a confidential enclave on request, for interviews under an ethics or IRB protocol; the hosted demo takes the public-domain sample interviews only
Checks
Every quote re-checked word for word against the transcript and against its speaker; participant counts computed in code; interviewer turns excluded; signed record of every model call
Output
Structured data · Signed record or verdict
Data
Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

Does it do the thematic analysis for me?

No. It proposes codes and themes; you rename, merge, split, delete and approve the codebook before anything is coded, and the themes are candidates for you to judge. The methods paragraph says so, and lists every edit you made.

How accurate is the coding?

On 1,000 passages from a published, human-coded interview study with 28 codes, it agreed with the published coding at Cohen's κ 0.52 (0.49 on the test split). A blind second coder agreed with the same coding at 0.62 on a 120-passage sample. With code names only, agreement fell to 0.29: good definitions matter.

Can I use it for interviews under an IRB or ethics protocol?

Run it on your own hardware, or ask for a confidential enclave; the hosted demo takes public-domain sample interviews only. It doesn't replace your ethics approval, but the page's data-flow statement and the signed record are written for your application and your file.

Are the quotes real?

Every quote is re-checked word for word against the transcript and against the participant it is credited to, and plays from its timestamp in the recording. A quote that fails shows why. Quotes come from the transcript, so a speech-recognition slip can still be in one: play it.

Can I open the results in my QDA software?

Yes: export a REFI-QDA project (.qdpx) or codebook (.qdc), which most QDA programs import, plus CSVs of the codebook and the coded segments.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Interview themes

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.