Turn a family interview into a film
A trailer, a film and a book in her own voice over her own photos, with English under her words and every quote traced to the second.
- For
- Grandchildren and adult children keeping a grandparent's stories
- Instead of
- The recording sits on a phone: nobody has the hours to find the best lines, time them, translate them and cut a film.
- Time per task41 stypical (median) on the sample; slowest 1 in 20: 71 s
- Cost per task~$0.80 per 100 filmsmeasured, at list price
- Accuracy46 / 46Quotes shown that are hers, at the right time (dev set)All results and caveats
Make the family film
Try it in one click
Watch Rosa's stories become a film
A 3½-minute interview in Italian and eight old photos go in. About a minute later: her chapters, her best lines (each one traced to the second she said it), English subtitles under her own words, and a trailer in her voice.
Rosa is fictional and her voice is a designed synthetic voice; the photos are public-domain pictures standing in for a family album.
Your family
Before you record
Studio writes an interview guide in her language, chapter by chapter, with gentle follow-ups. Then she says yes in her own words, and you record.
Watch
Recorded sessions are for the live audio tools. For the studio, the gallery under Try it shows finished renders.
Get an API key
- Call the family interview film API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on CPU for the edit; speech recognition, translation and the photo-back reader share one GPU (no video model).
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- family-film
Use the hosted API
# Decosa family interview film: use the hosted API
You are wiring Decosa's family interview film into this project. It takes a recorded interview with a parent or
grandparent (in her language), her consent in her own words and her old photos, and returns chapters of her life, her
best lines traced to the second she said them, English subtitles under her own words with a meaning check, the photo
backs read, and a trailer, a film and a book with QR codes that play her quotes. Her real voice and real photos only: no
voice clone, and no model changes a picture. Use only what is listed below. If you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`. Health check: `GET https://api.decosa.ai/healthz`.
- Record only someone who agrees. Her consent recording comes first; a refusal or a hesitation is refused (422).
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool's page. Keep it in `DECOSA_API_KEY`, never in code.
Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "studio"}` returns `{"token", ...}` (30 minutes).
3. Every tool call works inside a Studio project: `POST /studio/projects {"name": "..."}` returns `project.id` and a
`project_key`. Send the key as `?key=` on every project URL (the API's CORS allows only two headers). Keep it private:
the project key alone opens the project.
## Endpoints (all under a project; `?key=<project_key>`)
- `GET /studio/family/info`: languages, limits, models, rules (no token).
- `POST /studio/projects/{pid}/family/consent-text {"storyteller", "language"}`: the consent sentence to read to her.
- `POST /studio/projects/{pid}/family {"storyteller", "language", "interviewer"?, "consent_b64"}` → 201 the film with
`revocation.path`: her own link. Give it to her: it takes her consent back and deletes the project.
- `POST /studio/projects/{pid}/family/sample`: the synthetic Rosa sample (Italian, 207 s, public-domain photos).
- `POST /studio/family/{fid}/interview {"audio_b64"}` (up to 60 MB, 2 hours); `POST /studio/family/{fid}/photos
{"photos": [{"b64", "caption"?, "year"?, "place"?, "back_b64"?}]}` (up to 12 a call, 40 a film, 12 MB each). A picture is read as the back of a photo only
when no face is found in it.
- `GET /studio/family/{fid}`: poll every 5 s. `stage` runs consented, listening, speakers, words, chapters, tracing,
subtitles, photos, ready. Then `chapters[].quotes[]` (`text`, `en`, `start`, `end`, `traced`, `meaning`),
`films.trailer` / `films.film` (`status`, `url` relative to the base, `credential`), `tracing`, `speakers`, `receipts`.
- `POST /studio/family/{fid}/edit` (rename a chapter, leave a quote out, move a photo), then `POST /studio/family/{fid}/film`
to cut again. `GET /studio/family/{fid}/qr/{qid}.svg`: the QR code for a quote.
- `POST /studio/family/{fid}/share {"name"}` / `unshare`: private links for named people. `DELETE /studio/family/{fid}`.
## Example (Python, `pip install httpx`)
```python
import httpx, os, time
API = "https://api.decosa.ai"; H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
p = httpx.post(f"{API}/studio/projects", headers=H, json={"name": "Nonna Rosa"}).raise_for_status().json()
pid, key = p["project"]["id"], p["project_key"]
fid = httpx.post(f"{API}/studio/projects/{pid}/family/sample", params={"key": key}, headers=H, timeout=120).raise_for_status().json()["id"]
while True:
f = httpx.get(f"{API}/studio/family/{fid}", params={"key": key}, headers=H).json()
if f["status"] == "failed" or f["films"].get("trailer", {}).get("status") == "done":
break
time.sleep(5)
for ch in f["chapters"]:
print(ch["title_en"], [(q["en"], q["start"]) for q in ch["quotes"] if q["traced"]])
print(API + f["films"]["trailer"]["url"])
```
## Honest limits
- Measured on 5 synthetic two-voice interviews (it, es, pt, fr, de): all 46 shown quotes were hers at the right time;
real recordings with more relatives or room noise are not measured.
- The meaning check flags about half the English lines and most flags are nitpicks: read the English yourself.
- Quotes copy the speech recogniser's words, so a small slip ("né" for "née") can reach the film.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa family interview film: run it yourself (containers)
You are setting up Decosa's family interview film on this machine: speech recognition (Qwen3-ASR-1.7B), speaker
separation (ECAPA-TDNN on CPU), chapters and quotes (Qwen3.8-27B), subtitles (Hy-MT2-7B) with a meaning check, photo
backs (PaddleOCR-VL-1.6, only pictures with no face), and an FFmpeg film cut with a C2PA credential. Recordings and photos
stay here; nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images or models.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/family-film.zip (102 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py family-film` (the api image carries the same bundle under /app/rehearsal/family-film/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py family-film --bundle family-film.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "a refusal in her own words is refused", "the sample is ready", "at least 4 chapters of her life"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Licences first
All models are Apache-2.0 (Qwen3.8-27B, Qwen3-ASR-1.7B, Hy-MT2-7B, PaddleOCR-VL-1.6, speechbrain ECAPA-TDNN) or MIT (the
Ultra-Light face detector; the ACE-Step score cue). FFmpeg is LGPL/GPL. Record only someone who agrees; the consent
ledger keeps a voiceprint of her consent recording, which some laws treat as biometric data.
## Steps
1. Docker and the NVIDIA container toolkit: if `docker compose version` or
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi` fails, install them from the official
instructions. Plan on one card with about 45 GB free (speech, translation, reader, text model), or remote servers.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the
`llm`, `mt`, `reader`, `speech` and `api` services. For `api` set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LANG_SPEECH_URL`,
`DECOSA_LANG_MT_URL`, `DECOSA_DOCREADER_PARSER_URL` and `DECOSA_CONSENT_SPK_MODEL` (the ECAPA ONNX file from
`scripts/consent/export_ecapa_onnx.py`); use named volumes; bind every port to 127.0.0.1.
3. `docker compose pull && docker compose up -d`; wait for the health checks.
4. Smoke test: a token from `POST /demo/session {"vertical":"studio"}`, a project from `POST /studio/projects`, then
`POST /studio/projects/{pid}/family/sample?key=...` and poll `GET /studio/family/{fid}?key=...` until
`films.trailer.status` is `done`: expect 6-8 chapters, `tracing.rate` 1.0, `speakers.ok` true and a `c2pa` credential.
5. Report back: the signing key id (`GET /attest/signing-key`) and the time to the first trailer.
Off. Family recordings never go to the network.
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3-ASR-1.7B (language pack speech service) needs a GPU.
- GeForce RTX 4090Doesn't fit
Needs about 43.1 GB of GPU memory at the smallest settings; 24 GB available.
- GeForce RTX 5090Doesn't fit
Needs about 51.1 GB of GPU memory at the smallest settings; 32 GB available.
- 2x GeForce RTX 5090lite tierRuns with a smaller tier
The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits with changes.
- L40SDoesn't fit
Needs about 56.7 GB of GPU memory at the smallest settings; 48 GB available.
- H100 80 GB (SXM)lite tierRuns with a smaller tier
The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits with changes.
- RTX PRO 6000 Blackwell 96 GBlite tierRuns with a smaller tier
The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits.
- 2x RTX PRO 6000 Blackwell 96 GBlite tierRuns with a smaller tier
The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits.
- Apple M3 Ultra (Mac Studio), 96 GBCan't tell
Memory not known for Qwen3-ASR-1.7B (language pack speech service) has no mapped Apple Silicon build; Hy-MT2-7B (the language-pack block) has no mapped Apple Silicon build; Decosa document reader (Docling layout + PaddleOCR-VL-1.6) has no mapped Apple Silicon build.
- Apple M5 Max, 64 GBCan't tell
Memory not known for Qwen3-ASR-1.7B (language pack speech service) has no mapped Apple Silicon build; Hy-MT2-7B (the language-pack block) has no mapped Apple Silicon build; Decosa document reader (Docling layout + PaddleOCR-VL-1.6) has no mapped Apple Silicon build.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API, the speech model and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with both models loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "asr": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"family-film"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py family-film
Download the mock-data bundle (102 KB, 8 checks)expected.json
A refusal in the storyteller's own words must stop the project. Then the synthetic Rosa sample (a 207 s Italian interview with public-domain photos, bundled on the server) must reach ready with at least 4 chapters, every shown quote traced to the second she said it, her voice told from the interviewer's, and a trailer with a C2PA credential, every model call receipted. About 1-2 minutes.
What the rehearsal checks
- a refusal in her own words is refused
- the sample is ready
- at least 4 chapters of her life
- every shown quote is traced to the second she said it
- her voice is told from the interviewer's, with her consent recording as the reference
- the trailer renders
- the trailer carries a C2PA credential
- every model call has a signed receipt
Licence: inputs/refusal-it.wav: a synthetic refusal made with VoxCPM2 voice design (openbmb/VoxCPM2, Apache-2.0) from a text description; no person was recorded or cloned. The Rosa sample on the server: a synthetic interview (same method) and Library of Congress photos with no known restrictions.
Prompt for your coding agent
# Decosa family interview film: run it yourself (containers)
You are setting up Decosa's family interview film on this machine: speech recognition (Qwen3-ASR-1.7B), speaker
separation (ECAPA-TDNN on CPU), chapters and quotes (Qwen3.8-27B), subtitles (Hy-MT2-7B) with a meaning check, photo
backs (PaddleOCR-VL-1.6, only pictures with no face), and an FFmpeg film cut with a C2PA credential. Recordings and photos
stay here; nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images or models.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/family-film.zip (102 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py family-film` (the api image carries the same bundle under /app/rehearsal/family-film/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py family-film --bundle family-film.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "a refusal in her own words is refused", "the sample is ready", "at least 4 chapters of her life"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Licences first
All models are Apache-2.0 (Qwen3.8-27B, Qwen3-ASR-1.7B, Hy-MT2-7B, PaddleOCR-VL-1.6, speechbrain ECAPA-TDNN) or MIT (the
Ultra-Light face detector; the ACE-Step score cue). FFmpeg is LGPL/GPL. Record only someone who agrees; the consent
ledger keeps a voiceprint of her consent recording, which some laws treat as biometric data.
## Steps
1. Docker and the NVIDIA container toolkit: if `docker compose version` or
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi` fails, install them from the official
instructions. Plan on one card with about 45 GB free (speech, translation, reader, text model), or remote servers.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the
`llm`, `mt`, `reader`, `speech` and `api` services. For `api` set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LANG_SPEECH_URL`,
`DECOSA_LANG_MT_URL`, `DECOSA_DOCREADER_PARSER_URL` and `DECOSA_CONSENT_SPK_MODEL` (the ECAPA ONNX file from
`scripts/consent/export_ecapa_onnx.py`); use named volumes; bind every port to 127.0.0.1.
3. `docker compose pull && docker compose up -d`; wait for the health checks.
4. Smoke test: a token from `POST /demo/session {"vertical":"studio"}`, a project from `POST /studio/projects`, then
`POST /studio/projects/{pid}/family/sample?key=...` and poll `GET /studio/family/{fid}?key=...` until
`films.trailer.status` is `done`: expect 6-8 chapters, `tracing.rate` 1.0, `speakers.ok` true and a `c2pa` credential.
5. Report back: the signing key id (`GET /attest/signing-key`) and the time to the first trailer.
Off. Family recordings never go to the network.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Doesn't fitFamily interview film on GeForce RTX 5090
Needs about 51.1 GB of GPU memory at the smallest settings; 32 GB available.
Lite · chapters, quotes, subtitles and the film, no photo reading: what changesuses estimates
- Needs about 51.1 GB of GPU memory at the smallest settings; 32 GB available.
Memory per component
- Speech recognition in her language, one call...: Qwen3-ASR-1.7B (language pack speech service). ~5.1 GB, weights 3.4 GB (estimate). Estimate: 1.7B parameters at 2 bytes (BF16) per weight is about 3.4 GB, plus 20% working memory and 1 GB of runtime. Not measured.
- Who is speaking: ECAPA-TDNN (speechbrain/spkrec-ecapa-voxceleb, ONNX export). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Chapters of her life and her best lines: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
- English subtitles under her own words, senten...: Hy-MT2-7B (the language-pack block). ~18 GB (from stack.json). vram_gb 18 in stack.json.
- A quiet score under the film: ACE-Step 1.5 cue (music-gen-cleared library). CPU. Runs on CPU (vram_gb 0 in stack.json).
- The film: decosa-api family_film module + FFmpeg + c2pa-python. CPU. Runs on CPU (vram_gb 0 in stack.json).
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Family interview film, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Family interview film on my hardware Fetch https://decosa.ai/prompts/family-film-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=family-film) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Lite · chapters, quotes, subtitles and the film, no photo reading (lite). Fit check: doesn't fit; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Speech recognition in her language, one call...: Qwen3-ASR-1.7B (language pack speech service) (Qwen/Qwen3-ASR-1.7B), 5.1 GB - Who is speaking: ECAPA-TDNN (speechbrain/spkrec-ecapa-voxceleb, ONNX export) (speechbrain/spkrec-ecapa-voxceleb), CPU - Chapters of her life and her best lines: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB - English subtitles under her own words, senten...: Hy-MT2-7B (the language-pack block) (tencent/Hy-MT2-7B), 18 GB - A quiet score under the film: ACE-Step 1.5 cue (music-gen-cleared library) (ACE-Step/Ace-Step1.5), CPU - The film: decosa-api family_film module + FFmpeg + c2pa-python, CPU Warning: the fit check says this tier does not fit: Needs about 51.1 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further. During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/family-film-assemble.md
Get an API key
- Call the family interview film API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on CPU for the edit; speech recognition, translation and the photo-back reader share one GPU (no video model).
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Record Grandma telling her stories in her language; get a trailer, a film and a book in her own voice, subtitled in English.
A grandchild records a parent or grandparent on a phone, in their language, and adds old photos. Studio tells her voice from the interviewer's, finds the chapters of her life and her best lines, re-finds every line at the second she said it, puts English under her own words with a meaning check, reads the handwriting on the backs of the photos, and cuts a trailer, a film and a book whose QR codes play her quotes. Her real voice and her real photos only: no voice clone, and no model draws or changes a picture.
- Deployment
- Hosted or self-host
- Regulatory
- Not legal advice; written 29 Sep 2026. Consent: the storyteller records her consent in her own words before anything else happens; the meaning of what she said is checked (a refusal or a hesitation stops the project), and the consent is kept in the consent ledger with her own revocation link, which deletes the project. Her voice: the ledger keeps a voiceprint of her consent recording so the film can tell her voice from the interviewer's; some laws treat voiceprints as biometric data (for example the Illinois Biometric Information Privacy Act, 740 ILCS 14, which asks for a written release), so self-host or get the consent your law requires. Photos: fronts of photos are shown, panned and captioned only; nothing is generated from them. Children: children are paused in Studio. A family photo in which anyone looks under 18, or whose caption or writing says so, is refused and not kept, and Studio never makes a child the subject of a generated picture or video. A family mode, where a parent verifies and consents for their children, is planned. Every film carries a C2PA credential; nothing is used for training.
Text description
Her interview (a phone recording in her language), her consent in her own words and her old photos go in. Voice activity and ECAPA-TDNN (Apache-2.0) speaker embeddings tell her voice from the interviewer's, with her consent recording as the reference; Qwen3-ASR-1.7B (Apache-2.0) hears the words. A face detector (UltraFace, MIT) checks each picture: only pictures with no face (the backs of photos) go to the document reader (PaddleOCR-VL, Apache-2.0). Before any photo is kept, Qwen3.8-27B looks at it, and its caption is read, for anyone under 18 (children are paused). Qwen3.8-27B (Apache-2.0) finds chapters and her best lines, and code re-finds each line at its timestamp. Hy-MT2-7B (Apache-2.0) puts English under her words, with a meaning check on every line. FFmpeg on CPU cuts a trailer, a film and a book with QR codes: her real voice over her real photos, maps and dates, and a library music cue. Private family links; her own link takes the consent back. Receipts on every model call and a C2PA credential on every film. Self-hosted, everything stays on your machine.
At a glance
- Data retention
- The interview, photos, transcript and films stay in your private Studio project until you delete it or she takes her consent back with her link (which deletes everything). Her consent entry expires after three years. Logs hold ids, counts and hashes, never her words.
- What leaves the box
- On the hosted route, speech, translation, photo backs and the text model all run on Decosa's hosted service. Self-hosted, nothing leaves your machine.
- Her voice, her photos
- No voice clone and no generated faces: the film uses her real voice and shows her real photos. Only pictures with no face in them (the backs of photos) are read by a model.
- Cost per film
- A few cents or less per film in the eval, at list prices for the text model and GPU time.
- Typical time
- Chapters and a first trailer arrive within about a minute of upload for a short interview, longer for a longer one (measured).
- Outputs
- A trailer and a film (MP4, 16:9) with English subtitles under her words and a C2PA credential; subtitle files; a book page per chapter with a QR code per quote that plays her saying it; private share links you can turn off.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
chapters, quotes, subtitles and the film, no photo reading
Everything except reading the backs of photos: add places and years by hand instead.
- Models
- Qwen3-ASR-1.7B (language pack speech service)
- ECAPA-TDNN (speechbrain/spkrec-ecapa-voxceleb, ONNX export)
- Qwen3.8-27B (NVFP4)
- Hy-MT2-7B (the language-pack block)
- ACE-Step 1.5 cue (music-gen-cleared library)
- decosa-api family_film module + FFmpeg + c2pa-python
- Hardware
- One GPU for speech recognition and translation (Hy-MT2 served at 18 GB), the text model (about 20 GB, or a remote server), CPU for the edit
- Quality evidence
- Shown quotes inside her own turn, at the right time (5 synthetic interviews)46 / 46decosa-api docs/evals/family-film.md, 2026-09-29
- Speaker labels right79 / 80 segmentsdecosa-api docs/evals/family-film.md, 2026-09-29
- Latency
- measured: a trailer about a minute after upload on our server
- Verification
- Proof: partialSelf-host only
- In the hosted demo
Standard
the hosted demo, with photo backs read
Adds the face guard and the document reader, so handwriting on the backs of photos dates and places each picture.
- Models
- Qwen3-ASR-1.7B (language pack speech service)
- ECAPA-TDNN (speechbrain/spkrec-ecapa-voxceleb, ONNX export)
- Qwen3.8-27B (NVFP4)
- Hy-MT2-7B (the language-pack block)
- Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)
- Decosa document reader (Docling layout + PaddleOCR-VL-1.6)
- ACE-Step 1.5 cue (music-gen-cleared library)
- decosa-api family_film module + FFmpeg + c2pa-python
- Hardware
- 1x RTX PRO 6000 Blackwell 96 GB (shared) for speech, translation and the reader; the text model on a second card or remote; CPU for the edit
- Quality evidence
- Shown quotes inside her own turn, at the right time (5 synthetic interviews)46 / 46decosa-api docs/evals/family-film.md, 2026-09-29
- Photo backs read (synthetic handwriting)3 / 3decosa-api docs/evals/family-film.md, 2026-09-29
- Meaning check: lines flagged / clear mistranslations among them27 of 46 flagged; 2 clear mistranslationsdecosa-api docs/evals/family-film.md, 2026-09-29; the builder read every flag
- Latency
- measured on our server: chapters and a trailer within about a minute of upload
- Verification
- Proof: partial
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Speech recognition in her language, one call per speech segment (so every word keeps its time)Qwen3-ASR-1.7B (language pack speech service)Qwen/Qwen3-ASR-1.7B on Hugging Face (opens in a new tab) 1.7BProof: partial | LiteStandard | 1.7B | Proof: partial | |
| ||||
Who is speaking: voice activity by energy, then ECAPA-TDNN voice embeddings per segment in two clusters; the cluster nearest her consent recording is hersECAPA-TDNN (speechbrain/spkrec-ecapa-voxceleb, ONNX export)speechbrain/spkrec-ecapa-voxceleb on Hugging Face (opens in a new tab) 0 GBNo proof yet | LiteStandard | 0 GB | No proof yet | |
| ||||
Chapters of her life and her best lines (copied exactly from the transcript, then re-found in it by code), film titles, and the translation fallbackQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | LiteStandard | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
English subtitles under her own words, sentence by sentence, then a meaning check (back-translation compared with the source) on every lineHy-MT2-7B (the language-pack block)tencent/Hy-MT2-7B on Hugging Face (opens in a new tab) 7.5B · 18 GBProof: partial | LiteStandard | 7.5B · 18 GB | Proof: partial | |
| ||||
Face detection only (boxes): a picture is read as the back of a photo only when no face is found; photo pans drift toward facesUltra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320) 0 GBNo proof yet | Standard | 0 GB | No proof yet | |
| ||||
Reads the handwriting on the backs of photos (place, year, names) to date and place each pictureDecosa document reader (Docling layout + PaddleOCR-VL-1.6)PaddlePaddle/PaddleOCR-VL-1.6 on Hugging Face (opens in a new tab) Proof: partial | Standard | n/a | Proof: partial | |
| ||||
A quiet score under the film: a pre-rendered cue from the cleared music library (no model runs per film)ACE-Step 1.5 cue (music-gen-cleared library)ACE-Step/Ace-Step1.5 on Hugging Face (opens in a new tab) 2.39B DiT + 1.85B LM · 0 GBProof: partial | LiteStandard | 2.39B DiT + 1.85B LM · 0 GB | Proof: partial | |
| ||||
The film (CPU): her voice over her photos with slow pans, chapter cards, maps (Natural Earth) and dates, subtitles, the trailer and the book with a QR code per quote; C2PA credential per filedecosa-api family_film module + FFmpeg + c2pa-python 0 GBProof: partial | LiteStandard | 0 GB | Proof: partial | |
| ||||
Tools, services and hardware
Tools
- Natural Earth (opens in a new tab)Public domain
Coastlines and places for the journey maps.
- Library of Congress, Prints and Photographs (no known restrictions) (opens in a new tab)Public domain / no known restrictions
The sample's old photos.
- FFmpeg with libass (opens in a new tab)LGPL-2.1+ / GPL for some builds
Cuts the trailer and the film, burns in subtitles, mixes her voice and the score.
- c2pa-python (opens in a new tab)MIT OR Apache-2.0
The C2PA content credential on every film.
Example of a law that treats voiceprints as biometric identifiers. Unverified: the page refused our automated fetch on 29 Sep 2026, so check the current text.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /studio/family/info; POST /studio/projects/{pid}/family (her consent recording), /family/sample; POST /studio/family/{fid}/interview, /photos, /edit, /film, /share, /unshare, /studio/family/revoke; GET /studio/family/{fid}, /media/{name}, /qr/{qid}.svg. Needs FFmpeg, fonts and c2pa-python in the image.
- decosa-lang (MT and speech):8491
Hy-MT2-7B translation (:8491) and the Qwen3-ASR speech service (:8492).
- decosa document reader:8497
Reads photo backs (only pictures with no face).
- vLLM:8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4, behind our gateway (hosted) or called directly (self-host).
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB, shared Fits
Measured on our server 2026-09-29: speech recognition, translation and the document reader on GPU0 beside other services; the text model on GPU1; the film cut on CPU.
- CPU only Does not fit
Speech recognition and translation need a GPU (or remote servers); the speaker model and the edit run on CPU.
Latency per lane
- Chapters and quotes ready after upload (1-3.5 minute interview)29.9 s
Measuredmeasured on our server 2026-09-29: 26-55 s over 5 interviews (p50 29.9 s)
- First trailer after upload40.8 s
Measuredmeasured on our server 2026-09-29: 37-71 s over 5 interviews (p50 40.8 s)
- Full film cut (181 s film from a 207 s interview)33.0 s
Measuredmeasured on our server 2026-09-29: 33 s of CPU
Notes
- One film style for now: her voice over her photos with slow pans, chapter cards, maps and dates.
- Speaker separation was measured on two voices (her and one interviewer); a table of 4-6 relatives is not measured.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa family interview film on this machine
You are setting up a tool that turns a recorded family interview into a film: it tells the storyteller's voice from the
interviewer's, hears her words in her language, finds the chapters of her life and her best lines (each re-found at the
second she said it), puts English subtitles under her own words with a meaning check, reads the handwriting on the backs
of photos, and cuts a trailer, a film and a book with QR codes that play her quotes. Her real voice and her real photos
only: no voice clone, and no model draws or changes a picture. Work step by step, show me each command before you run
anything with `sudo`, and stop to ask if a check fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/family-film.zip (102 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py family-film` (the api image carries the same bundle under /app/rehearsal/family-film/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py family-film --bundle family-film.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "a refusal in her own words is refused", "the sample is ready", "at least 4 chapters of her life"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- Models: Qwen3.8-27B (chapters, quotes, titles; Apache-2.0), Qwen3-ASR-1.7B (speech; Apache-2.0), Hy-MT2-7B
(translation; Apache-2.0), PaddleOCR-VL-1.6 (photo backs; Apache-2.0), ECAPA-TDNN from
`speechbrain/spkrec-ecapa-voxceleb` (speakers, CPU, ONNX; Apache-2.0), and the Ultra-Light face detector (MIT, bundled
in the api image). The score is a pre-rendered ACE-Step 1.5 cue (MIT) shipped with the api.
- Record only someone who agrees. The api asks for her consent in her own words first, checks what she means (a
refusal or a hesitation stops the project), keeps it in the consent ledger with her own revocation link, and deletes
the project when she uses it. The ledger keeps a voiceprint of her consent recording; some laws treat voiceprints as
biometric data, so get the consent your law asks for.
- The fronts of photos never go to any model. A picture is read only when the face detector finds no face in it.
- Bind every port to 127.0.0.1. Logs carry ids, counts and hashes, never her words.
## 1. Check the machine
1. `nvidia-smi`: one card with about 45 GB free runs speech (about 5 GB while loaded), translation (vLLM at
`--gpu-memory-utilization 0.18`, about 18 GB), the photo-back reader (0.04) and, with room to spare, the text model
(about 20 GB, NVFP4 needs a Blackwell card; use a server you already run otherwise). We measured it on a shared RTX
PRO 6000 96 GB.
2. `docker --version`, `docker compose version`; if Docker or the NVIDIA container toolkit is missing, ask me, then
install them from the official repositories and check `docker run --rm --gpus all ubuntu nvidia-smi`.
3. Disk: about 60 GB (Qwen 20 GB, Hy-MT2 15 GB, ASR 4 GB, reader 2 GB, images).
## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
`git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out the newest release that contains
`decosa_api/studio_core/family_film/` (`main` until one does), and `docker build -f docker/api/Dockerfile -t decosa-api:family .`.
The image carries FFmpeg (libass), fonts, onnxruntime and c2pa-python.
- The speaker model: in a throwaway venv with `speechbrain` and CPU torch, run
`python scripts/consent/export_ecapa_onnx.py ./ecapa` from the same checkout. It downloads
`speechbrain/spkrec-ecapa-voxceleb`, writes `ecapa/ecapa.onnx` once and checks parity; the api needs only onnxruntime.
- `vllm/vllm-openai:v0.29.0` for the text model (`nvidia/Qwen3.8-27B-NVFP4`, revision `482ca0f3832238542f8f5295dde86b5f22711d80`),
translation (`tencent/Hy-MT2-7B`) and the reader (`PaddlePaddle/PaddleOCR-VL-1.6`, revision `c5630ab`).
- The speech service: `services/lang/server.py` from the checkout in its own Python 3.12 venv (we run torch 2.8.0+cu128,
`qwen-asr` 0.0.6, fastapi, uvicorn, soundfile and librosa 1.0.0), with `Qwen/Qwen3-ASR-1.7B` (revision `7278e1e70fe206f11671096ffdd38061171dd6e5`)
in `LANG_MODELS/Qwen3-ASR-1.7B/`. Run it with `LANG_HOST=127.0.0.1 LANG_PORT=8492 LANG_DEVICE=cuda:0`.
## 3. docker-compose.yml
Write this in `~/decosa/family/`. `network_mode: host` lets the api reach the speech service on 127.0.0.1:8492.
```yaml
services:
llm:
image: vllm/vllm-openai:v0.29.0
command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--host", "127.0.0.1", "--port", "8114",
"--max-model-len", "32768", "--enable-prefix-caching", "--gpu-memory-utilization", "0.30"]
network_mode: host
volumes: ["hf-cache:/root/.cache/huggingface"]
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
healthcheck: { test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8114/v1/models', timeout=4)"], interval: 30s, retries: 30 }
mt:
image: vllm/vllm-openai:v0.29.0
command: ["--model", "tencent/Hy-MT2-7B", "--served-model-name", "hy-mt2-7b", "--host", "127.0.0.1", "--port", "8491",
"--model-impl", "transformers", "--enforce-eager", "--gpu-memory-utilization", "0.18", "--max-model-len", "8192"]
network_mode: host
volumes: ["hf-cache:/root/.cache/huggingface"]
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
reader:
image: vllm/vllm-openai:v0.29.0
command: ["--model", "PaddlePaddle/PaddleOCR-VL-1.6", "--revision", "c5630ab", "--served-model-name", "PaddlePaddle/PaddleOCR-VL-1.6",
"--trust-remote-code", "--host", "127.0.0.1", "--port", "8498", "--gpu-memory-utilization", "0.04", "--max-model-len", "8192"]
network_mode: host
volumes: ["hf-cache:/root/.cache/huggingface"]
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
api:
image: decosa-api:family
network_mode: host
environment:
DECOSA_HOST: 127.0.0.1
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://127.0.0.1:8114/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LANG_SPEECH_URL: http://127.0.0.1:8492
DECOSA_LANG_MT_URL: http://127.0.0.1:8491/v1
DECOSA_LANG_MT_MODEL: hy-mt2-7b
DECOSA_DOCREADER_PARSER_URL: http://127.0.0.1:8498/v1
DECOSA_CONSENT_SPK_MODEL: /models/ecapa/ecapa.onnx
DECOSA_STUDIO_GPU_USD_PER_H: "1.32" # what your GPU hour costs you, for the estimates
DECOSA_PROVENANCE_DIR: /provenance
volumes: ["decosa-data:/data", "decosa-provenance:/provenance", "./ecapa:/models/ecapa:ro"]
depends_on: { llm: { condition: service_healthy } }
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/studio/family/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
decosa-data:
decosa-provenance:
hf-cache:
```
Use named volumes for state (the api runs as uid 10001). `docker compose up -d`, start the speech service, and wait for
the health checks. Create the C2PA signing material once, then restart: `docker compose exec api python
scripts/provenance_devcert.py && docker compose restart api` (a development CA: public validators show the signature as
valid and the issuer as untrusted). Back up the box's Ed25519 key: `docker compose cp api:/data/attest ./attest-backup`.
## 4. Smoke test
```bash
B=http://127.0.0.1:8445
curl -s $B/studio/family/info | jq '{languages: [.languages[].id], models}'
T=$(curl -s -XPOST $B/demo/session -H 'content-type: application/json' -d '{"vertical":"studio"}' | jq -r .token)
P=$(curl -s -XPOST $B/studio/projects -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"name":"[TEST] family"}')
PID=$(echo "$P" | jq -r .project.id); K=$(echo "$P" | jq -r .project_key)
F=$(curl -s -XPOST "$B/studio/projects/$PID/family/sample?key=$K" -H "authorization: Bearer $T" | jq -r .id)
until curl -s "$B/studio/family/$F?key=$K" -H "authorization: Bearer $T" | jq -e '.films.trailer.status=="done" or .status=="failed"' >/dev/null; do sleep 10; done
curl -s "$B/studio/family/$F?key=$K" -H "authorization: Bearer $T" | jq '{status, chapters: (.chapters|length), tracing, speakers: .speakers.ok, meaning: {ok: .meaning.ok, check: .meaning.check, error: .meaning.error}, trailer: .films.trailer.credential}'
```
Expect `ready`, 6-8 chapters, `tracing.rate` 1.0, `speakers: true`, English on every quote, and `trailer: "c2pa"`,
about a minute after the sample starts on our server. Open the trailer URL from `.films.trailer.url` and watch it.
Report the times you measure.
## 5. Your own interview
In the app (or the API): create a project, record her consent (`POST /studio/projects/{pid}/family` with
`storyteller`, `language` and `consent_b64`, her own words on the consent text from `/family/consent-text`), then
upload the interview (`/studio/family/{fid}/interview`, up to 60 MB, 2 hours) and up to 40 photos. Keep the
revocation link the consent step returns: give it to her.
## 6. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`. Contract: `API_CONTRACT.md`, the Changes entry
"Studio: family interview film, music video starring you, our story film" (29 Sep 2026).What it does, in shortWho it's for, where it runs and the key results
Record Grandma telling her stories, in her language, and add her old photos. Studio hears who is speaking, finds the chapters of her life and her best lines (each traced to the second she said it), subtitles her in English under her own words, reads the backs of her photos, and cuts a trailer and a film: her real voice over her real photos, with maps, dates and a quiet score. A book view puts a QR code next to every quote that plays her saying it. Her consent is recorded in her own words and her own link takes it back. No model ever draws, animates or changes a photo, and there is no voice clone.
In short
Last reviewed
- What it is
- A family interview film from one phone recording: Studio finds the chapters and her best lines, subtitles her in English under her own words, and cuts a trailer, a film and a book in her real voice over her real photos.
- Who it's for
- Grandchildren and adult children who want a parent's or grandparent's stories kept, in their own voice and language.
- Where it runs
- Hosted, private by default; self-host for recordings that shouldn't leave the family
- Key numbers
- 46 / 46 Shown quotes inside her own turn, at the right time (dev (tuned on), n = 46)
- 44 / 46 Shown quotes that match her words (fuzzy 0.9) (dev (tuned on), n = 46)
- 2 / 48 Quotes dropped by the re-hearing check (dev (tuned on), n = 48)
- 40.8 s Median end-to-end run, hosted (QA sweep 2026-09-29)
How we tested itEnd-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 29 Sep 2026 · measured 29 Sep 2026: · p50 41 s · p95 71 s (5 runs) · ~$0.009 per run · 11 receipts
Loading the nightly status…
Self-host: not yet verified
Measured cost to run: about $0.80 per 100 films (hosted, 29 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
Known limits (5)
- Measured on 5 synthetic interviews with two voices each, voiced by an open voice-design model; real family recordings (noise, overlapping talk, more relatives) are not measured.
- The meaning check flags about half the lines, and most flags are nitpicks: read the English yourself for a language you know.
- Quotes copy the speech recogniser's words: a one-letter slip ("né" for "née") reached a shown quote.
- Chapters follow the order she told them, not the years, and can't be reordered yet.
- One film style.
Eval results, nightly checks and cost per run · eval not held out
Technical detailsModels, where it runs, labels, what it is built from
- Models
- Qwen3-ASR-1.7B · ECAPA-TDNN (speakers) · Qwen3.8-27B · Hy-MT2-7B · PaddleOCR-VL-1.6 (photo backs) · ACE-Step 1.5 (score)
- Where
- Hosted, private by default; self-host for recordings that shouldn't leave the family
- Checks
- Every quote re-heard at its timestamp; meaning check on every English line; receipts per model call; C2PA in every film
- Industry
- Creative and media · Film, TV and games
- Input
- Voice, live audio · Files and media
- Output
- Media · Notes, reports and drafts
- Data
- Personal data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Part of
- Decosa Studio: Film
- Runs in
- Decosa hosted · Self-host
- Built from
- Live speech to text · Studio render · Content credentials
Every result carries a signed record of which model produced it, so you can check it later. How that works
Questions people ask
How do I make a family interview film from a phone recording?
Record her consent in her own words, upload the interview and a few photos, and Studio finds chapters and quotes in about 30 s and a first trailer in about 40 s for a 1-minute interview (71 s for 3.5 minutes, measured).
Does it clone her voice or animate her photos?
No. The film uses her real voice and shows her real photos with slow pans. No model draws, animates or changes a picture, and only pictures with no face in them (the backs of photos) are read by a model.
How accurate are the quotes and subtitles?
In 5 synthetic interviews, all 46 shown quotes were inside her own turn at the right time, and 2 lines with recognition slips were dropped. The meaning check flagged 27 of 46 English lines, only 2 of them clear mistranslations, so read the English yourself.
Who can see the film?
Only people you name: each gets a private link you can turn off. Her own link takes her consent back and deletes the project.
Which languages work?
Measured on Italian, Spanish, Portuguese, French and German interviews, with English subtitles. Other languages the speech and translation models support may work but are not measured.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Family interview film
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…