170 · Creative and media · Film, TV and games · live
Family interview film
Eval results
Not held outRun 29 Sep 2026Eval write-up (decosa-api, access required)
- Shown quotes inside her own turn, at the right time46 / 46dev (tuned on)n = 465 languages; a quote counts when it overlaps her true turn by at least 80% of its span.
- Shown quotes that match her words (fuzzy 0.9)44 / 46dev (tuned on)n = 46Both misses write the year as digits where the script spells it out; one also has "né" for "née" (a real slip).
- Quotes dropped by the re-hearing check2 / 48dev (tuned on)n = 48Both had speech-recognition slips; dropped quotes are never shown.
- Speaker labels right79 / 80dev (tuned on)n = 80The miss: a 0.66 s interviewer segment labelled hers; no quote came from it.
- Meaning check: lines flagged / clear mistranslations27 / 46 flagged; 2 cleardev (tuned on)n = 46The builder read every flag: 2 clear mistranslations, 25 nitpicks or misreadings. Misses not measured.
- Consent refusals and hesitations refused2 / 2syntheticn = 2
- Photo backs read3 / 3syntheticn = 3Synthetic handwriting on the sample's photo backs.
- First trailer after upload (p50)40.8 sdev (tuned on)n = 537-71 s; interviews of 56-207 s, pre-release server.
- Cost per film$0.0085-0.02dev (tuned on)n = 5List prices for the text model and GPU time.
- Blind granddaughter: keeps it / shares it as is / would payyes / no / $39syntheticn = 1An Opus sub-agent on the Italian sample: about 40 min to fix in the tool against 6+ hours by hand (its estimate). Two of its problems fixed since: held error lines, no borrowed photos.
Dataset
5 synthetic interviews written by the builder (Italian 207 s, Spanish 70 s, Portuguese 68 s, French 56 s, German 81 s), voiced by VoxCPM2 voice design (Apache-2.0; no real person recorded or cloned), with public-domain Library of Congress photos and synthetic photo backs.
Caveats
- Synthetic interviews with two clean voices; real recordings are not measured.
- The same builder wrote the scripts, the pipeline and the eval, and fixed bugs found in an earlier run on the same set.
- Quote correctness is a fuzzy text match against the script, so years written as digits count as misses.
- The meaning-check judgement is the builder's own reading, not a blind or native-speaker review.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 29 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 41 s
- Receipts
- 11
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.009
Self-host verification
Not yet verified on a fresh self-host setup.
Rehearsal bundle: family-film.zip (102 KB, 8 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Measured on 5 synthetic interviews with two voices each, voiced by an open voice-design model; real family recordings (noise, overlapping talk, more relatives) are not measured.
- The meaning check flags about half the lines, and most flags are nitpicks: read the English yourself for a language you know.
- Quotes copy the speech recogniser's words: a one-letter slip ("né" for "née") reached a shown quote.
- Chapters follow the order she told them, not the years, and can't be reordered yet.
- One film style.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Speech recognition in her language, one call per speech segment (so every word keeps its time)Qwen3-ASR-1.7B (language pack speech service)Apache-2.0
- Who is speaking: voice activity by energy, then ECAPA-TDNN voice embeddings per segment in two clusters; the cluster nearest her consent recording is hersECAPA-TDNN (speechbrain/spkrec-ecapa-voxceleb, ONNX export)Apache-2.0
- Chapters of her life and her best lines (copied exactly from the transcript, then re-found in it by code), film titles, and the translation fallbackQwen3.8-27B (NVFP4)Apache-2.0
- English subtitles under her own words, sentence by sentence, then a meaning check (back-translation compared with the source) on every lineHy-MT2-7B (the language-pack block)Apache-2.0
- Face detection only (boxes): a picture is read as the back of a photo only when no face is found; photo pans drift toward facesUltra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)MIT
- Reads the handwriting on the backs of photos (place, year, names) to date and place each pictureDecosa document reader (Docling layout + PaddleOCR-VL-1.6)Apache-2.0
- A quiet score under the film: a pre-rendered cue from the cleared music library (no model runs per film)ACE-Step 1.5 cue (music-gen-cleared library)MIT
- The film (CPU): her voice over her photos with slow pans, chapter cards, maps (Natural Earth) and dates, subtitles, the trailer and the book with a QR code per quote; C2PA credential per filedecosa-api family_film module + FFmpeg + c2pa-pythonAGPL-3.0-or-later
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · chapters, quotes, subtitles and the film, no photo reading (2)
- Shown quotes inside her own turn, at the right time (5 synthetic interviews): 46 / 46decosa-api docs/evals/family-film.md, 2026-09-29
- Speaker labels right: 79 / 80 segmentsdecosa-api docs/evals/family-film.md, 2026-09-29
Standard · the hosted demo, with photo backs read (3)
- Shown quotes inside her own turn, at the right time (5 synthetic interviews): 46 / 46decosa-api docs/evals/family-film.md, 2026-09-29
- Photo backs read (synthetic handwriting): 3 / 3decosa-api docs/evals/family-film.md, 2026-09-29
- Meaning check: lines flagged / clear mistranslations among them: 27 of 46 flagged; 2 clear mistranslationsdecosa-api docs/evals/family-film.md, 2026-09-29; the builder read every flag