Skip to content
decosa

170 · Creative and media · Film, TV and games · live

Family interview film

Open the toolJSON

Eval results

Not held outRun 29 Sep 2026Eval write-up (decosa-api, access required)

  • Shown quotes inside her own turn, at the right time46 / 46dev (tuned on)n = 465 languages; a quote counts when it overlaps her true turn by at least 80% of its span.
  • Shown quotes that match her words (fuzzy 0.9)44 / 46dev (tuned on)n = 46Both misses write the year as digits where the script spells it out; one also has "né" for "née" (a real slip).
  • Quotes dropped by the re-hearing check2 / 48dev (tuned on)n = 48Both had speech-recognition slips; dropped quotes are never shown.
  • Speaker labels right79 / 80dev (tuned on)n = 80The miss: a 0.66 s interviewer segment labelled hers; no quote came from it.
  • Meaning check: lines flagged / clear mistranslations27 / 46 flagged; 2 cleardev (tuned on)n = 46The builder read every flag: 2 clear mistranslations, 25 nitpicks or misreadings. Misses not measured.
  • Consent refusals and hesitations refused2 / 2syntheticn = 2
  • Photo backs read3 / 3syntheticn = 3Synthetic handwriting on the sample's photo backs.
  • First trailer after upload (p50)40.8 sdev (tuned on)n = 537-71 s; interviews of 56-207 s, pre-release server.
  • Cost per film$0.0085-0.02dev (tuned on)n = 5List prices for the text model and GPU time.
  • Blind granddaughter: keeps it / shares it as is / would payyes / no / $39syntheticn = 1An Opus sub-agent on the Italian sample: about 40 min to fix in the tool against 6+ hours by hand (its estimate). Two of its problems fixed since: held error lines, no borrowed photos.

Dataset

5 synthetic interviews written by the builder (Italian 207 s, Spanish 70 s, Portuguese 68 s, French 56 s, German 81 s), voiced by VoxCPM2 voice design (Apache-2.0; no real person recorded or cloned), with public-domain Library of Congress photos and synthetic photo backs.

Caveats

  • Synthetic interviews with two clean voices; real recordings are not measured.
  • The same builder wrote the scripts, the pipeline and the eval, and fixed bugs found in an earlier run on the same set.
  • Quote correctness is a fuzzy text match against the script, so years written as digits count as misses.
  • The meaning-check judgement is the builder's own reading, not a blind or native-speaker review.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
29 Sep 2026
Latency, this run
n/a
p50 over passed runs
41 s
Receipts
11
Model calls
n/a
Tokens
n/a
Cost per run
$0.009

Self-host verification

Not yet verified on a fresh self-host setup.

Rehearsal bundle: family-film.zip (102 KB, 8 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Measured on 5 synthetic interviews with two voices each, voiced by an open voice-design model; real family recordings (noise, overlapping talk, more relatives) are not measured.
  • The meaning check flags about half the lines, and most flags are nitpicks: read the English yourself for a language you know.
  • Quotes copy the speech recogniser's words: a one-letter slip ("né" for "née") reached a shown quote.
  • Chapters follow the order she told them, not the years, and can't be reordered yet.
  • One film style.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Speech recognition in her language, one call per speech segment (so every word keeps its time)Qwen3-ASR-1.7B (language pack speech service)Apache-2.0
  • Who is speaking: voice activity by energy, then ECAPA-TDNN voice embeddings per segment in two clusters; the cluster nearest her consent recording is hersECAPA-TDNN (speechbrain/spkrec-ecapa-voxceleb, ONNX export)Apache-2.0
  • Chapters of her life and her best lines (copied exactly from the transcript, then re-found in it by code), film titles, and the translation fallbackQwen3.8-27B (NVFP4)Apache-2.0
  • English subtitles under her own words, sentence by sentence, then a meaning check (back-translation compared with the source) on every lineHy-MT2-7B (the language-pack block)Apache-2.0
  • Face detection only (boxes): a picture is read as the back of a photo only when no face is found; photo pans drift toward facesUltra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)MIT
  • Reads the handwriting on the backs of photos (place, year, names) to date and place each pictureDecosa document reader (Docling layout + PaddleOCR-VL-1.6)Apache-2.0
  • A quiet score under the film: a pre-rendered cue from the cleared music library (no model runs per film)ACE-Step 1.5 cue (music-gen-cleared library)MIT
  • The film (CPU): her voice over her photos with slow pans, chapter cards, maps (Natural Earth) and dates, subtitles, the trailer and the book with a QR code per quote; C2PA credential per filedecosa-api family_film module + FFmpeg + c2pa-pythonAGPL-3.0-or-later

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · chapters, quotes, subtitles and the film, no photo reading (2)
  • Shown quotes inside her own turn, at the right time (5 synthetic interviews): 46 / 46decosa-api docs/evals/family-film.md, 2026-09-29
  • Speaker labels right: 79 / 80 segmentsdecosa-api docs/evals/family-film.md, 2026-09-29
Standard · the hosted demo, with photo backs read (3)
  • Shown quotes inside her own turn, at the right time (5 synthetic interviews): 46 / 46decosa-api docs/evals/family-film.md, 2026-09-29
  • Photo backs read (synthetic handwriting): 3 / 3decosa-api docs/evals/family-film.md, 2026-09-29
  • Meaning check: lines flagged / clear mistranslations among them: 27 of 46 flagged; 2 clear mistranslationsdecosa-api docs/evals/family-film.md, 2026-09-29; the builder read every flag

How we measure · All tools