Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: turn a family interview into a film (170)

Run on our server, 2026-09-29, against a pre-release test instance of this branch (text model on the direct route, priced at list). Raw results in docs/evals/family-film/eval_family.json (keys and ids removed). Script: scripts/family_film_eval.py. Everything below was measured. There is no held-out split: the same builder wrote the interviews, the pipeline and this eval, and an earlier run on the same five interviews found two bugs that were fixed before this run (see "Fixed from an earlier run").

Data and licences

Set What Licence Used for
5 synthetic interviews (builder-written scripts) Italian (Rosa, 207 s), Spanish (Carmen, 70 s), Portuguese (Graça, 68 s), French (Odette, 56 s), German (Hanne, 81 s); a storyteller and an interviewer; a truth file with every turn's speaker, text and times ours; voices made by VoxCPM2 voice design (openbmb/VoxCPM2, Apache-2.0) from text descriptions: no person was recorded or cloned every metric below
Library of Congress Prints and Photographs 8 old photos for the Rosa sample no known restrictions / public domain photo layout, face guard
Synthetic photo backs 3 pictures of handwriting (place, year, names) ours photo-back reading
Natural Earth coastlines and places public domain journey maps

1. Quotes: hers, at the right time

The film may only show a quote that is (a) her words and (b) at the time she said them. The pipeline copies each quote from the transcript, re-finds it in the storyteller's segments by code, and then transcribes the quote's own clip again: if less than 80% of its words are heard there, the quote is dropped and never shown.

The eval scores every quote the film may show against the truth file: on time when the quote overlaps a storyteller turn by at least 80% of its span; her words when that turn's text contains it (fold + fuzzy match >= 0.9) and no interviewer turn does.

Interview Quotes Shown (traced) On time, in her turn Her words (fuzzy 0.9)
it-rosa 18 18 18 18
es-carmen 6 6 6 6
pt-graca 8 7 7 6
fr-odette 6 5 5 4
de-hanne 10 10 10 10
All 48 46 46 / 46 44 / 46
  • The 2 dropped quotes had speech-recognition slips ("Jala" for "já lá"; a line cut short). The re-hearing check caught both, so neither is shown.
  • The 2 "her words" misses are shown quotes that overlap her turn by more than 96%. The transcript writes the year as digits ("1942", "1937") where the script spells it out, which the fuzzy match counts against. One of them also has a real one-letter slip: "Je suis né" where she said "née". Quotes copy the recogniser's words, so a slip like that can reach the film.

2. Who is speaking

Each speech segment is labelled by ECAPA-TDNN embeddings in two clusters. The cluster nearest her consent recording is hers. A doubtful segment goes to the interviewer (the safe side: dropped from quotes, never put in her mouth).

it es pt fr de All
Segments labelled right 29/29 12/12 11/12 12/12 15/15 79/80

The miss is a 0.66 s interviewer segment labelled hers. No quote came from it. Two clean voices is the easy case: a table of four to six relatives, or overlapping talk, is not measured.

3. Subtitles and the meaning check

Every shown quote got English subtitles (46/46). The earlier run lost 13 of 17 Italian lines when the translation server restarted mid-run; Qwen3.8 now translates when Hy-MT2 is down. The language pack's meaning check (back-translation plus a judge) marks each line ok, check or error, and repairs an error once.

ok check error repaired
All 46 lines 18 25 2 3

I read all 27 flagged lines (the builder, not a blind or native reviewer):

  • 2 clear mistranslations: "receipt, meaning received" for Rosa's "receipt, ricevuta" (the first English word she learned), and "the same ballroom church" for "la stessa chiesa del ballo".
  • 25 nitpicks or misreadings, for example "on the stove" for "sul fuoco", "Lombardi family's bakery" for "il forno dei signori Lombardi", or the judge reading the source and target backwards. The English on those lines reads correctly.
  • Misses (bad lines marked ok) were not measured.

So the check is useful for pointing at the two real problems but noisy: it flags about half of all lines. The page tells families to read the English themselves in a language they know.

4. Consent, photos and the face guard

  • Consent: a refusal ("No, non voglio che registriate le mie storie") and a hesitation ("Mah, non lo so... forse") in her designed voice were both refused (422, code consent). Consent recordings took 1.9-2.5 s to check.
  • Photo backs: the document reader read all 3 synthetic photo backs exactly (place, year, names), about 0.12 s each.
  • Face guard: a picture with a face is never sent to the reader (tested: blocked). The fronts of photos only go through FFmpeg (pan, crop, caption).

5. Time and cost

Interview Audio Chapters ready First trailer Trailer length Receipts Cost
it-rosa 207 s 55.1 s 70.5 s 70.5 s 22 $0.0200
es-carmen 70 s 26.3 s 40.8 s 55.5 s 8 $0.0088
pt-graca 68 s 26.4 s 40.3 s 49.0 s 11 $0.0092
fr-odette 56 s 29.9 s 37.2 s 29.2 s 11 $0.0085
de-hanne 81 s 32.9 s 48.5 s 52.3 s 14 $0.0096

p50: chapters 29.9 s, first trailer 40.8 s. A full 181 s film of the Rosa interview cut in 33 s of CPU. Cost is the text model at $0.30/$1.50 per million tokens plus speech, translation and reader GPU time at $1.32/h (a shared card, so an upper bound) plus CPU at $0.10/h. Measured on a shared GPU0 with other workloads running.

6. A blind grandchild (sub-agent)

A blind Claude Code sub-agent (Opus 5.5) played the grandchild on the Italian sample. It saw only the transcript and the tool's output, never the truth file.

  • By hand: 6 chapters and 14 quotes in 93 s of agent time; its estimate for a person doing the same was 120 minutes.
  • Reviewing our draft: 27 changes (0 wrong words, 3 time overlaps, 15 subtitle notes, 13 of them the MT-outage blanks since fixed), 73 s of agent time, estimate 60 minutes. Verdict: "yes, would use the draft, but check every subtitle".

That review ran before the subtitle fix, and the estimates are the agent's own.

7. A blind viewer (sub-agent), after the fixes

A second blind Claude Code sub-agent (Opus 5.5) played Giulia, a granddaughter who reads little Italian, looking at the Rosa trailer and film (frame sheets), the chapters and the transcript. Verdict (family-film/blind_viewer_verdict.json): keeps it, would not share it as is, would pay and make a second film; about 40 minutes to fix in the tool against 6+ hours in a video editor. Most moving: "I thought it was smaller than I imagined, but then I saw my mother crying and realized it was huge for her." Its 14 problems, and what we did:

  • "Receipt, meaning received" was marked an error by the tool and still in the film and trailer: fixed, a line the meaning check still calls an error after repair is now held out of both, with "Left out: the English looks wrong. Read it, and put it back if it's right." (test added).
  • Bakery lines over the ship photo: fixed, a chapter never borrows another chapter's photo; it gets a quiet title card instead. (In the re-run the bakery lines sat over the Court Street photo.)
  • The punchline "L'ha fatto" cut off: in the re-run the quote ended "... I'll marry you." He did it." Quote boundaries vary run to run; not fixed as such.
  • "Sold us the oven" (forno = the bakery) and "the same church with the ballroom": still flagged "check", not fixed.
  • Chapters follow the order she told them, not the years (1966 before 1957): not changed. A family can rename chapters and leave quotes out, but not reorder them yet.
  • A sample photo back ("io e Lucia") sits on a landscape postcard: a quirk of our synthetic sample, not of the tool.

Fixed from an earlier run

  • Subtitles went blank when Hy-MT2 restarted mid-run: Qwen3.8 now translates as a fallback, with the same meaning check.
  • A non-ASCII person label ("Graça") broke the consent ledger entry: labels are ASCII-folded.
  • The frame check ran out of tokens at 8 pictures and failed closed (shared packs/safety.py).

Expected properties of the sample run (for the rehearsal kit)

The Rosa sample (POST /studio/projects/{pid}/family/sample):

  1. a refusal in her own words is refused (422, code consent);
  2. the sample reaches status: ready with at least 4 chapters;
  3. every shown quote is traced (tracing.rate 1.0 on this sample; at least 0.9 expected);
  4. the speaker step reports ok with her consent voice as the reference;
  5. the trailer renders (films.trailer.status = done) with a C2PA credential;
  6. every model call carries a signed receipt.

Not measured

  • Real family recordings: room noise, phone compression, overlapping talk, more than two voices, dialects.
  • Languages beyond the five above (the speech and translation models list more).
  • Misses of the meaning check, and any native-speaker review of the subtitles.
  • Whether families like the film. Only one blind sub-agent reviewed it, and no real family has.