Eval: turn a family interview into a film (170)
Run on our server, 2026-09-29, against a pre-release test instance of this branch (text model on the direct route, priced
at list). Raw results in docs/evals/family-film/eval_family.json (keys and ids removed). Script:
scripts/family_film_eval.py. Everything below was measured. There is no held-out split: the same builder wrote the
interviews, the pipeline and this eval, and an earlier run on the same five interviews found two bugs that were fixed
before this run (see "Fixed from an earlier run").
Data and licences
| Set | What | Licence | Used for |
|---|---|---|---|
| 5 synthetic interviews (builder-written scripts) | Italian (Rosa, 207 s), Spanish (Carmen, 70 s), Portuguese (Graça, 68 s), French (Odette, 56 s), German (Hanne, 81 s); a storyteller and an interviewer; a truth file with every turn's speaker, text and times | ours; voices made by VoxCPM2 voice design (openbmb/VoxCPM2, Apache-2.0) from text descriptions: no person was recorded or cloned | every metric below |
| Library of Congress Prints and Photographs | 8 old photos for the Rosa sample | no known restrictions / public domain | photo layout, face guard |
| Synthetic photo backs | 3 pictures of handwriting (place, year, names) | ours | photo-back reading |
| Natural Earth | coastlines and places | public domain | journey maps |
1. Quotes: hers, at the right time
The film may only show a quote that is (a) her words and (b) at the time she said them. The pipeline copies each quote from the transcript, re-finds it in the storyteller's segments by code, and then transcribes the quote's own clip again: if less than 80% of its words are heard there, the quote is dropped and never shown.
The eval scores every quote the film may show against the truth file: on time when the quote overlaps a storyteller turn by at least 80% of its span; her words when that turn's text contains it (fold + fuzzy match >= 0.9) and no interviewer turn does.
| Interview | Quotes | Shown (traced) | On time, in her turn | Her words (fuzzy 0.9) |
|---|---|---|---|---|
| it-rosa | 18 | 18 | 18 | 18 |
| es-carmen | 6 | 6 | 6 | 6 |
| pt-graca | 8 | 7 | 7 | 6 |
| fr-odette | 6 | 5 | 5 | 4 |
| de-hanne | 10 | 10 | 10 | 10 |
| All | 48 | 46 | 46 / 46 | 44 / 46 |
- The 2 dropped quotes had speech-recognition slips ("Jala" for "já lá"; a line cut short). The re-hearing check caught both, so neither is shown.
- The 2 "her words" misses are shown quotes that overlap her turn by more than 96%. The transcript writes the year as digits ("1942", "1937") where the script spells it out, which the fuzzy match counts against. One of them also has a real one-letter slip: "Je suis né" where she said "née". Quotes copy the recogniser's words, so a slip like that can reach the film.
2. Who is speaking
Each speech segment is labelled by ECAPA-TDNN embeddings in two clusters. The cluster nearest her consent recording is hers. A doubtful segment goes to the interviewer (the safe side: dropped from quotes, never put in her mouth).
| it | es | pt | fr | de | All | |
|---|---|---|---|---|---|---|
| Segments labelled right | 29/29 | 12/12 | 11/12 | 12/12 | 15/15 | 79/80 |
The miss is a 0.66 s interviewer segment labelled hers. No quote came from it. Two clean voices is the easy case: a table of four to six relatives, or overlapping talk, is not measured.
3. Subtitles and the meaning check
Every shown quote got English subtitles (46/46). The earlier run lost 13 of 17 Italian lines when the translation server restarted mid-run; Qwen3.8 now translates when Hy-MT2 is down. The language pack's meaning check (back-translation plus a judge) marks each line ok, check or error, and repairs an error once.
| ok | check | error | repaired | |
|---|---|---|---|---|
| All 46 lines | 18 | 25 | 2 | 3 |
I read all 27 flagged lines (the builder, not a blind or native reviewer):
- 2 clear mistranslations: "receipt, meaning received" for Rosa's "receipt, ricevuta" (the first English word she learned), and "the same ballroom church" for "la stessa chiesa del ballo".
- 25 nitpicks or misreadings, for example "on the stove" for "sul fuoco", "Lombardi family's bakery" for "il forno dei signori Lombardi", or the judge reading the source and target backwards. The English on those lines reads correctly.
- Misses (bad lines marked ok) were not measured.
So the check is useful for pointing at the two real problems but noisy: it flags about half of all lines. The page tells families to read the English themselves in a language they know.
4. Consent, photos and the face guard
- Consent: a refusal ("No, non voglio che registriate le mie storie") and a hesitation ("Mah, non lo so... forse")
in her designed voice were both refused (422, code
consent). Consent recordings took 1.9-2.5 s to check. - Photo backs: the document reader read all 3 synthetic photo backs exactly (place, year, names), about 0.12 s each.
- Face guard: a picture with a face is never sent to the reader (tested: blocked). The fronts of photos only go through FFmpeg (pan, crop, caption).
5. Time and cost
| Interview | Audio | Chapters ready | First trailer | Trailer length | Receipts | Cost |
|---|---|---|---|---|---|---|
| it-rosa | 207 s | 55.1 s | 70.5 s | 70.5 s | 22 | $0.0200 |
| es-carmen | 70 s | 26.3 s | 40.8 s | 55.5 s | 8 | $0.0088 |
| pt-graca | 68 s | 26.4 s | 40.3 s | 49.0 s | 11 | $0.0092 |
| fr-odette | 56 s | 29.9 s | 37.2 s | 29.2 s | 11 | $0.0085 |
| de-hanne | 81 s | 32.9 s | 48.5 s | 52.3 s | 14 | $0.0096 |
p50: chapters 29.9 s, first trailer 40.8 s. A full 181 s film of the Rosa interview cut in 33 s of CPU. Cost is the text model at $0.30/$1.50 per million tokens plus speech, translation and reader GPU time at $1.32/h (a shared card, so an upper bound) plus CPU at $0.10/h. Measured on a shared GPU0 with other workloads running.
6. A blind grandchild (sub-agent)
A blind Claude Code sub-agent (Opus 5.5) played the grandchild on the Italian sample. It saw only the transcript and the tool's output, never the truth file.
- By hand: 6 chapters and 14 quotes in 93 s of agent time; its estimate for a person doing the same was 120 minutes.
- Reviewing our draft: 27 changes (0 wrong words, 3 time overlaps, 15 subtitle notes, 13 of them the MT-outage blanks since fixed), 73 s of agent time, estimate 60 minutes. Verdict: "yes, would use the draft, but check every subtitle".
That review ran before the subtitle fix, and the estimates are the agent's own.
7. A blind viewer (sub-agent), after the fixes
A second blind Claude Code sub-agent (Opus 5.5) played Giulia, a granddaughter who reads little Italian, looking at the
Rosa trailer and film (frame sheets), the chapters and the transcript. Verdict (family-film/blind_viewer_verdict.json):
keeps it, would not share it as is, would pay and make a second film; about 40 minutes to fix in the tool against
6+ hours in a video editor. Most moving: "I thought it was smaller than I imagined, but then I saw my mother crying and
realized it was huge for her." Its 14 problems, and what we did:
- "Receipt, meaning received" was marked an error by the tool and still in the film and trailer: fixed, a line the meaning check still calls an error after repair is now held out of both, with "Left out: the English looks wrong. Read it, and put it back if it's right." (test added).
- Bakery lines over the ship photo: fixed, a chapter never borrows another chapter's photo; it gets a quiet title card instead. (In the re-run the bakery lines sat over the Court Street photo.)
- The punchline "L'ha fatto" cut off: in the re-run the quote ended "... I'll marry you." He did it." Quote boundaries vary run to run; not fixed as such.
- "Sold us the oven" (forno = the bakery) and "the same church with the ballroom": still flagged "check", not fixed.
- Chapters follow the order she told them, not the years (1966 before 1957): not changed. A family can rename chapters and leave quotes out, but not reorder them yet.
- A sample photo back ("io e Lucia") sits on a landscape postcard: a quirk of our synthetic sample, not of the tool.
Fixed from an earlier run
- Subtitles went blank when Hy-MT2 restarted mid-run: Qwen3.8 now translates as a fallback, with the same meaning check.
- A non-ASCII person label ("Graça") broke the consent ledger entry: labels are ASCII-folded.
- The frame check ran out of tokens at 8 pictures and failed closed (shared
packs/safety.py).
Expected properties of the sample run (for the rehearsal kit)
The Rosa sample (POST /studio/projects/{pid}/family/sample):
- a refusal in her own words is refused (422, code
consent); - the sample reaches
status: readywith at least 4 chapters; - every shown quote is traced (
tracing.rate1.0 on this sample; at least 0.9 expected); - the speaker step reports
okwith her consent voice as the reference; - the trailer renders (
films.trailer.status=done) with a C2PA credential; - every model call carries a signed receipt.
Not measured
- Real family recordings: room noise, phone compression, overlapping talk, more than two voices, dialects.
- Languages beyond the five above (the speech and translation models list more).
- Misses of the meaning check, and any native-speaker review of the subtitles.
- Whether families like the film. Only one blind sub-agent reviewed it, and no real family has.