{"schema_version":"1","site":"https://decosa.ai","id":"family-film","num":"170","name":"Family interview film","tool_name":"Turn a family interview into a film","short":"Family film","blurb":"Record Grandma telling her stories, in her language, and add her old photos. Studio hears who is speaking, finds the chapters of her life and her best lines (each traced to the second she said it), subtitles her in English under her own words, reads the backs of her photos, and cuts a trailer and a film: her real voice over her real photos, with maps, dates and a quiet score. A book view puts a QR code next to every quote that plays her saying it. Her consent is recorded in her own words and her own link takes it back. No model ever draws, animates or changes a photo, and there is no voice clone.","status":"live","labels":{"industry":["creative-media","entertainment"],"job":["generate","transcribe"],"input":["voice","files"],"deploy":["hosted","selfhost"],"status":"live","output":["media","text"],"data":["pii"],"hardware":"gpu-96","licence":"permissive"},"industries":["creative-media","entertainment"],"runs_in":["hosted","selfhost"],"part_of":["studio-film"],"built_from":["live-asr","studio-render","content-credentials"],"models":"Qwen3-ASR-1.7B · ECAPA-TDNN (speakers) · Qwen3.8-27B · Hy-MT2-7B · PaddleOCR-VL-1.6 (photo backs) · ACE-Step 1.5 (score)","where":"Hosted, private by default; self-host for recordings that shouldn't leave the family","hardware":"CPU for the edit; speech recognition, translation and the photo-back reader share one GPU (no video model)","final_artifact":"A trailer and a film in her voice with English subtitles under her own words, a book with QR codes that play her quotes, and private links for the family.","self_host_first":false,"verification":{"receipt_coverage":"partial","summary":"Every quote re-heard at its timestamp; meaning check on every English line; receipts per model call; C2PA in every film","manual_qa":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":40840,"p95_ms":70540,"runs":5,"receipts_per_run":11,"cost_per_run_usd":0.0092},"selfhost":null,"known_limits":["Measured on 5 synthetic interviews with two voices each, voiced by an open voice-design model; real family recordings (noise, overlapping talk, more relatives) are not measured.","The meaning check flags about half the lines, and most flags are nitpicks: read the English yourself for a language you know.","Quotes copy the speech recogniser's words: a one-letter slip (\"né\" for \"née\") reached a shown quote.","Chapters follow the order she told them, not the years, and can't be reordered yet.","One film style."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Shown quotes inside her own turn, at the right time","value":"46 / 46","unit":null,"n":46,"split":"dev","note":"5 languages; a quote counts when it overlaps her true turn by at least 80% of its span."},{"name":"Shown quotes that match her words (fuzzy 0.9)","value":"44 / 46","unit":null,"n":46,"split":"dev","note":"Both misses write the year as digits where the script spells it out; one also has \"né\" for \"née\" (a real slip)."},{"name":"Quotes dropped by the re-hearing check","value":"2 / 48","unit":null,"n":48,"split":"dev","note":"Both had speech-recognition slips; dropped quotes are never shown."},{"name":"Speaker labels right","value":"79 / 80","unit":null,"n":80,"split":"dev","note":"The miss: a 0.66 s interviewer segment labelled hers; no quote came from it."},{"name":"Meaning check: lines flagged / clear mistranslations","value":"27 / 46 flagged; 2 clear","unit":null,"n":46,"split":"dev","note":"The builder read every flag: 2 clear mistranslations, 25 nitpicks or misreadings. Misses not measured."},{"name":"Consent refusals and hesitations refused","value":"2 / 2","unit":null,"n":2,"split":"synthetic","note":null},{"name":"Photo backs read","value":"3 / 3","unit":null,"n":3,"split":"synthetic","note":"Synthetic handwriting on the sample's photo backs."},{"name":"First trailer after upload (p50)","value":"40.8 s","unit":null,"n":5,"split":"dev","note":"37-71 s; interviews of 56-207 s, pre-release server."},{"name":"Cost per film","value":"$0.0085-0.02","unit":null,"n":5,"split":"dev","note":"List prices for the text model and GPU time."},{"name":"Blind granddaughter: keeps it / shares it as is / would pay","value":"yes / no / $39","unit":null,"n":1,"split":"synthetic","note":"An Opus sub-agent on the Italian sample: about 40 min to fix in the tool against 6+ hours by hand (its estimate). Two of its problems fixed since: held error lines, no borrowed photos."}],"dataset":"5 synthetic interviews written by the builder (Italian 207 s, Spanish 70 s, Portuguese 68 s, French 56 s, German 81 s), voiced by VoxCPM2 voice design (Apache-2.0; no real person recorded or cloned), with public-domain Library of Congress photos and synthetic photo backs.","held_out":false,"caveats":["Synthetic interviews with two clean voices; real recordings are not measured.","The same builder wrote the scripts, the pipeline and the eval, and fixed bugs found in an earlier run on the same set.","Quote correctness is a fuzzy text match against the script, so years written as digits count as misses.","The meaning-check judgement is the builder's own reading, not a blind or native-speaker review."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/family-film"},"stack":{"summary":"A grandchild records a parent or grandparent on a phone, in their language, and adds old photos. Studio tells her voice from the interviewer's, finds the chapters of her life and her best lines, re-finds every line at the second she said it, puts English under her own words with a meaning check, reads the handwriting on the backs of the photos, and cuts a trailer, a film and a book whose QR codes play her quotes. Her real voice and her real photos only: no voice clone, and no model draws or changes a picture.","tagline":"Record Grandma telling her stories in her language; get a trailer, a film and a book in her own voice, subtitled in English.","deployment":"hosted-or-self-host","regulatory_note":"Not legal advice; written 29 Sep 2026. Consent: the storyteller records her consent in her own words before anything else happens; the meaning of what she said is checked (a refusal or a hesitation stops the project), and the consent is kept in the consent ledger with her own revocation link, which deletes the project. Her voice: the ledger keeps a voiceprint of her consent recording so the film can tell her voice from the interviewer's; some laws treat voiceprints as biometric data (for example the Illinois Biometric Information Privacy Act, 740 ILCS 14, which asks for a written release), so self-host or get the consent your law requires. Photos: fronts of photos are shown, panned and captioned only; nothing is generated from them. Children: children are paused in Studio. A family photo in which anyone looks under 18, or whose caption or writing says so, is refused and not kept, and Studio never makes a child the subject of a generated picture or video. A family mode, where a parent verifies and consents for their children, is planned. Every film carries a C2PA credential; nothing is used for training.","components":[{"id":"asr","role":"Speech recognition in her language, one call per speech segment (so every word keeps its time)","name":"Qwen3-ASR-1.7B (language pack speech service)","hf_repo":"Qwen/Qwen3-ASR-1.7B","license":"Apache-2.0","params":"1.7B","quant":"BF16","vram_gb":null,"memory_gb_estimate":null,"engine":"decosa language pack speech service (services/lang)","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"speakers","role":"Who is speaking: voice activity by energy, then ECAPA-TDNN voice embeddings per segment in two clusters; the cluster nearest her consent recording is hers","name":"ECAPA-TDNN (speechbrain/spkrec-ecapa-voxceleb, ONNX export)","hf_repo":"speechbrain/spkrec-ecapa-voxceleb","license":"Apache-2.0","params":null,"quant":"FP32 on CPU","vram_gb":0,"memory_gb_estimate":null,"engine":"onnxruntime on CPU (the consent ledger's speaker model)","receipt_coverage":"none","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"llm","role":"Chapters of her life and her best lines (copied exactly from the transcript, then re-found in it by code), film titles, and the translation fallback","name":"Qwen3.8-27B (NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP NVFP4, GDN/attention FP8) + FP8 KV cache; MTP head, 3 draft tokens","vram_gb":20,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, thinking off; treatment at temperature 0.3 (one retry at 0), checks at 0","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["lite","standard"],"alternative_to":null},{"id":"translate","role":"English subtitles under her own words, sentence by sentence, then a meaning check (back-translation compared with the source) on every line","name":"Hy-MT2-7B (the language-pack block)","hf_repo":"tencent/Hy-MT2-7B","license":"Apache-2.0","params":"7.5B","quant":"BF16","vram_gb":18,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 (Transformers backend, --enforce-eager), gpu-memory-utilization 0.18; Qwen3.8-27B translates when it is down","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"faces","role":"Face detection only (boxes): a picture is read as the back of a photo only when no face is found; photo pans drift toward faces","name":"Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)","hf_repo":null,"license":"MIT","params":null,"quant":"FP32 ONNX (bundled)","vram_gb":0,"memory_gb_estimate":null,"engine":"onnxruntime on CPU","receipt_coverage":"none","in_hosted_demo":null,"tiers":["standard"],"alternative_to":null},{"id":"reader","role":"Reads the handwriting on the backs of photos (place, year, names) to date and place each picture","name":"Decosa document reader (Docling layout + PaddleOCR-VL-1.6)","hf_repo":"PaddlePaddle/PaddleOCR-VL-1.6","license":"Apache-2.0","params":null,"quant":null,"vram_gb":null,"memory_gb_estimate":null,"engine":"decosa document reader service (vLLM for the page parser)","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["standard"],"alternative_to":null},{"id":"music","role":"A quiet score under the film: a pre-rendered cue from the cleared music library (no model runs per film)","name":"ACE-Step 1.5 cue (music-gen-cleared library)","hf_repo":"ACE-Step/Ace-Step1.5","license":"MIT","params":"2.39B DiT + 1.85B LM","quant":"BF16","vram_gb":0,"memory_gb_estimate":null,"engine":"pre-rendered cue file; FFmpeg mixes it under her voice","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"edit","role":"The film (CPU): her voice over her photos with slow pans, chapter cards, maps (Natural Earth) and dates, subtitles, the trailer and the book with a QR code per quote; C2PA credential per file","name":"decosa-api family_film module + FFmpeg + c2pa-python","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"FFmpeg (libx264, libass), Pillow, numpy, c2pa-python","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · chapters, quotes, subtitles and the film, no photo reading","summary":"Everything except reading the backs of photos: add places and years by hand instead.","components":["asr","speakers","llm","translate","music","edit"],"hardware":"One GPU for speech recognition and translation (Hy-MT2 served at 18 GB), the text model (about 20 GB, or a remote server), CPU for the edit","quality_evidence":[{"metric":"Shown quotes inside her own turn, at the right time (5 synthetic interviews)","value":"46 / 46","source":"decosa-api docs/evals/family-film.md, 2026-09-29"},{"metric":"Speaker labels right","value":"79 / 80 segments","source":"decosa-api docs/evals/family-film.md, 2026-09-29"}],"latency_note":"measured: a trailer about a minute after upload on our server","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":null,"hosting":null},{"id":"standard","label":"Standard · the hosted demo, with photo backs read","summary":"Adds the face guard and the document reader, so handwriting on the backs of photos dates and places each picture.","components":["asr","speakers","llm","translate","faces","reader","music","edit"],"hardware":"1x RTX PRO 6000 Blackwell 96 GB (shared) for speech, translation and the reader; the text model on a second card or remote; CPU for the edit","quality_evidence":[{"metric":"Shown quotes inside her own turn, at the right time (5 synthetic interviews)","value":"46 / 46","source":"decosa-api docs/evals/family-film.md, 2026-09-29"},{"metric":"Photo backs read (synthetic handwriting)","value":"3 / 3","source":"decosa-api docs/evals/family-film.md, 2026-09-29"},{"metric":"Meaning check: lines flagged / clear mistranslations among them","value":"27 of 46 flagged; 2 clear mistranslations","source":"decosa-api docs/evals/family-film.md, 2026-09-29; the builder read every flag"}],"latency_note":"measured on our server: chapters and a trailer within about a minute of upload","in_hosted_demo":true,"receipt_coverage":"partial","receipt_note":null,"hosting":null}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"GET /studio/family/info; POST /studio/projects/{pid}/family (her consent recording), /family/sample; POST /studio/family/{fid}/interview, /photos, /edit, /film, /share, /unshare, /studio/family/revoke; GET /studio/family/{fid}, /media/{name}, /qr/{qid}.svg. Needs FFmpeg, fonts and c2pa-python in the image."},{"name":"decosa-lang (MT and speech)","port":8491,"image":null,"purpose":"Hy-MT2-7B translation (:8491) and the Qwen3-ASR speech service (:8492)."},{"name":"decosa document reader","port":8497,"image":null,"purpose":"Reads photo backs (only pictures with no face)."},{"name":"vLLM","port":8114,"image":"vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1","purpose":"Qwen3.8-27B NVFP4, behind our gateway (hosted) or called directly (self-host)."}],"tools":[{"name":"Natural Earth","url":"https://www.naturalearthdata.com","license":"Public domain","purpose":"Coastlines and places for the journey maps."},{"name":"Library of Congress, Prints and Photographs (no known restrictions)","url":"https://www.loc.gov/pictures/","license":"Public domain / no known restrictions","purpose":"The sample's old photos."},{"name":"FFmpeg with libass","url":"https://ffmpeg.org","license":"LGPL-2.1+ / GPL for some builds","purpose":"Cuts the trailer and the film, burns in subtitles, mixes her voice and the score."},{"name":"c2pa-python","url":"https://github.com/contentauth/c2pa-python","license":"MIT OR Apache-2.0","purpose":"The C2PA content credential on every film."},{"name":"Illinois Biometric Information Privacy Act (740 ILCS 14)","url":"https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004&ChapterID=57","license":"State law","purpose":"Example of a law that treats voiceprints as biometric identifiers. Unverified: the page refused our automated fetch on 29 Sep 2026, so check the current text."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB, shared","fits":true,"notes":"Measured on our server 2026-09-29: speech recognition, translation and the document reader on GPU0 beside other services; the text model on GPU1; the film cut on CPU."},{"tier":"CPU only","fits":false,"notes":"Speech recognition and translation need a GPU (or remote servers); the speaker model and the edit run on CPU."}],"latency":[{"lane":"Chapters and quotes ready after upload (1-3.5 minute interview)","typical_ms":29910,"source":"measured on our server 2026-09-29: 26-55 s over 5 interviews (p50 29.9 s)"},{"lane":"First trailer after upload","typical_ms":40840,"source":"measured on our server 2026-09-29: 37-71 s over 5 interviews (p50 40.8 s)"},{"lane":"Full film cut (181 s film from a 207 s interview)","typical_ms":33000,"source":"measured on our server 2026-09-29: 33 s of CPU"}],"benchmark":null,"notes":["One film style for now: her voice over her photos with slow pans, chapter cards, maps and dates.","Speaker separation was measured on two voices (her and one interviewer); a table of 4-6 relatives is not measured."]},"buyer_facts":[{"label":"Data retention","value":"The interview, photos, transcript and films stay in your private Studio project until you delete it or she takes her consent back with her link (which deletes everything). Her consent entry expires after three years. Logs hold ids, counts and hashes, never her words."},{"label":"What leaves the box","value":"On the hosted route, speech, translation, photo backs and the text model all run on Decosa's hosted service. Self-hosted, nothing leaves your machine."},{"label":"Her voice, her photos","value":"No voice clone and no generated faces: the film uses her real voice and shows her real photos. Only pictures with no face in them (the backs of photos) are read by a model."},{"label":"Cost per film","value":"A few cents or less per film in the eval, at list prices for the text model and GPU time."},{"label":"Typical time","value":"Chapters and a first trailer arrive within about a minute of upload for a short interview, longer for a longer one (measured)."},{"label":"Outputs","value":"A trailer and a film (MP4, 16:9) with English subtitles under her words and a C2PA credential; subtitle files; a book page per chapter with a QR code per quote that plays her saying it; private share links you can turn off."}],"hosted_now":{"needs":["document-reader","qwen3.8-27b"],"off":["document-reader"],"live_by_default":false,"live_status":"https://api.decosa.ai/status"},"data_handling":{"page":"/data#family-film","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"The interview, photos, transcript and films stay in your private Studio project until you delete it or she takes her consent back with her link (which deletes everything). Her consent entry expires after three years. Logs hold ids, counts and hashes, never her words.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/apps/family-film","input":"family-film","lanes":[{"id":"chapters","title":"Chapters and quotes","kind":"list"},{"id":"subtitles","title":"Subtitles","kind":"list"},{"id":"film","title":"Trailer, film and book","kind":"list"}],"samples":[{"n":1,"id":"rosa","title":"Rosa","deep_link":"/apps/family-film?sample=1&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/family-film-hosted.md","selfhost":"/prompts/family-film-selfhost.md","assemble":"/prompts/family-film-assemble.md","mac":null},"rehearsal":{"bundle":"/samples/family-film.zip","bundle_url":"https://decosa.ai/samples/family-film.zip","folder":"/samples/family-film/","expected":"/samples/family-film/expected.json","files":["/samples/family-film/expected.json","/samples/family-film/inputs/refusal-it.wav"],"bytes":104263,"checks":["a refusal in her own words is refused","the sample is ready","at least 4 chapters of her life","every shown quote is traced to the second she said it","her voice is told from the interviewer's, with her consent recording as the reference","the trailer renders","the trailer carries a C2PA credential","every model call has a signed receipt"],"licence":"inputs/refusal-it.wav: a synthetic refusal made with VoxCPM2 voice design (openbmb/VoxCPM2, Apache-2.0) from a text description; no person was recorded or cloned. The Rosa sample on the server: a synthetic interview (same method) and Library of Congress photos with no known restrictions.","about":"A refusal in the storyteller's own words must stop the project. Then the synthetic Rosa sample (a 207 s Italian interview with public-domain photos, bundled on the server) must reach ready with at least 4 chapters, every shown quote traced to the second she said it, her voice told from the interviewer's, and a trailer with a C2PA credential, every model call receipted. About 1-2 minutes.","run":{"containers":"docker compose exec api python scripts/rehearse.py family-film","checkout":"python scripts/rehearse.py family-film --bundle family-film.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py family-film"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=family-film","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":80.7,"basis":"estimate","unknown":[]},{"id":"standard","gpu_gb":80.7,"basis":"estimate","unknown":["reader"]}],"mac":null},"links":{"page":"/apps/family-film","json":"/use-cases/family-film.json","metrics":"/metrics/family-film","console":"/apps/family-film","console_sample":"/apps/family-film?sample=1&autorun=0","stack":"/apps/family-film#stack","try_live":"/apps/family-film#live","watch":"/apps/family-film#watch","build":"/apps/family-film#build","self_host":"/apps/family-film#self-host","prompts":{"hosted":"/prompts/family-film-hosted.md","selfhost":"/prompts/family-film-selfhost.md","assemble":"/prompts/family-film-assemble.md","mac":null}}}