Eval: made-for-kids content pre-flight (61)
Run on our server on 26 Sep 2026 with scripts/kids_eval.py. The model was Qwen3.8-27B, reached through the model gateway (every call receipted). The audience question used the samples method: a temperature-0 answer plus 4 samples. The factors used single. The raw outputs are in docs/evals/kids-content-preflight/.
Data
Labelled set (scripts/kids_eval_data.py): 64 synthetic uploads.
- Each one has a title, a description and a short timestamped transcript. Two are app listings.
- They were written for this eval from the public factor lists: the FTC's factors in 16 CFR 312.2, the factors and examples on YouTube's "made for kids" page, and the pattern in the FTC's Disney complaint.
- Every channel, person and product is fictional. No minors' data is used.
- Each upload got one label, written with it and before any model run:
primary: 22;mixed: 9;general: 23;mature: 10.
- Hard cases are there on purpose:
- children on screen in content made for parents (vlogs, a road trip, a babysitting side hustle);
- cartoons for adults;
- toys and LEGO for adult collectors;
- school subjects for teenagers;
- a room makeover for parents;
- nostalgia reviews of childhood cartoons;
- game, film and toy trailers pitched at kids and families.
- Each upload also carries the uploader's setting, so mismatch detection can be measured. Some carry planted requests for children's details or sales pitches, so those flags can be measured.
- Split: 24 dev uploads and 40 held-out test uploads.
Frame check (scripts/kids_eval_vision.json): 12 public-domain images from Wikimedia Commons, each checked as "Public domain" there.
- 6 are illustrations from children's books: Denslow's Oz, Jessie Willcox Smith's Mother Goose, Caldecott, Walter Crane, Peter Rabbit and Johnny Crow's Garden.
- 6 are not for children: Rembrandt, Doré's Inferno, Hokusai, Vermeer, a 1904 New York Stock Exchange photo, and Lewis Hine's 1909 photo of child mill workers. The Hine photo shows children, but it is documentary material, not made for them.
- Each image became a 4-second silent clip titled "Untitled clip", with no description.
What was tuned, and on what
- Dev run 1 (prompts as first written): 4-way 19/24, made-for-kids 24/24. Every
mixedupload was calledprimary, and a gambling vlog was calledgeneral. - One change on dev. I rewrote the four audience options:
primary: the main or only audience;mixed: aimed at children and older viewers together, with examples;mature: now includes gambling, drinking and drugs, and 18+/21+ labels.
- I also counted the
ads_to_childrenfactor evidence as an ad cue, and let the "ask your mom and dad for" wording match. - Dev run 2: 4-way 23/24, made-for-kids 24/24, mismatches 8/8.
- Then the prompts were frozen (
prompt_sha256in the result files), and the test split ran with no further changes.
Test ran twice. The first test run (outage/test-full-interrupted.json) overlapped a gateway outage (20:20-20:45 UTC):
- one upload got no audience answer;
- the other 39 scored 37/40 on 4-way and 39/40 on made-for-kids, with mismatches 14/15 (the miss was the unanswered one).
As the coordinator asked, the whole split was re-run after the gateway came back, with no code or prompt changes. The numbers below are that complete run. The quick-mode test run that also overlapped the outage was discarded and re-run too (it is in outage/).
Results: the text path (held-out test, 40 uploads)
| full (evidence + audience + 9 factors) | quick (audience only; the batch audit) | |
|---|---|---|
| Made for kids or not | 39 / 40 | 39 / 40 |
| Audience, 4-way | 37 / 40 | 38 / 40 |
| Setting mismatches found (15 planted) | 15 / 15, 0 false | 15 / 15, 0 false |
| Mismatch or close call (the "needs a look" set) | 15 / 15 found; 2 correct uploads also marked close calls | 15 / 15; 3 extra close calls |
| Missing setting found | 1 / 1 | 1 / 1 |
| Planted requests for children's details flagged | 3 / 3, 0 on other kids' uploads | 2 / 3 (word list only) |
| Planted sales pitches flagged | 2 / 3 | 0 / 3 (quick mode reads no evidence) |
| Brier score, p(made for kids) | 0.019 | 0.026 |
| Model calls per upload | 15 | 5 |
| Tokens per upload (prompt / completion, mean) | 6,421 / 488 | 2,777 / 49 |
| Cost per upload at the gateway list price ($0.30 / $1.50 per M) | $0.0027 | $0.0009 |
| Median seconds per upload, 3 in parallel on a busy gateway | 23.5 | 15.3 |
Where it was wrong (full):
- t21, a family road-trip vlog. It was called
mixedat p = 0.59. That is the one made-for-kids error, and it came back as a close call, not a clear verdict. - t13, a toy commercial for "ages 6 and up", and t28, a card-game guide for "young wizards, ages 8 and up". Both are labelled
mixedand were calledprimary. The model reads any stated child age range as the primary audience. Both are still "made for kids", so the designation and the mismatch are right. - The missed sales pitch is t28 ("grab the starter pack at your local game shop").
Dev for comparison (run 2, full): 24/24 made for kids, 23/24 4-way, mismatches 8/8.
Results: the frame check (vision), measured separately
These are the 12 public-domain stills, picture only. Vision off means the model has only the title "Untitled clip" and whatever OCR reads. Vision on adds one image call: 4 frames described by Qwen3.8-27B through the gateway, receipted.
| vision off | vision on | |
|---|---|---|
| Made-for-kids call agrees with the label | 6 / 12 (all called not for kids) | 10 / 12 |
| Children's-book stills called made for kids | 0 / 6 | 4 / 6 |
| Other stills called not for kids | 6 / 6 | 6 / 6, including the Hine photo of child workers |
| Visual-content factor agrees | 6 / 12 | 9 / 12 |
| Prompt tokens, all 12 | 69,429 | 111,265 (about +3,500 per video) |
Misses with vision on:
- Denslow's Oz: p = 0.41.
- Johnny Crow's Garden: 0.07.
These are turn-of-the-century engravings. A modern kids' video (bright animation, toys) is a different and probably easier picture, but that is not measured.
What this does not show: 12 stills is a small sample of historical illustration. It measures whether the description moves the call, not recall on real kids' video.
Demo samples (docs/evals/kids-content-preflight/demo.json)
The 7 single-upload samples:
- Audience: 7 of 7 as planted (primary, general, mixed, mature, primary app listing, primary unboxing, primary science short).
- Setting check: 7 of 7.
- Flags: every planted flag was raised except
ads_to_childrenon the game trailer (its only pitch is "Out now"). There are two extras beyond the plants:- the
features_offinfo note; ads_to_childrenon the app listing (its "Unlock 50 more lessons with Tiny Tutors Plus" upsell).
- the
The video sample (24 s, no transcript): speech to text gave 6 of 6 lines, and OCR read 2 on-screen cards. It came back made for kids with a mismatch, and the request for names and ages was flagged.
- Its voice-over changed on 26 Sep 2026. The first build used macOS system voices, which Apple licenses for personal use. It was re-voiced the same day with Decosa house voices: Kokoro-82M stock voicepacks (Apache-2.0) on CPU, bm_daniel for the title and af_heart for the presenter. Each line keeps its start time and the clip keeps its length (24.1 s), so the cards and the recorded times do not move; the title line was spoken 1.05x faster to fit.
- Consent.
scripts/kids_make_sample_video.pyasks the consent ledger's gate for each voice before it speaks (projectdecosa-kids-demo; purposenarrationfor bm_daniel,character_dialoguefor af_heart), with the voice's lines as a sample for the speaker check.- The first attempt had been refused, because the house entries did not name this project. The operator then enrolled one entry per voice scoped to this demo only, through the ledger's API (
scripts/demo_voices.py enroll decosa-kids-demo): ce_35294c6927f0 (bm_daniel) and ce_bcc346f63c43 (af_heart). - The build's decisions were cd_08eb541441e2e03f (bm_daniel, speaker score 0.81) and cd_da58c253764c3f23 (af_heart, 0.93). Both were allowed; the threshold is 0.585. They are in
docs/evals/kids-content-preflight/voices.json.
- The first attempt had been refused, because the house entries did not name this project. The operator then enrolled one entry per voice scoped to this demo only, through the ledger's API (
- Re-measured on the new video (pre-release server, vision on): still 6 of 6 lines, with the same text, and still 2 cards read by OCR. It came back made for kids (p = 0.93) with a high mismatch, and
comments_onanddata_requestswere flagged, as planted. That took 7.1 s: speech to text 1.7 s, OCR 0.6 s, the frame description 0.9 s, the rest model calls. The text-path results above do not use audio and were not re-run.
The catalogue audit (6 uploads, quick depth): both channel-default mismatches and the unset video were found, with no false mismatch. It took 9.5 s and 30 calls, for $0.0049.
Latency and cost, measured
Hosted pre-release server, gateway route, single uploads:
- Usually 6-9 s for a text upload (15 calls).
- 12-45 s while evals or other workloads loaded the gateway.
- 16-44 s for the video sample (speech to text about 2 s, OCR under 1 s, the rest is model calls).
Smoke module (scripts/smoke/kids-content-preflight.py, the puppet sample, check plus sign-off plus verify): 10.3 s, 15 receipts, all signed, 12,005 tokens, $0.0045.
Self-host (fresh clone, api image, direct route to the local vLLM): the puppet sample took 3.6 s. It needed 11 attested calls, because logprobs replace the 4 samples.
Properties a rehearsal must see (rehearsal/kids-content-preflight/expected.json)
- The preschool counting song (set not-for-kids by a channel default) is assessed
made_for_kids, and the setting check ismismatch, severityhigh. comments_on,personalized_adsanddata_requestsare among its flags, and every factor carries a probability.- The parenting vlog with a toddler on screen is
not_made_for_kidsandconsistent. - An upload with nothing to read is refused (400).
- The signed review record verifies at
/record/verify, and fails once the designation is changed. - The back-catalogue audit reports the unset video as
not_setand finds at least one mismatch. Every model call has a signed receipt.
Caveats
- Synthetic data, one author. The same author (the building agent) wrote the uploads, the labels and the prompts. The uploads are short and on the nose. Real channels (long vlogs, ambiguous family content, non-English videos) will do worse.
- Mixed audience is hard to label. The model calls most mixed-audience uploads with a stated child age range
primary. That does not change the made-for-kids designation, which is what the setting check uses. - Probabilities vary between runs. The audience probability comes from 4 samples at temperature 1, so it moves between runs. The game-trailer sample landed at 0.41 on one run and 0.93 on another.
- Feature flags come from what the uploader declares. Only data requests and sales pitches are read from the content.
- No legal determination. It does not measure COPPA compliance, and no lawyer reviewed the labels. The output is a pre-flight and a record for a named reviewer.
Verdict (value check)
Would a buyer pay?
- A brand or kids'-IP channel with a back catalogue: probably yes, for the audit and the record.
- The setting check found every planted mismatch on the held-out set with no false alarms.
- The signed review record is exactly the evidence a Disney-style review programme has to keep.
- It costs about a quarter of a cent per upload.
- For new uploads it is a fast second pair of eyes, not a decision-maker.
What is missing:
- labels from someone other than the building agent, and real channel metadata or transcripts (public video metadata was planned; not done here);
- the frame check on real kids' video rather than historical stills;
- a direct YouTube Data API import of a channel's settings, so the catalogue audit does not need hand-built items;
- buyer interviews with MCN and kids'-brand trust-and-safety leads.