Eval: script to animatic (use case 50), 26 Sep 2026
Measured on our server with the branch API (the pre-release branch, port 8457): Qwen3.8-27B through the model gateway
(receipted), Wan2.2-VACE-Fun-A14B in the shared studio ComfyUI on GPU0 (one job at a time, alongside other workloads'
renders), Kokoro-82M on CPU. Script: scripts/animatic_eval.py; raw results in docs/evals/animatic-studio/
(shotlists.json, planted.json, consistency.json, consistency-frames.json). After these numbers the shot-list prompt gained one rule (original looks, with shape, colour and silhouette for
non-humans); coverage was not re-measured with it. Development used the Last Lamp sample:
the prompt rule against adding staging details and the context lines for the judge were added after a first run on it
(26 Sep, 01:40), before these numbers. Nothing was tuned after them; one fix found here (parentheticals shown to the
judge) went in afterwards and is not re-measured.
Data (all synthetic or public domain)
| Script | Kind | Lines needing a shot | Licence |
|---|---|---|---|
| The Last Lamp | screenplay, 2 characters (one V.O.) | 15 | original, CC0 |
| Hollin Oats: Night Shift | 30 s ad script, VO and SUPER, made-up brand | 11 | original, CC0 |
| Salvage | game cutscene, pilot and robot | 11 | original, CC0 |
| The Importance of Being Earnest, Act I (handbag scene) | stage play, directions inside speeches | 18 | Wilde 1895, public domain (Gutenberg #844) |
| Trifles (opening) | stage play, 4 speakers, long directions | 23 | Glaspell 1916, public domain in the US (Gutenberg #10623) |
1. Shot-list coverage and invented dialogue (10 runs: 5 scripts x 2)
- Coverage by the model's own list, before any code repair: 100% in 10 of 10 runs (78 required lines per pass). Code would attach a missed line to the nearest shot and mark it; it never had to.
- Invalid line citations 0, shots out of script order 0, a dialogue line cited in two shots 0, truncated JSON 0.
- Invented dialogue: 0 by construction. The model cites dialogue by line id and never writes it; code inserts the
script's words. A quoted phrase in an action that is not in the cited lines: 0 in 10 runs. A human edit cannot add
dialogue either:
validate_shotsrebuilds it from line ids (test). - Grounding verdicts on the model's shot actions (151): 128 supported, 17 partial (adds a detail, e.g. "on the table"), 3 unsupported, 3 contradicted. I read the 6 hard flags: 2 were real (Trifles: the action named the County Attorney for a line the Sheriff speaks; the ad: "Dana continues walking through the corridor" after she has walked out), 4 were over-strict ("Kettle speaks to Riva" twice, "pointing towards the rocker" where the stage direction says so, and "Dana speaks to herself", which the judge could not see because parentheticals were left out of its evidence; fixed).
- Latency per script (one shot-list call plus one grounding call per shot, gateway under load): 7.9-29.6 s. Cost at the gateway list price ($0.30 / $1.50 per Mtok): $0.0035-0.013 per shot list (8-25 receipted calls).
2. Does the grounding check catch invented action? (planted)
Two shot actions per script were replaced with invented events ("A fire breaks out and Margit runs outside", "Police sirens wail as two officers burst through the door", ...): 10 plants over the 5 scripts. 10 of 10 were judged unsupported. The 66 unplanted shots in the same pass: 57 supported, 9 partial, 0 unsupported or contradicted. The plants are blunt on purpose; a subtle invention ("she smiles") lands in partial at best. Plants and checker have the same author.
3. Character consistency
18 single-character shots (6 each from Last Lamp, Salvage and Earnest), three arms, same seed per shot: unlocked = the frame prompt names the character only; locked = the hosted default, every frame repeats the character's look word for word; reference = locked plus the character sheet as a Wan2.2-VACE reference image (strength 0.5). Scored with CLIP ViT-L/14 on CPU, on the raw model frames (before the sketch filter): identity = cosine(frame, that character's sheet); ident-acc = the nearest sheet is the right character (Salvage and Earnest only, 12 frames; chance about 0.42); within = mean pairwise cosine between frames of the same character; adherence = 100 x CLIP text-image score against the shot's framing and action (without the look). Whole-frame CLIP is confounded by background and framing: read these as differences between arms, not absolute truth.
| Arm | identity | ident-acc | within | adherence | median s per frame (shared GPU) |
|---|---|---|---|---|---|
| unlocked | 0.528 | 0.33 | 0.624 | 16.6 | 24 |
| locked (default) | 0.690 | 0.75 | 0.775 | 18.3 | 27 |
| reference | 0.812 | 1.00 | 0.845 | 14.9 | 42 |
Locking the look is what makes characters recognisable at all (identity +0.16, ident-acc 0.33 to 0.75). The reference arm scores highest on identity because it copies the sheet: same pose, same framing and a textured artefact background where the scene should be (see the grid), so adherence drops. It is not used; an edit model conditioned on the sheet is the "wanted" tier. Sheet, unlocked, locked and reference for four shots:

Seen in the frames and not in the scores: from a one-line description ("a squat maintenance robot with one round lens"), the robot KETTLE came out close to a well-known film robot. Those frames were withdrawn and are not published anywhere (site, Watch replay, this doc); the grid above leaves KETTLE out, though its two frames are still in the scores. The sample now spells KETTLE out (an orange kettle-shaped drone on six spider legs, one red lens on the spout; no treads, no eyes) and the shot-list prompt asks for original looks with shape, colour and silhouette for non-humans. The re-rendered Salvage cut shows a clearly original drone. Frames aren't checked against existing characters; review for resemblance before sharing.
4. Render time and cost
Full animatics through the API (studio queue; GPU0 shared with other workloads' renders throughout):
| Sample | Shots (motion) | Animatic | Render | GPU minutes per minute of animatic |
|---|---|---|---|---|
| Last Lamp, sketch | 15 (1) | 49 s | 359 s | 7.3 |
| Earnest, sketch | 18 (0) | 154 s | 409 s | 2.7 |
| Salvage, colour (first design, withdrawn) | 11 (1) | 35 s | 557 s | 15.8 |
| Salvage, colour (re-rendered, GPU busy) | 12 (1) | 39 s | 1161 s | 29.9 |
| Night Shift ad, grey | 7 (0) | 25 s | 726 s | 28.5 (GPU busy with other jobs) |
| Night Shift ad, browser run | 8 (0) | 27 s | 531 s | 19.3 |
| Night Shift ad, self-host container | not recorded | not recorded | 576 s |
- Per frame: about 13 s with the model loaded and the GPU free (6 steps, cfg 2, 1280x720); 15-88 s per frame in these runs depending on other jobs. A motion clip (81 frames, 832x480): 75-81 s. Character sheets: 30-90 s each. Temp voices: 5-15 s on CPU. Assembly: 4-19 s. Re-rendering one edited shot and re-cutting: 47 s and 45 s (two runs).
- Short scripts cost the most GPU time per minute of animatic (sheets and model loads are fixed costs).
- Cost: the model calls ($0.0035-0.013 per script). The render is 6-12 minutes of a shared RTX PRO 6000; it is not billed in the demo and we do not quote a GPU price.
Frames from the four recorded cuts (the Watch replay on the site):

5. End-to-end checks
- Four samples rendered through the API: consent decisions signed (house voices and fictional cast allowed; Ines Okafor
refused with
project_not_in_scope, her line subtitled), MP4 with a C2PA credential, CSV, storyboard PDF, a record that verifies at/record/verify; the disclosure pre-flight OCR read the AI label on 3-6 of 6 sampled frames per cut. - Browser (private headless Playwright): samples, empty and oversized scripts (Run disabled with a message), a one-line
script, deep link
?sample=night-shift-ad&autorun=1, an edited action, a render (531 s, voice refused shown, label read on 6 of 6 frames), CSV and PDF downloads, a single-shot re-render (45 s), Watch replay with the video and 16 frames loading, no horizontal scroll at 390 px (after a fix). - Self-host: fresh clone, the api image plus the Kokoro layer, compose with named volumes, the local Qwen server (direct route) and ComfyUI: the rehearsal bundle passed 9 of 9 in 590 s (C2PA stamped) after a fix below; 0 script words in the container logs; torn down.
Bugs found and fixed: in a container, frames could not be moved from the work volume to the data volume (EXDEV; now
shutil.move, regression test); the CSV prefixed every empty cell with ' (test); the console stopped polling after one
failed request (now retries); the regulatory note's long links overflowed at 390 px (the Stack panel now breaks words);
the judge lacked the lines before a shot, so "she" could not be resolved, and lacked parentheticals (both added, test);
sketch styling blew out motion clips (now the same filter frame by frame).
Expected properties of the sample run (rehearsal kit, rehearsal/animatic-studio/)
POST /animatic/shotliston the Night Shift ad:coverage.pct= 100 andcoverage.dialogue_spoken_once= true.- Every
shots[].dialogue[].textequals the script's line text exactly. - At least 1 + (number of shots) receipts: one shot-list call and one grounding call per shot.
- Rendering with
{"VO": "id_house-af-heart", "DANA": "id_demo-ines-okafor"}: VO allowed, DANA refused (project_not_in_scope), each with adecision_id. - The finished version has
credential: "c2pa"and a record of more than 10 entries that verifies at/record/verify, with entries of origin human, ai and code. - The CSV export starts with
shot,tc_in,tc_outand has one row per shot.