Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: script to animatic (use case 50), 26 Sep 2026

Measured on our server with the branch API (the pre-release branch, port 8457): Qwen3.8-27B through the model gateway (receipted), Wan2.2-VACE-Fun-A14B in the shared studio ComfyUI on GPU0 (one job at a time, alongside other workloads' renders), Kokoro-82M on CPU. Script: scripts/animatic_eval.py; raw results in docs/evals/animatic-studio/ (shotlists.json, planted.json, consistency.json, consistency-frames.json). After these numbers the shot-list prompt gained one rule (original looks, with shape, colour and silhouette for non-humans); coverage was not re-measured with it. Development used the Last Lamp sample: the prompt rule against adding staging details and the context lines for the judge were added after a first run on it (26 Sep, 01:40), before these numbers. Nothing was tuned after them; one fix found here (parentheticals shown to the judge) went in afterwards and is not re-measured.

Data (all synthetic or public domain)

Script Kind Lines needing a shot Licence
The Last Lamp screenplay, 2 characters (one V.O.) 15 original, CC0
Hollin Oats: Night Shift 30 s ad script, VO and SUPER, made-up brand 11 original, CC0
Salvage game cutscene, pilot and robot 11 original, CC0
The Importance of Being Earnest, Act I (handbag scene) stage play, directions inside speeches 18 Wilde 1895, public domain (Gutenberg #844)
Trifles (opening) stage play, 4 speakers, long directions 23 Glaspell 1916, public domain in the US (Gutenberg #10623)

1. Shot-list coverage and invented dialogue (10 runs: 5 scripts x 2)

  • Coverage by the model's own list, before any code repair: 100% in 10 of 10 runs (78 required lines per pass). Code would attach a missed line to the nearest shot and mark it; it never had to.
  • Invalid line citations 0, shots out of script order 0, a dialogue line cited in two shots 0, truncated JSON 0.
  • Invented dialogue: 0 by construction. The model cites dialogue by line id and never writes it; code inserts the script's words. A quoted phrase in an action that is not in the cited lines: 0 in 10 runs. A human edit cannot add dialogue either: validate_shots rebuilds it from line ids (test).
  • Grounding verdicts on the model's shot actions (151): 128 supported, 17 partial (adds a detail, e.g. "on the table"), 3 unsupported, 3 contradicted. I read the 6 hard flags: 2 were real (Trifles: the action named the County Attorney for a line the Sheriff speaks; the ad: "Dana continues walking through the corridor" after she has walked out), 4 were over-strict ("Kettle speaks to Riva" twice, "pointing towards the rocker" where the stage direction says so, and "Dana speaks to herself", which the judge could not see because parentheticals were left out of its evidence; fixed).
  • Latency per script (one shot-list call plus one grounding call per shot, gateway under load): 7.9-29.6 s. Cost at the gateway list price ($0.30 / $1.50 per Mtok): $0.0035-0.013 per shot list (8-25 receipted calls).

2. Does the grounding check catch invented action? (planted)

Two shot actions per script were replaced with invented events ("A fire breaks out and Margit runs outside", "Police sirens wail as two officers burst through the door", ...): 10 plants over the 5 scripts. 10 of 10 were judged unsupported. The 66 unplanted shots in the same pass: 57 supported, 9 partial, 0 unsupported or contradicted. The plants are blunt on purpose; a subtle invention ("she smiles") lands in partial at best. Plants and checker have the same author.

3. Character consistency

18 single-character shots (6 each from Last Lamp, Salvage and Earnest), three arms, same seed per shot: unlocked = the frame prompt names the character only; locked = the hosted default, every frame repeats the character's look word for word; reference = locked plus the character sheet as a Wan2.2-VACE reference image (strength 0.5). Scored with CLIP ViT-L/14 on CPU, on the raw model frames (before the sketch filter): identity = cosine(frame, that character's sheet); ident-acc = the nearest sheet is the right character (Salvage and Earnest only, 12 frames; chance about 0.42); within = mean pairwise cosine between frames of the same character; adherence = 100 x CLIP text-image score against the shot's framing and action (without the look). Whole-frame CLIP is confounded by background and framing: read these as differences between arms, not absolute truth.

Arm identity ident-acc within adherence median s per frame (shared GPU)
unlocked 0.528 0.33 0.624 16.6 24
locked (default) 0.690 0.75 0.775 18.3 27
reference 0.812 1.00 0.845 14.9 42

Locking the look is what makes characters recognisable at all (identity +0.16, ident-acc 0.33 to 0.75). The reference arm scores highest on identity because it copies the sheet: same pose, same framing and a textured artefact background where the scene should be (see the grid), so adherence drops. It is not used; an edit model conditioned on the sheet is the "wanted" tier. Sheet, unlocked, locked and reference for four shots:

consistency arms

Seen in the frames and not in the scores: from a one-line description ("a squat maintenance robot with one round lens"), the robot KETTLE came out close to a well-known film robot. Those frames were withdrawn and are not published anywhere (site, Watch replay, this doc); the grid above leaves KETTLE out, though its two frames are still in the scores. The sample now spells KETTLE out (an orange kettle-shaped drone on six spider legs, one red lens on the spout; no treads, no eyes) and the shot-list prompt asks for original looks with shape, colour and silhouette for non-humans. The re-rendered Salvage cut shows a clearly original drone. Frames aren't checked against existing characters; review for resemblance before sharing.

4. Render time and cost

Full animatics through the API (studio queue; GPU0 shared with other workloads' renders throughout):

Sample Shots (motion) Animatic Render GPU minutes per minute of animatic
Last Lamp, sketch 15 (1) 49 s 359 s 7.3
Earnest, sketch 18 (0) 154 s 409 s 2.7
Salvage, colour (first design, withdrawn) 11 (1) 35 s 557 s 15.8
Salvage, colour (re-rendered, GPU busy) 12 (1) 39 s 1161 s 29.9
Night Shift ad, grey 7 (0) 25 s 726 s 28.5 (GPU busy with other jobs)
Night Shift ad, browser run 8 (0) 27 s 531 s 19.3
Night Shift ad, self-host container not recorded not recorded 576 s
  • Per frame: about 13 s with the model loaded and the GPU free (6 steps, cfg 2, 1280x720); 15-88 s per frame in these runs depending on other jobs. A motion clip (81 frames, 832x480): 75-81 s. Character sheets: 30-90 s each. Temp voices: 5-15 s on CPU. Assembly: 4-19 s. Re-rendering one edited shot and re-cutting: 47 s and 45 s (two runs).
  • Short scripts cost the most GPU time per minute of animatic (sheets and model loads are fixed costs).
  • Cost: the model calls ($0.0035-0.013 per script). The render is 6-12 minutes of a shared RTX PRO 6000; it is not billed in the demo and we do not quote a GPU price.

Frames from the four recorded cuts (the Watch replay on the site):

frames from cuts

5. End-to-end checks

  • Four samples rendered through the API: consent decisions signed (house voices and fictional cast allowed; Ines Okafor refused with project_not_in_scope, her line subtitled), MP4 with a C2PA credential, CSV, storyboard PDF, a record that verifies at /record/verify; the disclosure pre-flight OCR read the AI label on 3-6 of 6 sampled frames per cut.
  • Browser (private headless Playwright): samples, empty and oversized scripts (Run disabled with a message), a one-line script, deep link ?sample=night-shift-ad&autorun=1, an edited action, a render (531 s, voice refused shown, label read on 6 of 6 frames), CSV and PDF downloads, a single-shot re-render (45 s), Watch replay with the video and 16 frames loading, no horizontal scroll at 390 px (after a fix).
  • Self-host: fresh clone, the api image plus the Kokoro layer, compose with named volumes, the local Qwen server (direct route) and ComfyUI: the rehearsal bundle passed 9 of 9 in 590 s (C2PA stamped) after a fix below; 0 script words in the container logs; torn down.

Bugs found and fixed: in a container, frames could not be moved from the work volume to the data volume (EXDEV; now shutil.move, regression test); the CSV prefixed every empty cell with ' (test); the console stopped polling after one failed request (now retries); the regulatory note's long links overflowed at 390 px (the Stack panel now breaks words); the judge lacked the lines before a shot, so "she" could not be resolved, and lacked parentheticals (both added, test); sketch styling blew out motion clips (now the same filter frame by frame).

Expected properties of the sample run (rehearsal kit, rehearsal/animatic-studio/)

  1. POST /animatic/shotlist on the Night Shift ad: coverage.pct = 100 and coverage.dialogue_spoken_once = true.
  2. Every shots[].dialogue[].text equals the script's line text exactly.
  3. At least 1 + (number of shots) receipts: one shot-list call and one grounding call per shot.
  4. Rendering with {"VO": "id_house-af-heart", "DANA": "id_demo-ines-okafor"}: VO allowed, DANA refused (project_not_in_scope), each with a decision_id.
  5. The finished version has credential: "c2pa" and a record of more than 10 entries that verifies at /record/verify, with entries of origin human, ai and code.
  6. The CSV export starts with shot,tc_in,tc_out and has one row per shot.