Expert-to-SOP eval (27 Sep 2026)
Six narrated screen recordings of scripted procedures in three fictional desktop apps, run through the real pipeline on the
metered gateway route (Qwen3.8-27B with video input; the narration from MOSS-Transcribe-Diarize). Script:
scripts/eval_expert_to_sop.py; clips built by scripts/sop_clips/ (the apps, the procedures with their ground truth,
the Kokoro narration and the recorder). Results: docs/evals/expert-to-sop/results-v2.json and one run-*-v2.json per clip.
Data
- Recordings: headless Chromium drives a fictional app (a label printer tool, a press operator panel, an inventory app) with a visible cursor; 32–49 s each, 1280×720. The ground truth is the recorder's own timeline (start and end of each step).
- Narration: Kokoro-82M stock voice
am_michael(Apache-2.0; a synthetic voice, not a real person's), placed at each step's start, then transcribed by the diarizer like any upload (the eval uses the diarizer's transcript, errors included). - Planted in every recording: one step that is said but never done (e.g. "wear cut-resistant gloves", "check the tax ID against the W-9") and one that is done but never said (e.g. choosing the label size). 6–8 steps per recording.
- Split: dev =
labeldesk-add-printer,hmi-crimp-calibration(the two demo samples; used while writing the prompts). Test =labeldesk-return-label,hmi-recipe-change,tally-cycle-count,tally-new-supplier, run once with frozen prompts. - Licence: everything is synthetic and ours (CC0 recordings; the generator is AGPL-3.0-or-later with decosa-api). No real people, faces or voices.
Metrics
A drafted step matches a planted step when its text has one of the step's keywords and its time range overlaps the truth (±2 s; ±6 s for said-only steps). One-to-one, greedy by start time.
| Metric | Dev (2 recordings) | Test (4 recordings) |
|---|---|---|
| Planted steps done on screen found as seen steps | 14 / 14 | 21 / 25 (84%) |
| Found steps in the right order (pairs) | 100% | 100% |
| Cited start within 3 s of the truth | 14 / 14 | 19 / 21 (90.5%), median error 0.7 s |
| Said-but-not-done steps marked needs confirmation | 2 / 2 | 4 / 5 |
| Done-but-not-said steps marked not narrated | 2 / 2 | 3 / 4 |
| Said-and-done steps marked confirmed | 10 / 12 | 16 / 21 |
| Seen steps wrongly marked needs confirmation | 0 | 0 |
| Finer sub-steps (e.g. "type the PIN" and "click OK" for one planted step) | 7 | 9 |
| Invented steps (no planted action at that time) | 0 | 1 |
| Time per 40 s recording (all calls, gateway under load) | 17.2 s median | 18.5 s median |
| Receipts (all gateway-signed) | 27 | 43 |
What went wrong on the test split
- One recording's clock was wrong (
tally-cycle-count). The model listed all eight steps in the right order but cited them in 2-second slots (0:00, 0:02, 0:04…) across the first 19 s of a 32 s recording, instead of reading the times. Four of the "not found" steps, the missed done-not-said step and four of the five unconfirmed steps come from this one recording: with the wrong times, the narration around each step was the wrong narration. The same failure shows on long synthetic clips in the video-input eval. Not fixed by tuning (it was seen on the test split): the pipeline now emits a timing warning when the cited times cover under 60% of a recording over 20 s (it would have fired here and on no other recording), and every step has a keyframe at its cited time, so a reviewer sees the mismatch. - One said-only step was missed (
labeldesk-return-label): "Check the customer name and the ship date match the request" was not raised; the said-only pass treated it as covered by the search step next to it. That recording's other planted instruction ("photograph the damaged box") was flagged. - Two cited starts more than 3 s off (
tally-new-supplier: 3.7 s and 5.3 s, on the typing steps whose ranges the model drew wide). - One "invented" step (
labeldesk-return-label): "Click the Search button", which is real (the recorder clicks it within the planted search step) but falls just outside the matcher's window. Counted against us. - Unconfirmed said-and-done steps outside that recording: the safety guard toggled on and off in
hmi-recipe-change(the narration says "open the guard"; the model called both clicks unnarrated), and on dev the printer name, which the ASR heard as "DocB Thermal", came back as narration differs. That is the intended behaviour for a real mismatch, and here it is a speech-to-text error; the reviewer decides.
Changes made on dev (before the test split ran)
- Narration checks run four at a time (160 s → 17 s per recording).
- A said-only instruction that the second video look finds next to a drafted step is joined to that step instead of being added again (dev had "Give the printer a name…" duplicated beside "Type Dock B thermal").
- The metric separates finer sub-steps from invented ones (the first dev run counted 11 "invented" steps that were all real sub-actions). After the test run: the timing warning above (it changes no step or status).
Expected properties of the demo samples (for the rehearsal bundle)
rehearsal/expert-to-sop (the crimp calibration recording with its transcript) checks, against any install:
- the gloves instruction (said, never shown) comes back
needs_confirmation; - the offset typed on screen but never said comes back
not_narrated; - zeroing the height sensor comes back
confirmed; - every step has a keyframe, and there is no timing warning;
- at least six steps are drafted;
- the signed revision record verifies, and fails once changed; every model call has a signed receipt.
Passed 10/10 against the pre-release server (18 s). The smoke module (
scripts/smoke/expert-to-sop.py) runs the same sample: 14 steps, 16 signed receipts, 45,779 tokens, about $0.016 at list price, 24 s.
Limits
- Six short recordings of simple desktop tasks, made and labelled by the building agent; clean UI, a clear narrator and one planted discrepancy of each kind. Bench recordings (hands, tools, parts), long procedures, several speakers and real shop-floor audio are not measured.
- Steps come from what the video shows at one frame a second: a click that leaves no trace on screen for a second can be missed; past 4 minutes frames are sparser.
- The narration check is a model judging one step against nearby narration; it can call a paraphrase a conflict (and the ASR can mishear names).
- A draft for a named person to review. It does not judge whether the procedure is safe, correct or compliant.