Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Expert-to-SOP eval (27 Sep 2026)

Six narrated screen recordings of scripted procedures in three fictional desktop apps, run through the real pipeline on the metered gateway route (Qwen3.8-27B with video input; the narration from MOSS-Transcribe-Diarize). Script: scripts/eval_expert_to_sop.py; clips built by scripts/sop_clips/ (the apps, the procedures with their ground truth, the Kokoro narration and the recorder). Results: docs/evals/expert-to-sop/results-v2.json and one run-*-v2.json per clip.

Data

  • Recordings: headless Chromium drives a fictional app (a label printer tool, a press operator panel, an inventory app) with a visible cursor; 32–49 s each, 1280×720. The ground truth is the recorder's own timeline (start and end of each step).
  • Narration: Kokoro-82M stock voice am_michael (Apache-2.0; a synthetic voice, not a real person's), placed at each step's start, then transcribed by the diarizer like any upload (the eval uses the diarizer's transcript, errors included).
  • Planted in every recording: one step that is said but never done (e.g. "wear cut-resistant gloves", "check the tax ID against the W-9") and one that is done but never said (e.g. choosing the label size). 6–8 steps per recording.
  • Split: dev = labeldesk-add-printer, hmi-crimp-calibration (the two demo samples; used while writing the prompts). Test = labeldesk-return-label, hmi-recipe-change, tally-cycle-count, tally-new-supplier, run once with frozen prompts.
  • Licence: everything is synthetic and ours (CC0 recordings; the generator is AGPL-3.0-or-later with decosa-api). No real people, faces or voices.

Metrics

A drafted step matches a planted step when its text has one of the step's keywords and its time range overlaps the truth (±2 s; ±6 s for said-only steps). One-to-one, greedy by start time.

Metric Dev (2 recordings) Test (4 recordings)
Planted steps done on screen found as seen steps 14 / 14 21 / 25 (84%)
Found steps in the right order (pairs) 100% 100%
Cited start within 3 s of the truth 14 / 14 19 / 21 (90.5%), median error 0.7 s
Said-but-not-done steps marked needs confirmation 2 / 2 4 / 5
Done-but-not-said steps marked not narrated 2 / 2 3 / 4
Said-and-done steps marked confirmed 10 / 12 16 / 21
Seen steps wrongly marked needs confirmation 0 0
Finer sub-steps (e.g. "type the PIN" and "click OK" for one planted step) 7 9
Invented steps (no planted action at that time) 0 1
Time per 40 s recording (all calls, gateway under load) 17.2 s median 18.5 s median
Receipts (all gateway-signed) 27 43

What went wrong on the test split

  • One recording's clock was wrong (tally-cycle-count). The model listed all eight steps in the right order but cited them in 2-second slots (0:00, 0:02, 0:04…) across the first 19 s of a 32 s recording, instead of reading the times. Four of the "not found" steps, the missed done-not-said step and four of the five unconfirmed steps come from this one recording: with the wrong times, the narration around each step was the wrong narration. The same failure shows on long synthetic clips in the video-input eval. Not fixed by tuning (it was seen on the test split): the pipeline now emits a timing warning when the cited times cover under 60% of a recording over 20 s (it would have fired here and on no other recording), and every step has a keyframe at its cited time, so a reviewer sees the mismatch.
  • One said-only step was missed (labeldesk-return-label): "Check the customer name and the ship date match the request" was not raised; the said-only pass treated it as covered by the search step next to it. That recording's other planted instruction ("photograph the damaged box") was flagged.
  • Two cited starts more than 3 s off (tally-new-supplier: 3.7 s and 5.3 s, on the typing steps whose ranges the model drew wide).
  • One "invented" step (labeldesk-return-label): "Click the Search button", which is real (the recorder clicks it within the planted search step) but falls just outside the matcher's window. Counted against us.
  • Unconfirmed said-and-done steps outside that recording: the safety guard toggled on and off in hmi-recipe-change (the narration says "open the guard"; the model called both clicks unnarrated), and on dev the printer name, which the ASR heard as "DocB Thermal", came back as narration differs. That is the intended behaviour for a real mismatch, and here it is a speech-to-text error; the reviewer decides.

Changes made on dev (before the test split ran)

  1. Narration checks run four at a time (160 s → 17 s per recording).
  2. A said-only instruction that the second video look finds next to a drafted step is joined to that step instead of being added again (dev had "Give the printer a name…" duplicated beside "Type Dock B thermal").
  3. The metric separates finer sub-steps from invented ones (the first dev run counted 11 "invented" steps that were all real sub-actions). After the test run: the timing warning above (it changes no step or status).

Expected properties of the demo samples (for the rehearsal bundle)

rehearsal/expert-to-sop (the crimp calibration recording with its transcript) checks, against any install:

  1. the gloves instruction (said, never shown) comes back needs_confirmation;
  2. the offset typed on screen but never said comes back not_narrated;
  3. zeroing the height sensor comes back confirmed;
  4. every step has a keyframe, and there is no timing warning;
  5. at least six steps are drafted;
  6. the signed revision record verifies, and fails once changed; every model call has a signed receipt. Passed 10/10 against the pre-release server (18 s). The smoke module (scripts/smoke/expert-to-sop.py) runs the same sample: 14 steps, 16 signed receipts, 45,779 tokens, about $0.016 at list price, 24 s.

Limits

  • Six short recordings of simple desktop tasks, made and labelled by the building agent; clean UI, a clear narrator and one planted discrepancy of each kind. Bench recordings (hands, tools, parts), long procedures, several speakers and real shop-floor audio are not measured.
  • Steps come from what the video shows at one frame a second: a click that leaves no trace on screen for a second can be missed; past 4 minutes frames are sparser.
  • The narration check is a model judging one step against nearby narration; it can call a paraphrase a conflict (and the ASR can mishear names).
  • A draft for a named person to review. It does not judge whether the procedure is safe, correct or compliant.