Skip to content
decosa

76 · Field and trades · Any industry · live

Expert-to-SOP

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)

  • Planted on-screen steps found21 / 25test splitn = 25All four misses in one recording whose times came back in 2-second slots. Dev: 14 / 14.
  • Found steps in the right order100%test splitn = 21
  • Cited start within 3 s19 / 21test splitn = 21Median error 0.7 s
  • Said-but-not-shown steps flagged needs confirmation4 / 5test splitn = 5
  • Shown-but-not-said steps marked not narrated3 / 4test splitn = 4
  • Seen steps wrongly flagged0test splitn = 25
  • Invented steps1test splitA real click just outside the matcher's window; counted against us

Dataset

6 synthetic narrated screen recordings (32-49 s) of scripted procedures in three fictional desktop apps, a stock synthetic voice, one said-only and one shown-only step planted in each; 2 dev (the demo samples), 4 held out.

Caveats

  • The same author wrote the apps, the procedures, the ground truth and the prompts; the two dev recordings are the demo samples.
  • Screen recordings of simple desktop tasks with a clean synthetic voice; bench footage, long procedures and real speech are not measured.
  • One held-out recording's times came back in 2-second slots; a timing warning was added after the test run and changes no step or status.
  • A draft for a named reviewer; it does not judge whether a procedure is safe or correct.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
27 Sep 2026
Latency, this run
n/a
p50 over passed runs
20 s
Receipts
16
Model calls
n/a
Tokens
n/a
Cost per run
$0.016

Self-host verification

Verified on 27 Sep 2026: Fresh clone of the pre-release branch into a clean directory, docker build of the api image (32 s; ffmpeg 7.1.5 inside), the api with a named volume on host networking against the running local vLLM (Qwen3.8-27B with video input) and diarizer on the direct route; then torn down.

The rehearsal bundle passed 10 of 10 in 12.3 s; the labeldesk sample with speech to text inside the container drafted 12 steps (1 needs confirmation) in 18.6 s with 14 attested receipts and a model-call receipt for the ASR; the revision record verified; no step, narration or title text in the logs. The local vLLM was the production unit with the 32k video-token override, not the compose default (12,288); the model server's own startup was not re-verified (no new GPU load).

Rehearsal bundle: expert-to-sop.zip (693 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route, live diarizer); production gets this tool when the branch merges.
  • Measured on six synthetic screen recordings made and labelled by the building agent; bench footage and real narration are not measured.
  • On 1 of 4 held-out recordings the model cited times in 2-second slots instead of reading them; the timing warning catches that case, the times themselves are not corrected.
  • The narration check can call a paraphrase or a speech-to-text error a difference (the printer name heard as "DocB Thermal").

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Watches the recording (one video part, sampled at 1 frame a second) and lists each step with its time range; judges each step against the narration near its time (the grounding block's judge); lists instructions said but not drafted, then checks them back against the videoQwen3.8-27B (NVIDIA NVFP4), video inputApache-2.0
  • The narration as timed lines, with a model-call receipt (audio hash in, transcript hash out) signed by the instanceMOSS-Transcribe-Diarize 0.9BApache-2.0
  • Clip preparation (same timeline, no audio, at most 1280 px), keyframes at each cited time, statuses, sign-off and the signed revision record (no model; CPU)decosa-api video block and SOP module (decosa_api/video, decosa_api/verticals/sop)AGPL-3.0-or-later

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · steps and keyframes, no narration check (1)
  • held-out test, 4 recordings: planted on-screen steps found / in order / start within 3 s: 21 of 25 / 100% / 19 of 21 (the draft is the same call as Standard)decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route
Standard · steps checked against the narration (hosted demo) (4)
  • held-out test, 4 recordings: said-but-not-shown steps marked needs confirmation / shown-but-not-said marked not narrated: 4 of 5 / 3 of 4decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route
  • held-out test: said-and-shown steps confirmed / seen steps wrongly flagged / invented steps: 16 of 21 / 0 / 1decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route
  • held-out test: recordings where the model cited times in 2-second slots instead of reading them (a timing warning is shown): 1 of 4decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route
  • demo samples (dev, used while writing the prompts): planted steps found / flags as planted: 14 of 14 / 4 of 4decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route

How we measure · All tools