Skip to content
decosa
LabsHostedSelf-host

Turn an expert video into an SOP

A draft SOP whose every step cites a time range and a keyframe, with steps said but not shown, or shown but not said, flagged for a reviewer.

Held-out test21 / 25On-screen steps found in test recordings (held-out test)
On production20 smedian on production (2026-09-27); slower when the service is busy
List price~$0.016 per recordingmeasured, at list price

Built on: Video understanding, Speaker diarization, Grounding, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the expert-to-sop API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on Qwen3.8-27B with video input and MOSS-Transcribe-Diarize: one 96 GB GPU (measured on two shared ones); ffmpeg, keyframes and the record on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
expert-to-sop

Use the hosted API

# Decosa Expert-to-SOP: use the hosted API

You are wiring Decosa's Expert-to-SOP into this project (a training portal, a work-instruction library or an internal
runbook tool). It takes a recording of an expert doing a procedure (screen or bench, with narration) and returns a draft
SOP: each step cites a time range and a keyframe from the recording and is checked against what was said; steps said
but never shown come back as "needs confirmation". A named reviewer decides the flagged steps and signs the revision into
a record anyone can verify. Use only what is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- It is a draft for a qualified person to review, not an approved procedure and not safety advice. Never label an SOP
  "approved" or "compliant" from this output alone.
- Send recordings you have the right to upload, made with the knowledge of the people in them. Confidential recordings
  belong on a self-hosted box.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page, kept in `DECOSA_API_KEY`, never in code.
   Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "expert-to-sop"}` returns
   `{"token", "expires_at", "budget"}`. Sessions per IP are limited (HTTP 429 with `Retry-After`); a demo token runs one
   draft at a time (409).
3. A draft needs about 7,700 generated tokens of budget before it starts (402 otherwise); a 40 s recording uses far less.

## Endpoints
- `POST /sop/draft` (token). Body: `{"video_b64": "<MP4, MOV or WebM as base64>" | "sample": "<id>", "title"?: "...", "transcript"?: [{"start": 2.1, "end": 3.8, "text": "..."}], "stream"?: true}`.
  - Recordings up to 64 MB and 10 minutes. Without `transcript`, the narration is transcribed on Decosa's hosted service (a
    model-call receipt). The model sees the recording at one frame a second (one frame every 2.5 s past 4 minutes).
  - JSON response: `{run_id, status: "draft_awaiting_signoff", totals, report, budget}`. `report.steps`:
    `[{n, action, start_s, end_s, range, evidence, seen, status: "confirmed"|"not_narrated"|"narration_differs"|"needs_confirmation", narration: {verdict, quote, reason}, keyframe: {t_s, sha256, jpeg_b64}}]`;
    `report.timing_warning` is a string when the cited times look wrong (check the keyframes), else null.
  - With `Accept: text/event-stream` (or `"stream": true`): `video`, `transcript`, `receipt`, `draft`, `step`,
    `said_only`, `timing_warning`, `keyframes`, `report`, `done`, `budget`.
- `POST /sop/runs/{run_id}/signoff` (same token) `{"name", "role", "revision"?: "1", "decisions"?: {"<step n>": "keep"|"confirmed_by_expert"|"remove"}, "default_decision"?, "confirm": true}`
  → `{record, check, signoff}`. Every step marked needs_confirmation or narration_differs must be decided. Runs are kept for one hour.
- `GET /sop/runs/{run_id}/export?format=md|json|record`, `POST /record/verify` `{"record"}` (no token),
  `GET /sop/info`, `GET /sop/samples`, `GET /sop/samples/{id}/video`, `GET /attest/signing-key`.

## Example: draft, review the flagged steps, sign (Python, `pip install httpx`)
```python
import base64, httpx, json, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
video = base64.b64encode(open("procedure.mp4", "rb").read()).decode()
r = httpx.post(f"{API}/sop/draft", json={"video_b64": video, "title": "Calibrate the crimp press"}, headers=H, timeout=600).json()
rep = r["report"]
if rep["timing_warning"]:
    print("check the keyframes:", rep["timing_warning"])
decisions = {}
for s in rep["steps"]:
    print(s["n"], s["range"], s["status"], s["action"])
    if s["status"] in ("needs_confirmation", "narration_differs"):
        decisions[str(s["n"])] = input(f"step {s['n']}: keep / confirmed_by_expert / remove? ").strip()
signed = httpx.post(f"{API}/sop/runs/{r['run_id']}/signoff", headers=H, json={"name": "A. Reviewer", "role": "CI lead",
                    "revision": "1", "decisions": decisions, "confirm": True}).json()
open("sop.md", "w").write(httpx.get(f"{API}/sop/runs/{r['run_id']}/export?format=md", headers=H).text)
json.dump(signed["record"], open("sop-revision-1.record.json", "w"))   # anyone can re-check it at /record/verify
```

## Receipts
Every model call on the hosted route gets a gateway-signed receipt; the video calls' request hash covers the video's
sha256 and the sampling settings. The speech to text gets a model-call receipt signed by Decosa's instance key. The
revision record lists them all and holds hashes only.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Expert-to-SOP: run it yourself (containers)

You are setting up Decosa's Expert-to-SOP on this machine, so recordings of internal systems and shop floors never leave
it. It turns a narrated recording of a procedure into a draft SOP whose steps cite a time range and a keyframe, checked
against what was said, and seals a named reviewer's sign-off into a signed revision record. Nothing is sent to Decosa's
hosted API. It drafts; it does not approve procedures.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/expert-to-sop.zip (693 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py expert-to-sop` (the api image carries the same bundle under /app/rehearsal/expert-to-sop/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py expert-to-sop --bundle expert-to-sop.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the gloves instruction (said, never shown) needs confirmation", "the offset typed on screen but never said is marked not narrated", "zeroing the height sensor is seen and said"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM; it must run with `--limit-mm-per-prompt '{"image":4,"video":1}'`
   and `--media-io-kwargs '{"video":{"num_frames":240,"fps":1}}'`, not `--language-model-only`) and the `api` service.
   For the `api` service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`
   and bind every port to 127.0.0.1. For narration, keep a `diarize` service if the compose file has one and set
   `DECOSA_DIARIZE_URL=http://diarize:8092`; otherwise send the transcript with each recording. Never set the gateway
   route on this box.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Smoke test: get a token with `POST /demo/session {"vertical":"expert-to-sop"}` and send
   `{"sample": "hmi-crimp-calibration", "stream": false}` to `POST /sop/draft`. Expect 8-14 steps, each with a keyframe;
   the gloves instruction `needs_confirmation`; the offset step `not_narrated`; no `timing_warning`; every receipt
   `"attested"` (signed by this box). Then `POST /sop/runs/<run_id>/signoff` with
   `{name, role, default_decision: "remove", confirm: true}` and `POST /record/verify` with the record: `ok` must be true.
5. Report back: `GET /attest/signing-key` (the public key a reviewer pins), the smoke-test results and how long the draft took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
confidential recordings. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVIDIA NVFP4), video input needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (61.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (61.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API, the speech model and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with both models loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "asr": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"expert-to-sop"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py expert-to-sop

Download the mock-data bundle (693 KB, 10 checks)expected.json

A 40-second screen recording of a fictional press operator panel: maintenance mode with a PIN, zero the height sensor, three test crimps, a height offset, save, and a log note, narrated by a stock synthetic voice. The narrator also says to wear cut-resistant gloves, which the recording never shows, and never mentions typing the offset, which it does show. The draft must mark the gloves instruction as needs confirmation, the offset step as seen but not narrated, the zeroing step as seen and said, cite a keyframe for every step, and keep the cited times inside the recording. A reviewer signs; the revision record must verify, and fail once a step count is changed.

What the rehearsal checks
  • the gloves instruction (said, never shown) needs confirmation
  • the offset typed on screen but never said is marked not narrated
  • zeroing the height sensor is seen and said
  • every step has a keyframe
  • the cited times stay inside the recording (no timing warning)
  • at least six steps are drafted
  • an empty request is refused
  • the signed revision record verifies
  • the record fails once it is changed
  • every model call has a signed receipt

Licence: Synthetic: a fictional app written for Decosa, driven by a script and narrated by the Kokoro-82M stock voice am_michael (Apache-2.0). No real people, faces or voices. Recording CC0; part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa Expert-to-SOP: run it yourself (containers)

You are setting up Decosa's Expert-to-SOP on this machine, so recordings of internal systems and shop floors never leave
it. It turns a narrated recording of a procedure into a draft SOP whose steps cite a time range and a keyframe, checked
against what was said, and seals a named reviewer's sign-off into a signed revision record. Nothing is sent to Decosa's
hosted API. It drafts; it does not approve procedures.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/expert-to-sop.zip (693 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py expert-to-sop` (the api image carries the same bundle under /app/rehearsal/expert-to-sop/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py expert-to-sop --bundle expert-to-sop.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the gloves instruction (said, never shown) needs confirmation", "the offset typed on screen but never said is marked not narrated", "zeroing the height sensor is seen and said"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM; it must run with `--limit-mm-per-prompt '{"image":4,"video":1}'`
   and `--media-io-kwargs '{"video":{"num_frames":240,"fps":1}}'`, not `--language-model-only`) and the `api` service.
   For the `api` service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`
   and bind every port to 127.0.0.1. For narration, keep a `diarize` service if the compose file has one and set
   `DECOSA_DIARIZE_URL=http://diarize:8092`; otherwise send the transcript with each recording. Never set the gateway
   route on this box.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Smoke test: get a token with `POST /demo/session {"vertical":"expert-to-sop"}` and send
   `{"sample": "hmi-crimp-calibration", "stream": false}` to `POST /sop/draft`. Expect 8-14 steps, each with a keyframe;
   the gloves instruction `needs_confirmation`; the offset step `not_narrated`; no `timing_warning`; every receipt
   `"attested"` (signed by this box). Then `POST /sop/runs/<run_id>/signoff` with
   `{name, role, default_decision: "remove", confirm: true}` and `POST /record/verify` with the record: `ok` must be true.
5. Report back: `GET /attest/signing-key` (the public key a reviewer pins), the smoke-test results and how long the draft took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
confidential recordings. If I ask for it later, follow the Provide page instead of improvising.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsExpert-to-SOP on GeForce RTX 5090: use the Standard · steps checked against the narration (hosted demo) tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB: Not measured. The weights are about 20 GB; a maximum video request is 32k prompt tokens of KV cache.

Standard · steps checked against the narration (hosted demo): what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Watches the recording: Qwen3.8-27B (NVIDIA NVFP4), video input. ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)
  • The narration as timed lines, with a model-ca...: MOSS-Transcribe-Diarize 0.9B. ~4 GB, weights 1.8 GB (estimate). MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate.
  • Clip preparation: decosa-api video block and SOP module (decosa_api/video, decosa_api/verticals/sop). CPU. Runs on CPU (vram_gb 0 in stack.json).

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Expert-to-SOP, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Expert-to-SOP on my hardware

Fetch https://decosa.ai/prompts/expert-to-sop-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=expert-to-sop)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · steps checked against the narration (hosted demo) (standard). Fit check: runs with changes, about 32 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Watches the recording: Qwen3.8-27B (NVIDIA NVFP4), video input (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- The narration as timed lines, with a model-ca...: MOSS-Transcribe-Diarize 0.9B (OpenMOSS-Team/MOSS-Transcribe-Diarize), 4 GB
- Clip preparation: decosa-api video block and SOP module (decosa_api/video, decosa_api/verticals/sop), CPU

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4), video input ~28 GB (88%), MOSS-Transcribe-Diarize 0.9B ~4 GB (13%); about 0 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/expert-to-sop-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 27 Sep 2026 · measured 27 Sep 2026: · p50 20 s · ~$0.016 per run · 16 receipts

Loading the nightly status…

Self-host: verified 27 Sep 2026 · Fresh clone of the pre-release branch into a clean directory, docker build of the api image (32 s; ffmpeg 7.1.5 inside), the api with a named volume on host networking against the running local vLLM (Qwen3.8-27B with video input) and diarizer on the direct route; then torn down.

Measured cost to run: about $0.016 per recording (hosted, 27 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The rehearsal bundle passed 10 of 10 in 12.3 s; the labeldesk sample with speech to text inside the container drafted 12 steps (1 needs confirmation) in 18.6 s with 14 attested receipts and a model-call receipt for the ASR; the revision record verified; no step, narration or title text in the logs. The local vLLM was the production unit with the 32k video-token override, not the compose default (12,288); the model server's own startup was not re-verified (no new GPU load).

Known limits (4)
  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route, live diarizer); production gets this tool when the branch merges.
  • Measured on six synthetic screen recordings made and labelled by the building agent; bench footage and real narration are not measured.
  • On 1 of 4 held-out recordings the model cited times in 2-second slots instead of reading them; the timing warning catches that case, the times themselves are not corrected.
  • The narration check can call a paraphrase or a speech-to-text error a difference (the printer name heard as "DocB Thermal").

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the expert-to-sop API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on Qwen3.8-27B with video input and MOSS-Transcribe-Diarize: one 96 GB GPU (measured on two shared ones); ffmpeg, keyframes and the record on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

A recording of an expert doing a procedure, turned into a draft SOP whose steps cite a time range and a keyframe, checked against what they said.

An expert records the screen or the bench while doing a procedure and talking. Qwen3.8-27B watches the recording and lists each step it sees with its time range; a keyframe is cut at each cited moment; every step is checked against the narration. Steps said but never shown come back as needs confirmation, and steps shown but never said are marked. A named reviewer decides the flagged steps and signs the revision into a record anyone can verify. For CI, training and quality leads who keep work instructions current.

Deployment
Hosted or self-host
Regulatory
Not legal or safety advice; a draft for a qualified person to review, not a validated or approved procedure. Checked against the primary text on 27 Sep 2026: where OSHA's control of hazardous energy standard applies, 29 CFR 1910.147(c)(4)(i) says procedures 'shall be developed, documented and utilized', (c)(4)(ii) that they 'shall clearly and specifically outline the scope, purpose, authorization, rules, and techniques', and (c)(6)(i) requires a periodic inspection of each procedure at least annually (https://www.osha.gov/laws-regs/regulations/standardnumber/1910/1910.147). A drafted SOP does not meet those duties by itself: the employer writes, authorizes and inspects the procedure. Recording people at work can be personal data (for example under the GDPR where workers are identifiable); record with the knowledge of the people in the recording and keep faces out where you can. The demo recordings are synthetic.
Architecture
Text description

A recording goes to decosa-api. ffmpeg makes a clip with the same timeline and no audio, and extracts the audio. MOSS-Transcribe-Diarize turns the audio into timed narration lines with a model-call receipt. Qwen3.8-27B (Apache-2.0) watches the clip through our gateway and lists steps with time ranges; each step is judged against nearby narration; instructions said but not drafted are checked back against the video. ffmpeg cuts a keyframe at each cited time. The draft SOP shows each step's status; a reviewer decides flagged steps and signs a revision record. Gateway receipts carry the video's sha256. Self-hosted, everything stays on the box.

Architecture

At a glance

What it does
Turns one narrated recording of a procedure into numbered steps, each with a time range, a keyframe and a status: seen and said, seen but not narrated, narration differs, or said but not shown (needs confirmation). A named reviewer decides every flagged step and signs the revision; exports in Markdown and JSON.
What it does not do
It does not judge whether a procedure is safe, correct or compliant, and a draft is not an approved procedure. It does not hear sound effects or read audio cues other than speech, merge several recordings, or diff revisions. Recordings over 4 minutes are sampled more sparsely (one frame every 2.5 s at 10 minutes), so short actions can be missed.
Data retention
The recording and its clip are deleted when the run ends. The report, with its keyframes, is kept in memory for one hour for the token or key that made it. Logs carry run ids and counts, never titles, narration or step text. You keep the exports and the signed record, which holds hashes only.
What leaves the box (hosted demo)
The clip (no audio) and the narration text go to Qwen3.8-27B through our gateway, which Decosa operates; the audio is transcribed on Decosa's hosted service. The gateway's receipts hold hashes, not content. Self-hosted, nothing leaves.
Accuracy
On 4 held-out synthetic screen recordings: 21 of 25 planted on-screen steps found, all in order; 4 of 5 said-but-not-shown steps flagged; no seen step wrongly flagged. On 1 of the 4 the model cited times in 2-second slots instead of reading them; the draft then shows a timing warning.
Cost per recording
A few cents or less for a short recording at gateway list price (measured on the demo sample). Each run shows its own measured cost.
Output
A draft SOP with steps, time ranges, keyframes and statuses; after sign-off a decosa.record.v1 revision record with the recording hash, each step's hashes, the receipts and the reviewer (verify at /record/verify).
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    steps and keyframes, no narration check

    The steps the video shows, each with a time range and a keyframe, and the sign-off record. No speech to text, so nothing is checked against what was said and said-only steps are not found.

    Models
    • Qwen3.8-27B (NVIDIA NVFP4), video input
    • decosa-api video block and SOP module (decosa_api/video, decosa_api/verticals/sop)
    Hardware
    1x RTX PRO 6000 96 GB (measured)
    Quality evidence
    • held-out test, 4 recordings: planted on-screen steps found / in order / start within 3 s21 of 25 / 100% / 19 of 21 (the draft is the same call as Standard)decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route
    Latency
    measured: the draft call alone takes seconds for a short recording
    Verification
    Proof: strongSelf-host onlyGateway receipt per call on the hosted route.
  • In the hosted demo

    Standard

    steps checked against the narration (hosted demo)

    Speech to text for the narration, the steps from the video, each step judged against what was said near its time, and instructions said but not shown checked back against the video and marked needs confirmation.

    Models
    • Qwen3.8-27B (NVIDIA NVFP4), video input
    • MOSS-Transcribe-Diarize 0.9B
    • decosa-api video block and SOP module (decosa_api/video, decosa_api/verticals/sop)
    Hardware
    1x RTX PRO 6000 96 GB (measured); the diarizer on a second GPU in the hosted setup
    Quality evidence
    • held-out test, 4 recordings: said-but-not-shown steps marked needs confirmation / shown-but-not-said marked not narrated4 of 5 / 3 of 4decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route
    • held-out test: said-and-shown steps confirmed / seen steps wrongly flagged / invented steps16 of 21 / 0 / 1decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route
    • held-out test: recordings where the model cited times in 2-second slots instead of reading them (a timing warning is shown)1 of 4decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route
    • demo samples (dev, used while writing the prompts): planted steps found / flags as planted14 of 14 / 4 of 4decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route
    Latency
    measured: under half a minute and a few cents or less per short recording at list price
    Verification
    Proof: strongGateway receipt per model call (video calls bind the video's sha256 and sampling); model-call receipt for the speech to text; signed revision record.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Watches the recording (one video part, sampled at 1 frame a second) and lists each step with its time range; judges each step against the narration near its time (the grounding block's judge); lists instructions said but not drafted, then checks them back against the videoQwen3.8-27B (NVIDIA NVFP4), video inputnvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
The narration as timed lines, with a model-call receipt (audio hash in, transcript hash out) signed by the instanceMOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab)
0.9BProof: partialIn the hosted demo
Clip preparation (same timeline, no audio, at most 1280 px), keyframes at each cited time, statuses, sign-off and the signed revision record (no model; CPU)decosa-api video block and SOP module (decosa_api/video, decosa_api/verticals/sop)
0 GBProof: partialIn the hosted demo

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    GET /sop/info, /sop/samples; POST /sop/draft (SSE or JSON), POST /sop/runs/{id}/signoff, GET /sop/runs/{id}/export?format=md|json|record. No GPU; ffmpeg inside. The upload is deleted when the run ends; reports are kept in memory for one hour.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B with image and video input. Internal to the compose network.

  • decosa-diarize:8092

    MOSS-Transcribe-Diarize for the narration (DECOSA_DIARIZE_URL). No published image yet; built from services/diarize. Without it, send the transcript with the recording.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured: the hosted Qwen3.8-27B with video runs on one of these cards on our server (KV cache 60.17 GiB after video was turned on), the diarizer on the other.

  • 1x RTX 5090 32 GB

    Not measured. The weights are about 20 GB; a maximum video request is 32k prompt tokens of KV cache.

  • CPU only Does not fit

    The steps come from the video model. Clip preparation, keyframes, the record and verification run on CPU.

Latency per lane

  • 40 s narrated screen recording, full draft (speech to text, video, narration checks), hosted gateway route20.0 s

    Measuredmeasured on our server 2026-09-27: 16-21 s on the pre-release server and in the eval (median 17.2 s dev, 18.5 s test) while other workloads used the gateway

  • the video call alone, 120 s 1280x720 clip at the cap (32k tokens)9.8 s

    Measuredmeasured on our server 2026-09-27, direct to vLLM; 17-19 s each with four at once

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

expert-to-sop/assemble-prompt.md134 lines
# Assemble Decosa Expert-to-SOP on this machine

You are setting up Expert-to-SOP: it takes a recording of an expert doing a procedure (screen or bench, with narration)
and drafts an SOP. Each step cites a time range and a keyframe from the recording and is checked against what the expert
said; steps said but never shown come back as "needs confirmation". A named reviewer signs each revision into a record
signed by this box's own key. Work step by step, show me each command before you run anything with `sudo`, and stop to
ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/expert-to-sop.zip (693 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py expert-to-sop` (the api image carries the same bundle under /app/rehearsal/expert-to-sop/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py expert-to-sop --bundle expert-to-sop.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the gloves instruction (said, never shown) needs confirmation", "the offset typed on screen but never said is marked not narrated", "zeroing the height sensor is seen and said"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Models: Qwen3.8-27B (Apache-2.0) with video input, and MOSS-Transcribe-Diarize (Apache-2.0) for the narration. The
  API is decosa-api (AGPL-3.0-or-later); ffmpeg runs inside it as a separate program (LGPL/GPL).
- Recordings stay on this machine. Bind every port to 127.0.0.1. The API deletes each upload when its run ends and keeps
  reports in memory for one hour; logs carry counts only. Keep it that way.
- Be honest about what it does: it drafts from one recording for a person to review. It does not judge whether a
  procedure is safe or correct. Record people only with their knowledge.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 is about 20 GB of weights; a maximum video request adds
   about 32k tokens of KV cache). Measured on an RTX PRO 6000 96 GB. Blackwell cards run NVFP4; on older cards use the
   FP8 weights. The diarizer needs about 3 GB more (same card or a second one).
2. `docker --version`, `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them from
   the official Docker and NVIDIA repositories after asking me, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 35 GB free (model weights, images).

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` and `${DECOSA_REGISTRY}/decosa-llm:<tag>` (**publishing soon**). If a
  pull fails, build from source: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a
  release that contains `decosa_api/verticals/sop/`, and build `docker/api/Dockerfile` as `decosa-api:local` and
  `docker/llm` as `decosa-llm:local` (vLLM 0.29.0, pinned by digest).
- The diarizer has no published image yet: build `services/diarize` from the same checkout as `decosa-diarize:local`
  (weights `OpenMOSS-Team/MOSS-Transcribe-Diarize` at revision `704aa4a9c304e8520be88901e0d1960158ef5b15`). Without
  it, send the transcript with each recording (step 4.4).
- Weights `nvidia/Qwen3.8-27B-NVFP4` at revision `482ca0f3832238542f8f5295dde86b5f22711d80` download on the llm's
  first start into the named `hf-cache` volume.

## 3. docker-compose.yml
Write this in `~/decosa/sop/`:

```yaml
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:<tag>
    command: ["nvidia/Qwen3.8-27B-NVFP4", "--revision", "482ca0f3832238542f8f5295dde86b5f22711d80",
              "--served-model-name", "qwen3.8-27b", "--limit-mm-per-prompt", "{\"image\":4,\"video\":1}",
              "--media-io-kwargs", "{\"video\":{\"num_frames\":240,\"fps\":1}}", "--max-model-len", "65536",
              "--gpu-memory-utilization", "0.60", "--kv-cache-dtype", "fp8_e4m3", "--host", "0.0.0.0", "--port", "8000"]
    ipc: host
    volumes: ["hf-cache:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], interval: 15s, retries: 60 }
  diarize:
    image: decosa-diarize:local
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>
    ports: ["127.0.0.1:8445:8445"]
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_DIARIZE_URL: http://diarize:8092
    volumes: ["decosa-data:/data"]
    depends_on: { llm: { condition: service_healthy } }
    healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/sop/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
  decosa-data:
  hf-cache:
```

The api keeps its state (keys, receipts, this box's signing key) in the named volume `decosa-data`, not a host folder:
the image runs as uid 10001, and a host folder Docker creates is owned by root (`PermissionError ... /data/keys.sqlite`).
Start with `docker compose up -d`; the llm's first start downloads the weights and takes several minutes.

Video budget: with these flags the model's own processor caps one video at 12,288 video tokens. The hosted service raises
it to 32,768 by mounting a copy of the model's `processor_config.json` and `video_preprocessor_config.json` whose
`video_processor.size.longest_edge` is 67108864 (read-only, over the files in the model folder). Do that only with a
local model folder and room for the extra KV cache; do not use `--mm-processor-kwargs` for it (vLLM 0.29.0 fails at start).

On the first start the api creates this box's Ed25519 key in the volume (`/data/attest/`, mode 0600). Back it up with
`docker compose cp api:/data/attest ./attest-backup`, keep it private, never print it.

## 4. Smoke test
1. `curl -s localhost:8445/sop/info | jq '.serving, .steps.asr'`: 1 video, 1 fps, 240 frames; ASR "on".
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"expert-to-sop"}' | jq -r .token)`.
3. Draft the crimp calibration sample:
   `curl -s -XPOST localhost:8445/sop/draft -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample":"hmi-crimp-calibration","stream":false}' > draft.json`.
   Expect 8-14 steps with a keyframe each; the gloves instruction `needs_confirmation`; the offset step `not_narrated`;
   zeroing the sensor `confirmed`; no `timing_warning`. On our card it took about 20 s.
4. Your own recording without the diarizer: add `"transcript": [{"start": 2.1, "end": 3.8, "text": "..."}]` (seconds) and
   `"video_b64"` (the file as base64) instead of `"sample"`. MP4, MOV or WebM, up to 10 minutes and 64 MB.
5. Sign off: `curl -s -XPOST localhost:8445/sop/runs/$(jq -r .run_id draft.json)/signoff -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"name":"Your Name","role":"CI lead","default_decision":"remove","confirm":true}' > signed.json`,
   then `jq '{record}' signed.json | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @-`
   must say `"ok": true`; change one step's status in the record and verify again: it must fail.
6. Rehearse on the bundled mock recording before any real one: `python scripts/rehearse.py expert-to-sop --base http://127.0.0.1:8445`
   (from the decosa-api checkout) must print 10/10 checks passed.

Every model call on the direct route gets a receipt signed with this box's key (status `attested`): an attestation by
me, the operator, not a proof of computation.

## 5. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`, or call `POST /sop/draft` from your own
tooling and keep the Markdown export and the signed record with each SOP revision. Contract: `API_CONTRACT.md`,
section "video input + expert-to-sop".

Off by default. Joining serves other people's requests on this GPU; never do it on a box that holds confidential
recordings. Follow the provider guide at `/provide` on the site only if I ask, and do not enable it without my explicit yes.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

3 laws, rules and guidance pages cited; 3 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
SOP from video: an expert records the screen or the bench while doing the procedure and talking, and the draft comes back as numbered steps, each with the time it happens, a keyframe from that moment and a check against what was said.
Who it's for
CI, training and quality leads who keep work instructions current, and IT and ops teams who own runbooks.
Where it runs
Hosted or self-host
Key numbers

On 4 held-out synthetic screen recordings, 21 of 25 planted steps were found, all in order, and 4 of 5 steps said but never shown were flagged; on one recording the cited times were wrong and a warning says so.

  • 21 / 25 Planted on-screen steps found (test split, n = 25)
  • 100% Found steps in the right order (test split, n = 21)
  • 19 / 21 Cited start within 3 s (test split, n = 21)
  • 20.2 s Median end-to-end run, hosted (QA sweep 2026-09-27)
All results, datasets and caveats
Models
Qwen3.8-27B (watches the recording, checks each step) · MOSS-Transcribe-Diarize (the narration)
Where
Hosted or self-host
Checks
Receipt per model call, with the video's sha256 and sampling in the request hash; signed ASR receipt; signed revision record with step, keyframe and receipt hashes
Output
Notes, reports and drafts · Signed record or verdict
Data
Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

How does it turn an SOP from video into steps?

Qwen3.8-27B watches the recording at one frame a second and lists every action it sees with its time range. A keyframe is cut at each cited time, each step is checked against the narration around it, and instructions said but not drafted are checked back against the video.

What does needs confirmation mean?

The expert said it, but the recording does not show it being done: often a safety or quality check such as wearing gloves or checking a form. The reviewer decides whether to keep it, confirm it with the expert or remove it before signing.

How accurate is it?

On 4 held-out synthetic screen recordings it found 21 of 25 planted on-screen steps, all in the right order, flagged 4 of 5 steps said but not shown, and wrongly flagged none. On one recording the model cited times in 2-second slots instead of reading them; the draft shows a timing warning when that happens.

Is the draft an approved procedure?

No. It is a draft for a qualified person to review and sign. Where rules such as OSHA 29 CFR 1910.147 require documented procedures, the employer still writes, authorizes and inspects them.

Can recordings stay on our own hardware?

Yes. The API, Qwen3.8-27B with video input (Apache-2.0) and the speech-to-text model run in containers on one 96 GB GPU, so recordings of internal systems or shop floors never leave your machine.

What does it cost to run?

About $0.016 of model time for a 40-second recording at the gateway's list price (16 calls, measured on the demo sample), and about 20 seconds.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Expert-to-SOP

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.