Skip to content
decosa

Developers · Building blocks

Video understanding

One recording per request to Qwen3.8-27B: sampled at one frame a second, with a timestamp the model can cite, and a receipt whose request hash covers the video's sha256 and the sampling settings.

Off on the hosted API at launch: a hosted request that carries a video gets HTTP 400 with a plain message, never a silent drop. It runs on your own hardware with self-host (on request). The numbers below were measured on Decosa's own server on 27 Sep 2026, when it was on. Ask for self-host

Measured 2026-09-27. Full eval. First use: Expert-to-SOP.

Watch a real run

Loading the tool…

The cap

Videos per request1
Sampling1 frame a second, at most 240 frames (longer clips uniformly: one frame every 2.5 s at 10 minutes)
Video tokensat most 32,768 (about 9 more per 2-frame group for the timestamp)
Length and sizeup to 10 minutes and 32 MB (MP4 or QuickTime; decosa-api converts MOV/WebM and larger uploads)
Audionot given to the model; speech goes through the speech-to-text block

One video must not starve the serving model's cache, so the sampling is the low end of what the model supports (it defaults to two frames a second). The model sees each pair of frames with its time, which is what lets answers cite time ranges.

What it costs to serve

KV cache on the 96 GB card62.62 GiB → 60.17 GiB (1,766,176 → 1,697,345 tokens, −3.9%)
One maximum video32,189 prompt tokens, 1.9% of the KV cache
Latency, 120 s 720p clip at the cap9.8 s alone; 17-19 s each with four at once
Latency, 30 s clip through the API3.6 s (3,464 prompt tokens)

Results at this cap

Synthetic planted events: words read (60 in 12 clips, 60-480 s)60 / 60
Synthetic planted events: start within 2 s58 / 60 (53 / 60 on a first run)
Perception Test sample (CC-BY 4.0): multiple-choice QA, 36 questions28 / 36 with the video vs 26 / 36 with 4 still frames
Perception Test sample: action localisation, 37 segmentsmean IoU 0.456; start within 2 s: 27 / 37

The failure mode is the clock, not the content: on long clips the model sometimes stretches or compresses time (every card on one 480 s clip cited about 1.3x too late; on one held-out SOP recording steps came back in 2-second slots). Words and order were right. Callers should check cited times against the duration and show the frame at each cited time; Expert-to-SOP does both.

API

POST /v1/chat/completions (self-host)A user message with {"type": "video_url", "video_url": {"url": "data:video/mp4;base64,..."}} to `qwen3.8-27b` served by your own vLLM. The response's receipt lists the video's sha256, size, duration, frames, sampled frames and billed tokens. The hosted API refuses video at launch (HTTP 400).
decosa-api: decosa_api.videoprepare() makes the clip (same timeline, no audio, ≤1280 px, fits 32 MB); keyframe(); audio_pcm(); video_part(); ask(); seconds() and clamp_range() for time ranges. LLMClient.chat checks parts before sending.
Receipt request hashsha256(0xff || canonical JSON of the messages, with a version string), where the video part is {type, mime, bytes, sha256, sampling: {fps, max_frames, max_pixels}}; the receipt checker recomputes it from the request you sent.

Where it goes next

  • Expert-to-SOPbuilt

    Steps from a narrated recording with time ranges and keyframes, checked against the narration.

  • Video evidence summaryunblocked

    Timestamped events from dashcam, bodycam or site footage, each cited to a time range.

  • Walkthrough to quoteunblocked

    A contractor's walkthrough video to a line-item scope, each item cited to the moment it is shown.

  • Kids and disclosure pre-flightsupgrade

    Whole-video checks instead of four or six sampled frames.

Licences

Qwen3.8-27B (nvidia/Qwen3.8-27B-NVFP4)Apache-2.0
vLLM 0.29.0Apache-2.0
ffmpeg (run as a separate program)LGPL-2.1+ / GPL-2.0+
Perception Test sample split (eval only)CC-BY 4.0

What it does not do

  • Hour-long video: the cap is 10 minutes, and past 4 minutes frames are sparser than one a second.
  • Sound: the model does not hear the recording.
  • A large public benchmark: Video-MME, Charades-STA and similar sets are research-only or built on YouTube videos, so the public part is the 8-clip Perception Test sample.
All building blocks