Video understanding
One recording per request to Qwen3.8-27B: sampled at one frame a second, with a timestamp the model can cite, and a receipt whose request hash covers the video's sha256 and the sampling settings.
Off on the hosted API at launch: a hosted request that carries a video gets HTTP 400 with a plain message, never a silent drop. It runs on your own hardware with self-host (on request). The numbers below were measured on Decosa's own server on 27 Sep 2026, when it was on. Ask for self-host
Measured 2026-09-27. Full eval. First use: Expert-to-SOP.
Watch a real run
Loading the tool…
The cap
| Videos per request | 1 |
| Sampling | 1 frame a second, at most 240 frames (longer clips uniformly: one frame every 2.5 s at 10 minutes) |
| Video tokens | at most 32,768 (about 9 more per 2-frame group for the timestamp) |
| Length and size | up to 10 minutes and 32 MB (MP4 or QuickTime; decosa-api converts MOV/WebM and larger uploads) |
| Audio | not given to the model; speech goes through the speech-to-text block |
One video must not starve the serving model's cache, so the sampling is the low end of what the model supports (it defaults to two frames a second). The model sees each pair of frames with its time, which is what lets answers cite time ranges.
What it costs to serve
| KV cache on the 96 GB card | 62.62 GiB → 60.17 GiB (1,766,176 → 1,697,345 tokens, −3.9%) |
| One maximum video | 32,189 prompt tokens, 1.9% of the KV cache |
| Latency, 120 s 720p clip at the cap | 9.8 s alone; 17-19 s each with four at once |
| Latency, 30 s clip through the API | 3.6 s (3,464 prompt tokens) |
Results at this cap
| Synthetic planted events: words read (60 in 12 clips, 60-480 s) | 60 / 60 |
| Synthetic planted events: start within 2 s | 58 / 60 (53 / 60 on a first run) |
| Perception Test sample (CC-BY 4.0): multiple-choice QA, 36 questions | 28 / 36 with the video vs 26 / 36 with 4 still frames |
| Perception Test sample: action localisation, 37 segments | mean IoU 0.456; start within 2 s: 27 / 37 |
The failure mode is the clock, not the content: on long clips the model sometimes stretches or compresses time (every card on one 480 s clip cited about 1.3x too late; on one held-out SOP recording steps came back in 2-second slots). Words and order were right. Callers should check cited times against the duration and show the frame at each cited time; Expert-to-SOP does both.
API
| POST /v1/chat/completions (self-host) | A user message with {"type": "video_url", "video_url": {"url": "data:video/mp4;base64,..."}} to `qwen3.8-27b` served by your own vLLM. The response's receipt lists the video's sha256, size, duration, frames, sampled frames and billed tokens. The hosted API refuses video at launch (HTTP 400). |
| decosa-api: decosa_api.video | prepare() makes the clip (same timeline, no audio, ≤1280 px, fits 32 MB); keyframe(); audio_pcm(); video_part(); ask(); seconds() and clamp_range() for time ranges. LLMClient.chat checks parts before sending. |
| Receipt request hash | sha256(0xff || canonical JSON of the messages, with a version string), where the video part is {type, mime, bytes, sha256, sampling: {fps, max_frames, max_pixels}}; the receipt checker recomputes it from the request you sent. |
Where it goes next
Expert-to-SOPbuilt
Steps from a narrated recording with time ranges and keyframes, checked against the narration.
Video evidence summaryunblocked
Timestamped events from dashcam, bodycam or site footage, each cited to a time range.
Walkthrough to quoteunblocked
A contractor's walkthrough video to a line-item scope, each item cited to the moment it is shown.
Kids and disclosure pre-flightsupgrade
Whole-video checks instead of four or six sampled frames.
Licences
| Qwen3.8-27B (nvidia/Qwen3.8-27B-NVFP4) | Apache-2.0 |
| vLLM 0.29.0 | Apache-2.0 |
| ffmpeg (run as a separate program) | LGPL-2.1+ / GPL-2.0+ |
| Perception Test sample split (eval only) | CC-BY 4.0 |
What it does not do
- Hour-long video: the cap is 10 minutes, and past 4 minutes frames are sparser than one a second.
- Sound: the model does not hear the recording.
- A large public benchmark: Video-MME, Charades-STA and similar sets are research-only or built on YouTube videos, so the public part is the 8-clip Perception Test sample.