Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Video input: Qwen3.8-27B with video through the model gateway (27 Sep 2026)

What was switched on, at what cap, and what it buys. Measured on our server through the live metered route (the metered chat route, model qwen3.8-27b, gateway receipts) with scripts/eval_video_input.py. Raw results: docs/evals/video-input/results-synth.json (synthetic), results-synth-pt-mc-pt-loc.json (Perception Test; its synthetic rows are the first run, graded before the parser fix below).

The cap

Setting Value Where
Videos per request 1 vLLM --limit-mm-per-prompt {"image":4,"video":1}; the API's per-request video limit
Sampling 1 frame a second, at most 240 frames (longer clips uniformly: one frame every 2.5 s at 10 min) vLLM --media-io-kwargs {"video":{"num_frames":240,"fps":1}}
Video tokens at most 32,768 (+ about 9 timestamp tokens per 2-frame group) video processor size.longest_edge 67,108,864 px, set by a read-only override of the model's processor_config.json + video_preprocessor_config.json
Length / size ≤ 600 s, ≤ 32 MB decoded the API's video length and size limits
  • Page 48 suggested 1–2 fps and 32–64k tokens. We took the low end: 1 fps (the model's default is 2) and 32k tokens, because one video request should not starve the production KV cache. At 120 frames a 720p clip is seen at 992×544.
  • --mm-processor-kwargs '{"videos_kwargs":{"size":...}}' does not work in vLLM 0.29.0: the engine fails at startup with "Keyword argument size was passed two times" (vLLM passes a flat size too). A flat size/max_pixels would also change images. The config-file override changes the video processor only (images unchanged: a 6000×4000 image is still 16,224 tokens).
  • The model sees each 2-frame group with its time ("<12.5 seconds>"), so it can cite time ranges.

Serving cost (GPU1, RTX PRO 6000 96 GB)

Before (image only) After (video on)
KV cache 62.62 GiB, 1,766,176 tokens 60.17 GiB, 1,697,345 tokens (−3.9%)
Encoder cache budget 16,384 tokens (1 image profiled) 32,768 tokens (1 video profiled)
One maximum video n/a 32,189 prompt tokens = 1.9% of the KV pool
GPU memory used 94.4 GB (after a day up) 87.7 GB after start; 93.1 GB peak with 4 max-size videos at once; 95.8 GB steady after the eval runs (of 97.9 GB)

Latency (direct to vLLM, gateway load from other workloads at the time):

  • a 120 s 1280×720 clip at the cap (32,189 prompt tokens): 9.8 s alone, 17–19 s each with four at once;
  • a 30 s 640×360 clip (3,464 prompt tokens): 1.2 s direct, 3.6 s through the gateway;
  • synthetic eval clips through the gateway: median 9.5 s (60 s clips) to 14.7 s (240 s clips).

The 95.8 GB steady state is 1.4 GB above the image-only unit's; there were no out-of-memory errors. If it creeps up, lower --gpu-memory-utilization from 0.92 to 0.91 at the next restart.

1. Synthetic planted events (ours; CC0)

12 clips of 60, 120, 240 and 480 s (3 each), 1280×720, a moving gradient with a clock. In each, 5 black cards with one yellow word appear for 3 s at random known times. Question: list every card with the second it appears. A card counts as found when its word is read; its start is compared with the truth.

Clip length Events Found Start within 2 s Median start error Invented Median latency Prompt tokens
60 s 15 15 15 0.4 s 0 9.5 s 26,741
120 s 15 15 15 0.2 s 0 11.5 s 32,241
240 s 15 15 13 0.3 s 0 14.7 s 32,901
480 s 15 15 15 0.2 s 0 9.6 s 31,221
All 60 60 (100%) 58 (96.7%) 0.2 s 0 11.4 s
  • Run twice (temperature 0 is not deterministic under batching). Run 1: the same 60 words, 53 starts within 2 s.
  • The failure mode is a stretched clock on long clips, not missed events. On one 240 s clip, the last two cards came back 20 s and 40 s late, one of them after the end of the clip. On one 480 s clip in run 1, every time was about 1.3× too late (85 s for a card at 65 s). The words were always right. A caller should sanity-check cited times against the duration; Expert-to-SOP does (its timing warning) and shows a keyframe for every cited time.
  • Grader note: the first grader expected {"cards": [...]} and scored a bare JSON array as nothing found. It was fixed and the rows re-scored from the kept answers; no prompt changed.

2. Perception Test sample split (DeepMind, CC-BY 4.0)

The only licence-clean public video set with timestamps we could use. Page 48 names Video-MME, LongVideoBench, TimeScope and Charades-STA; to our knowledge (from memory, not re-checked today) Video-MME and Charades are for research or non-commercial use only, and LongVideoBench, MLVU, TimeScope and YouCook2 are built on YouTube videos whose terms we do not hold, so none of them is used here. The Perception Test's videos and annotations are CC-BY 4.0 (github.com/google-deepmind/perception_test, checked 27 Sep 2026). Only the sample split (8 videos, 26–35 s, 215 MB) was downloaded; the validation split is 70 GB. The videos show paid participants who consented to filming. The sample is small: read these as a sanity check, not a benchmark.

Task n Video part (at our cap) 4 still frames (the old way) Chance
Multiple-choice video QA 36 questions 28 / 36 (77.8%) 26 / 36 (72.2%) 33%
Action localisation (labels that occur once in a clip) 37 segments mean IoU 0.456 · R@IoU 0.3: 62.2% · R@IoU 0.5: 48.6% · start within 2 s: 73.0%, within 5 s: 86.5% · median start error 0.93 s n/a

MC latency: 1.2 s median with the video (15,166 prompt tokens) vs 0.9 s with four frames (960 tokens). Localisation asks "when does <label> happen" for each action label that occurs once in its clip (37 of 71 segments).

3. Receipts

Every call went through the metered route and came back with a gateway-signed receipt. For video requests the receipt's request_hash is sha256(0xff || canonical_json({"v": <the multimodal hash version>, "messages": [...]})) where the video part is {"type": "video", "mime", "bytes", "sha256", "sampling": {"fps": 1.0, "max_frames": 240, "max_pixels": 67108864}}; it was recomputed locally for text, image and video requests after each gateway restart and matched. Billing: the runtime's prompt tokens, capped at a ceiling from the sampling (30 s 640×360 clip: ceiling 3,482, billed 3,464).

What this does not show

  • Hour-long video: the cap is 10 minutes; past 4 minutes frames are sparser than one a second.
  • Audio: the model gets no sound. Speech goes through the ASR block (Expert-to-SOP transcribes the narration separately).
  • Real footage at scale: the public part is 8 short clips; the planted-event clips are ours and easy.
  • Reproduce: python scripts/eval_video_input.py --data <internal path> --out docs/evals/video-input (needs the Perception Test sample split in perception-test/; the synthetic clips are regenerated from a fixed seed).