Pull requests / #1609

#1609 Video input: `video_url` parts through strata-vision, in the Qwen3-VL layout

open · @Destrayon · 0 comments · View on GitHub

Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux

Description

## Title
Issue: none filed; DETAILS.md lists "no video" under the current limits.

## Summary
Qwen3.8-Flash-Next was trained on video (its template has `<|video_pad|>`, its mmproj is `qwen3vl_merger` with
temporal merge 2), but Strata reads pictures only. This adds video input through the existing image path - the engine
is unchanged.

A video is laid out as transformers' Qwen3-VL processor (`processing_qwen3_vl.py`) does it: frames sampled by ffmpeg,
frames 0+1, 2+3, ... merged into one temporal patch (an odd last frame doubled), each pair preceded by its time as
text, `<1.2 seconds>` (the mean of the two frames' real times), then `<|vision_start|>` image `<|vision_end|>`. Each
frame pair is an ordinary SVE1 image record with the usual M-RoPE positions, so `GENI` reads a video as that many
pictures. llama.cpp's own video helper (`mtmd_helper_video`) is not used for the layout: it writes a `Video:` label and
a `[0m5.00s]` timestamp every 5 s *after* the frame it marks, which leaves frame 0 unpaired and shifts every label;
with it, answers placed events about 1 s late. Its ffmpeg reader is used for the frames; its `fps` filter returns
frame k at about (k+0.5)/fps (measured: source frames 7, 22, 37, ... of a 30 fps clip at 2 fps), and the labels say so.

## What changed
**strata-vision: `ENCV <fps> <max_frames> <max_side> <pair_tokens> <total_tokens> <min_pair_tokens> <video> <out>`**
answers `OK <tokens> <pairs> <frames> <ms>` and `LAYOUT <T ids|I cells; ...>`. It plans from the video's probed
length within the whole-video budget: 2 fps at up to 768 tokens per pair (qwen-vl-utils' default) while it fits, then
smaller frames down to 128 per pair, then a lower frame rate spread over the whole video - so any length is one
request. It encodes 64 frames at a time (peak 859 MB for a 10-minute video at 158K tokens). `--ffmpeg-dir` names
ffmpeg/ffprobe when they are not on PATH. `ENC` is unchanged.

**serve:**
- `video_url` (OpenAI), `input_video` (Responses) and `video` blocks (Anthropic: `base64`, `url` or `path` source);
  `data:` URLs, `http(s)` URLs and local paths, as for images.
- The template's `<|vision_start|><|video_pad|><|vision_end|>` is replaced by the layout (text ids + one
  `<|image_pad|>` per cell); the pairs' records join the request's embeddings file in prompt order with the images.
  A literal `<|video_pad|>` in a message stays text, as #150 does for images.
- Budget: `video_total_tokens` 0 = 60% of the context (`video_context_share`), at most 224K (the model card's
  setting for hour-scale video), shared by the videos of one request. `video_fps`, `video_tokens`,
  `video_min_tokens`, `video_max_frames`, `video_max_side` in the `"vision"` section; a part's own `fps`,
  `max_frames`, `max_side`, `tokens`, `total_tokens` override them. The section's `max_tokens` caps a frame pair as it
  caps a picture.
- The image rules apply: a URL is downloaded outside the FIFO (`VIDEO_URL_MAX` 1 GiB), local files only from the
  server's own page or trusted origins (`_no_local_images`), names relative to the encoder's directory (#480), the
  combined file written after the context check, a request's own files never evicted from the encoder's cache
  (#1072), `strata_prefix` ignored as with images. A video's encoding gets its own read timeout,
  `STRATA_VISION_VIDEO_S` (30 min; the image's 300 s is too short for a long video on the CPU).
- `/health` has `"videos"`, `/v1/models` lists `"video"` when the encoder can read them.

**web:** the attach button and dropping a file take a video (up to 200 MB) when `/health` says videos are read; it
plays in the message and goes out as a `video_url` part.

**docs:** a Videos section in DETAILS.md (the layout, the keys, measurements) and a line in the README.

## Extra Notes
- **ffmpeg:** videos need `ffmpeg` and `ffprobe` installed, as llama.cpp's video support does (`MTMD_VIDEO`: "requires
  ffmpeg binary in PATH"); `"ffmpeg_dir"` names their folder otherwise. Setup does not install them. Without them a
  video is refused with `ERR cannot read the video ... (is ffmpeg/ffprobe installed?)`; pictures are unaffected.
- **For a release:** the ready-made `strata-vision` needs rebuilding for `ENCV`.
- **#1327** adds video another way: an engine-side media format, off by default, clips over 60 s or 128 frames
  refused. This one keeps the engine and the image path as they are and reads any length in one request by planning
  within the context. Either can be taken; they touch the same server code, so not both.

### Measured
Encoder on an RTX 5070 Ti (built with CUDA 12.8 for sm_120) or an i9-13900K (16 threads); the model IQ2_XS on the
5070 Ti, or IQ3_S on an RTX 2080 Ti with 94 GB of RAM; thinking off; at most 448 tokens per pair and labels
without the half-interval correction (both changed since, see Tested).

| Video | Budget | Frames | Tokens | Encode |
| --- | ---: | ---: | ---: | ---: |
| 10 s, 720x358 (`tools/mtmd/test-3.mp4`) | - | 20 | 1,078 | CPU 6.5 s, GPU 0.74 s |
| 36 s, 1080p | 39K (64K context) | 72 (2 fps) | 16,128 | GPU 8.6 s |
| 10 min, 540p | 39K (64K context) | 614 (~1 fps) | 36,840 | GPU 21 s |
| 10 min, 540p | 157K (256K context) | 1,200 (2 fps) | 158,400 | GPU 71 s |

- 10 s clip: the scene described in order; "a revolver ... at the 6-second mark" (in view from ~6 s).
- 10-minute synthetic video (20 scenes of 30 s, a 2-second banner at 7:13) on the 2080 Ti: the scene at 7:20 and
  the banner at 7:13 right, at 64K (38,589 prompt tokens, 68 s; follow-up 2.9 s from the conversation cache) and at
  256K (157,124 tokens, 328 s; follow-up 5.5 s), where the banner's 2-second length was also right.
- Small text (12 s 1080p, busy background, 20 px HUD, 18 px sign, 28 px subtitles): ~100 tokens per pair read 0/3
  subtitles; 448 per pair 3/3 at their exact seconds; `"fps": 1, "tokens": 1024` also every HUD value.
- Two videos in one request (a gameplay recording vs a copy with the minimap blacked out, 30% saturation, 1.25x
  speed): all three differences found, plus one that was not there; 42,606 tokens, 153 s.

### Tested
Rebased onto main fb58e0d. Since the measurements, two changes to match Qwen's reference: each label is the pair's
real time (it was 0.5/fps early: 0.25 s at 2 fps), and the per-pair default is 768 tokens, qwen-vl-utils'
`VIDEO_MAX_TOKEN_NUM` (was 448; long videos are sized by the budget either way).

- `python -m unittest serve.test_server`: 280 pass, 11 new (`VideoParts`: layout splicing, images and videos in
  prompt order, literal marker, request shapes for the three APIs, budget, the ENCV line, temp file removed on a
  broken pipe).
- This branch's strata-vision built with MSVC (CPU): `ENC` unchanged (300 tokens for `test-1.jpeg`), `ENCV` on the
  10 s clip (20 frames, 2,530 tokens; labels 0.5, 1.5, ... after the fix). The same ENCV code was built with CUDA
  12.8 (sm_120) for the measurements above.
- Before the rebase (on main 82f46a8), the server end to end with IQ2_XS on the 5070 Ti (engine 0.1.39), CPU
  encoder: the 10 s clip's revolver at 6 s right; on a 36 s HUD video the first HP change (87 -> 41 at 6 s) right,
  the later four not named.
- Not tested: HIP/SYCL, Linux, the CUDA 13 release build of the encoder.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.