Pull requests / #1609
#1609 Video input: `video_url` parts through strata-vision, in the Qwen3-VL layout
open · @Destrayon · 0 comentarios · En GitHub
Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux
Descripción
## Title Issue: none filed; DETAILS.md lists "no video" under the current limits. ## Summary Qwen3.8-Flash-Next was trained on video (its template has `<|video_pad|>`, its mmproj is `qwen3vl_merger` with temporal merge 2), but Strata reads pictures only. This adds video input through the existing image path - the engine is unchanged. A video is laid out as transformers' Qwen3-VL processor (`processing_qwen3_vl.py`) does it: frames sampled by ffmpeg, frames 0+1, 2+3, ... merged into one temporal patch (an odd last frame doubled), each pair preceded by its time as text, `<1.2 seconds>` (the mean of the two frames' real times), then `<|vision_start|>` image `<|vision_end|>`. Each frame pair is an ordinary SVE1 image record with the usual M-RoPE positions, so `GENI` reads a video as that many pictures. llama.cpp's own video helper (`mtmd_helper_video`) is not used for the layout: it writes a `Video:` label and a `[0m5.00s]` timestamp every 5 s *after* the frame it marks, which leaves frame 0 unpaired and shifts every label; with it, answers placed events about 1 s late. Its ffmpeg reader is used for the frames; its `fps` filter returns frame k at about (k+0.5)/fps (measured: source frames 7, 22, 37, ... of a 30 fps clip at 2 fps), and the labels say so. ## What changed **strata-vision: `ENCV <fps> <max_frames> <max_side> <pair_tokens> <total_tokens> <min_pair_tokens> <video> <out>`** answers `OK <tokens> <pairs> <frames> <ms>` and `LAYOUT <T ids|I cells; ...>`. It plans from the video's probed length within the whole-video budget: 2 fps at up to 768 tokens per pair (qwen-vl-utils' default) while it fits, then smaller frames down to 128 per pair, then a lower frame rate spread over the whole video - so any length is one request. It encodes 64 frames at a time (peak 859 MB for a 10-minute video at 158K tokens). `--ffmpeg-dir` names ffmpeg/ffprobe when they are not on PATH. `ENC` is unchanged. **serve:** - `video_url` (OpenAI), `input_video` (Responses) and `video` blocks (Anthropic: `base64`, `url` or `path` source); `data:` URLs, `http(s)` URLs and local paths, as for images. - The template's `<|vision_start|><|video_pad|><|vision_end|>` is replaced by the layout (text ids + one `<|image_pad|>` per cell); the pairs' records join the request's embeddings file in prompt order with the images. A literal `<|video_pad|>` in a message stays text, as #150 does for images. - Budget: `video_total_tokens` 0 = 60% of the context (`video_context_share`), at most 224K (the model card's setting for hour-scale video), shared by the videos of one request. `video_fps`, `video_tokens`, `video_min_tokens`, `video_max_frames`, `video_max_side` in the `"vision"` section; a part's own `fps`, `max_frames`, `max_side`, `tokens`, `total_tokens` override them. The section's `max_tokens` caps a frame pair as it caps a picture. - The image rules apply: a URL is downloaded outside the FIFO (`VIDEO_URL_MAX` 1 GiB), local files only from the server's own page or trusted origins (`_no_local_images`), names relative to the encoder's directory (#480), the combined file written after the context check, a request's own files never evicted from the encoder's cache (#1072), `strata_prefix` ignored as with images. A video's encoding gets its own read timeout, `STRATA_VISION_VIDEO_S` (30 min; the image's 300 s is too short for a long video on the CPU). - `/health` has `"videos"`, `/v1/models` lists `"video"` when the encoder can read them. **web:** the attach button and dropping a file take a video (up to 200 MB) when `/health` says videos are read; it plays in the message and goes out as a `video_url` part. **docs:** a Videos section in DETAILS.md (the layout, the keys, measurements) and a line in the README. ## Extra Notes - **ffmpeg:** videos need `ffmpeg` and `ffprobe` installed, as llama.cpp's video support does (`MTMD_VIDEO`: "requires ffmpeg binary in PATH"); `"ffmpeg_dir"` names their folder otherwise. Setup does not install them. Without them a video is refused with `ERR cannot read the video ... (is ffmpeg/ffprobe installed?)`; pictures are unaffected. - **For a release:** the ready-made `strata-vision` needs rebuilding for `ENCV`. - **#1327** adds video another way: an engine-side media format, off by default, clips over 60 s or 128 frames refused. This one keeps the engine and the image path as they are and reads any length in one request by planning within the context. Either can be taken; they touch the same server code, so not both. ### Measured Encoder on an RTX 5070 Ti (built with CUDA 12.8 for sm_120) or an i9-13900K (16 threads); the model IQ2_XS on the 5070 Ti, or IQ3_S on an RTX 2080 Ti with 94 GB of RAM; thinking off; at most 448 tokens per pair and labels without the half-interval correction (both changed since, see Tested). | Video | Budget | Frames | Tokens | Encode | | --- | ---: | ---: | ---: | ---: | | 10 s, 720x358 (`tools/mtmd/test-3.mp4`) | - | 20 | 1,078 | CPU 6.5 s, GPU 0.74 s | | 36 s, 1080p | 39K (64K context) | 72 (2 fps) | 16,128 | GPU 8.6 s | | 10 min, 540p | 39K (64K context) | 614 (~1 fps) | 36,840 | GPU 21 s | | 10 min, 540p | 157K (256K context) | 1,200 (2 fps) | 158,400 | GPU 71 s | - 10 s clip: the scene described in order; "a revolver ... at the 6-second mark" (in view from ~6 s). - 10-minute synthetic video (20 scenes of 30 s, a 2-second banner at 7:13) on the 2080 Ti: the scene at 7:20 and the banner at 7:13 right, at 64K (38,589 prompt tokens, 68 s; follow-up 2.9 s from the conversation cache) and at 256K (157,124 tokens, 328 s; follow-up 5.5 s), where the banner's 2-second length was also right. - Small text (12 s 1080p, busy background, 20 px HUD, 18 px sign, 28 px subtitles): ~100 tokens per pair read 0/3 subtitles; 448 per pair 3/3 at their exact seconds; `"fps": 1, "tokens": 1024` also every HUD value. - Two videos in one request (a gameplay recording vs a copy with the minimap blacked out, 30% saturation, 1.25x speed): all three differences found, plus one that was not there; 42,606 tokens, 153 s. ### Tested Rebased onto main fb58e0d. Since the measurements, two changes to match Qwen's reference: each label is the pair's real time (it was 0.5/fps early: 0.25 s at 2 fps), and the per-pair default is 768 tokens, qwen-vl-utils' `VIDEO_MAX_TOKEN_NUM` (was 448; long videos are sized by the budget either way). - `python -m unittest serve.test_server`: 280 pass, 11 new (`VideoParts`: layout splicing, images and videos in prompt order, literal marker, request shapes for the three APIs, budget, the ENCV line, temp file removed on a broken pipe). - This branch's strata-vision built with MSVC (CPU): `ENC` unchanged (300 tokens for `test-1.jpeg`), `ENCV` on the 10 s clip (20 frames, 2,530 tokens; labels 0.5, 1.5, ... after the fix). The same ENCV code was built with CUDA 12.8 (sm_120) for the measurements above. - Before the rebase (on main 82f46a8), the server end to end with IQ2_XS on the 5070 Ti (engine 0.1.39), CPU encoder: the 10 s clip's revolver at 6 s right; on a 36 s HUD video the first HP change (87 -> 41 at 6 s) right, the later four not named. - Not tested: HIP/SYCL, Linux, the CUDA 13 release build of the encoder. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
En el sitio
Enlaces a install, modelos, releases.