Pull requests / #1327

#1327 serve: opt-in finite video input

open · @constantindjonkam · 0 コメント · GitHub で見る

Server & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows

本文

## Summary

Adds opt-in finite video input on the existing Qwen4 vision path. Video remains disabled by default; existing text and image-only requests keep their current path. The encoder and engine must both advertise the same `qwen4_exp_16x2x2_2560_v1` video profile.

Supported user-message forms are OpenAI `video_url`, Anthropic/Responses `input_video`, and Qwen `video`. Per-request sampling and resolution overrides are rejected. `count_tokens` does not guess the expanded length of video content.

## Timestamp-based sampling

Serving now samples from presentation timestamps (PTS), rather than relying on the container's reported frame rate or a frame-index linspace. It places points on a fixed grid starting at the first frame, chooses the nearest actual PTS (earlier frame on an exact tie), and drops duplicate frame selections. Each chosen frame retains its actual timestamp for temporal grouping and labels. No constant-frame-rate transcode is needed. `sample_indices()` remains for the pinned index-based reference fixtures; it is not the serving path.

On the 12-frame VFR fixture, the sampler selected source indices `0, 2, 4, 6, 8, 9, 11` at `0, 0.5, 1.0, 1.5, 2.0, 3.0, 3.5` seconds. The selected RGB pixels matched the independently extracted reference images. Sampling rejects a clip that exceeds the frame budget rather than silently reducing the requested cadence.

## Runtime and safety

`ffmpeg` and `ffprobe` are executable programs, not linked libraries; their names on `PATH` are sufficient. They can be configured explicitly when needed. Pillow is required for resize. If the video tools are unavailable, image requests remain available and video is reported unavailable.

The defaults remain bounded: 2 FPS, 60 seconds, 128 selected frames, 8,192 source frames, 64 MiB source bytes, 256 MiB decoded RGB, 16,384 visual tokens, 256 MiB embeddings/cache, 512 MiB disk, and a 120-second deadline. Over-budget inputs fail closed; the server does not thin sampling to make them fit. Video is accepted only in user messages, and the existing image-only `ENC`/`SVE1`/`GENI` path is unchanged.

## Validation and scope

- Serving suite on this commit: **452 tests passed, 10 skipped**, no failures (`MEDIA_TEST_EXE` enabled).
- C++ media tests: **5,011 checks passed**, including ASan/UBSan runs.
- The real serving preparation path preserved/bound all **240 rows** for the four-group QA fixtures. Selected RGB, PTS/frame order, Python/C++ positions and row bindings were checked against independent host references.
- Native CPU vision rows were compared against the pinned checkpoint BF16 vision implementation (BF16 weights converted to FP32 for the comparison):

  | Fixture | Rows | Mean row cosine | Minimum row cosine | RMSE | Relative RMSE |
  |---|---:|---:|---:|---:|---:|
  | CFR | 240 | 0.999905609 | 0.993993108 | 0.000439710 | 1.22988% |
  | VFR A | 240 | 0.999924659 | 0.996627284 | 0.000321461 | 0.902571% |
  | VFR B | 240 | 0.999926929 | 0.996627284 | 0.000351521 | 0.985412% |

  These are fixture measurements, not a universal tolerance or proof of answer quality. Encoding the full clip or separate frame pairs produced bitwise-identical reference embeddings on these fixtures.
- Four vision-cache arms retained the prior text cache with zero evictions:

  | Hardware / vision placement | Result |
  |---|---|
  | RTX 3090, vision in RAM | Text cache kept: 53/60 after the clip, 39/60 before. Video follow-up restored 60/89. |
  | RTX 3090, vision on GPU | Same cache result; encoder resident at 1,230 MiB. |
  | RTX 3090 + RTX 5080, split 19 / pipeline 2, vision in RAM | Same cache result. |
  | RTX 3090 + RTX 5080, vision on GPU | Same cache result; 1,230 MiB on the 3090. |

  These runs used an **8,192-token context**, not the 262,144-token production context. The 128×64, five-frame, 2-FPS clip was a transport/cache smoke test, not a caption-quality result.

These checks cover host sampling, frame extraction, transport, encoder rows, positions and bounded cache behavior. They do **not** establish general temporal grounding or video-answer accuracy; focused temporal QA fixtures have not consistently returned the requested word.

Windows, HIP, SYCL, batched video requests, and helper-drafter operation are not advertised or validated by this change.

関連リンク

インストール・モデル・リリースへの站内リンク。