Pull requests / #994

#994 Isolate concurrent image positions and validate vision records

closed · @ilumn · 0 评论 · 在 GitHub 查看

NVIDIA / CUDAModels & quants

描述

With `--batch` and `--vision`, admitting another image or text prompt overwrites the device-wide MRoPE table still used by active slots. Their query/key and indexer rotations then use the new request's positions, so an unrelated admission can change an image reply.

Give each slot a stable position table on every GPU stage and capture that slot's pointer in its rotation/indexer graphs. Slot reuse updates the contents without changing captured addresses. The same image-admission path now rejects malformed SVE1 records and checks context bounds before constructing positions.

Validation against official main `6f32ec0`, using Swift IQ3_XXS on two RTX 5090s with CUDA 13.0:

- The unchanged serial/batch control had unequal argmax results and maximum logit difference 6.84. The fix matched all 95,354,880 FP32 vocabulary values in four configurations: one/two GPUs, pipeline groups, and FP16 KV at 262K allocated capacity. Comparisons used upstream's exactness controls and matched expert placement.
- Clean greedy/seeded generation matched six mixed requests; cancellation, slot reuse and invalid-request recovery passed.
- Added regression tests cover SVE1 validation and graph isolation across different image grids/text, including a wrong-table negative control.

The two added regression tests pass. The 262K run checks allocation and short-request correctness; full-length context quality and other models/backends remain outside qualification.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。