Pull requests / #1498
#1498 KV streaming under WSL: pin the K/V host copy with cudaHostRegister
open · @yjamil8 · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
描述
Under WSL2 the NVIDIA driver caps what `cudaHostAlloc(Mapped)` can pin at about 1 GiB in all. KV streaming allocates each QSA layer's K/V host copy that way (`src/core/layer.cpp`), so it cannot work under WSL, and setup turns it off there: the whole KV cache has to sit in VRAM beside the expert cache, which limits long contexts on WSL. `cudaHostRegister` of ordinary memory is not capped the same way. Strata already relies on that: the expert arena is registered `Portable|Mapped` in `pinned.cu`, and on my WSL machine the start log shows `expert arena: cudaHostRegister PORTABLE ok` for 46.84 GiB. This change adds `kv_host_alloc()` for the K/V host copy: - when `/dev/dxg` exists (WSL2, the same check as `under_wddm()` in `generate.cpp`) it maps anonymous memory and registers it `Portable|Mapped` before any page is touched. This also leaves WSL's small `cudaHostAlloc` budget to the staging buffers that are still allocated that way; - elsewhere `cudaHostAlloc` stays first; - either way the other method is the fallback when the first is refused, and `STRATA_KV_HOST_REGISTER=1` / `=0` forces the order; - the failure message no longer says WSL pins only 1 GiB. The buffer lives as long as the process, like the `cudaHostAlloc` copy it replaces. By inspection, native Linux and Windows behave as before unless `cudaHostAlloc` is refused; I have not run it there. **Tested** with this commit on one RTX 5090 (32 GB), Ryzen 9 9950X3D, WSL2 with 157 GiB, Windows driver 591.86, CUDA 13.1, IQ3_S, `--max-context 1048576 --rope-scaling yarn --rope-scale 4 --kv int8 --kv-resident 32768`: - start: `the host copy is pinned `KV streaming: 32768 of 1048576 cells per QSA layer in VRAM, the K/V in 12.38 GiB of pinned RAM`, with no environment variable set - `tools/needle_bench.py`: 4/4 found, 128K and 1M (1,015,070 prompt tokens) at depths 10 and 90 - decode 99-124 tok/s with 0.5-1M tokens in context (4 runs, greedy); medians of 150-191 tok/s from 55-token prompts up to 128K (single runs 113-196); about 4,200 tok/s prefill for aeam_parity` still passes; it allocates its own host pools, so it checks the streaming kernels, which this change does not touch **Not tested:** a layer split or remote experts under WSL (where t3253 applies to the expert arena, not to this copy), Docker Desktop, native Linux, Windows. **Not in this PR:** `setup.py` still never writes `--kv-resident` on WSL (`kv_streaming_wam configs there (`upgrade_config`), so until that changes WSL users have to add `--kv-resident` by hand, and `setup.py --update` removes it again. The docs also still say WSL cannot stream. I can add this PR, gated on `MIN_ENGINE`, or send it as a follow-up, whichever you prefer.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。