Pull requests / #659
#659 Fix WDDM host pinning for gfx1031 native inference
closed · @AxeSl · 0 Kommentare · Auf GitHub
AMD / HIPNVIDIA / CUDAModels & quantsWindows
Beschreibung
## Summary - Include commit `e9630cd` with gfx1031 support and Q5_1 PLE handling. - Cap fallback sliced `cudaHostRegister` on native Windows using the GPU's WDDM shared-memory budget when `STRATA_ARENA_PIN_GIB` is unset. - Preserve explicit numeric caps, `auto`, and `0` as the uncapped legacy override. - Document the WDDM pinning behavior and configuration controls. ## Root cause When whole-arena registration failed on Windows, the fallback registered layer-sized slices until the driver refused them. On the reported RX 6700M run this produced 46 registered slices / approximately 44 GiB while WDDM reported 9426 MiB of 10224 MiB budget in use. That consumed the shared-memory budget needed by later VRAM allocations and verify work, causing paging and stalls. The remaining arena is still kept resident through the working-set lock, but only the budget-safe prefix is CUDA-registered. ## Validation - `git diff --check` passed and the worktree is clean. - HIP configuration with `CMAKE_HIP_ARCHITECTURES=gfx1031` reaches the gfx1031 architecture warning and HIP compiler detection. - Full HIP build was unavailable in this checkout because `.rocm-win` is absent and the Visual Studio generator does not support HIP. - IQ pack test was unavailable because the llama.cpp `gguf-py` dependency is not present. - No commands accessed or modified `C:\\test12\\Strata`.
Mehr auf der Site
Links zu Install, Modellen, Releases.