Pull requests / #659

#659 Fix WDDM host pinning for gfx1031 native inference

closed · @AxeSl · 0 comentários · No GitHub

AMD / HIPNVIDIA / CUDAModels & quantsWindows

Descrição

## Summary

- Include commit `e9630cd` with gfx1031 support and Q5_1 PLE handling.
- Cap fallback sliced `cudaHostRegister` on native Windows using the GPU's WDDM shared-memory budget when `STRATA_ARENA_PIN_GIB` is unset.
- Preserve explicit numeric caps, `auto`, and `0` as the uncapped legacy override.
- Document the WDDM pinning behavior and configuration controls.

## Root cause

When whole-arena registration failed on Windows, the fallback registered layer-sized slices until the driver refused them. On the reported RX 6700M run this produced 46 registered slices / approximately 44 GiB while WDDM reported 9426 MiB of 10224 MiB budget in use. That consumed the shared-memory budget needed by later VRAM allocations and verify work, causing paging and stalls. The remaining arena is still kept resident through the working-set lock, but only the budget-safe prefix is CUDA-registered.

## Validation

- `git diff --check` passed and the worktree is clean.
- HIP configuration with `CMAKE_HIP_ARCHITECTURES=gfx1031` reaches the gfx1031 architecture warning and HIP compiler detection.
- Full HIP build was unavailable in this checkout because `.rocm-win` is absent and the Visual Studio generator does not support HIP.
- IQ pack test was unavailable because the llama.cpp `gguf-py` dependency is not present.
- No commands accessed or modified `C:\\test12\\Strata`.

No site

Links install, modelos, releases.