Pull requests / #247
#247 HIP: Windows and gfx1201 (RX 9070)
closed · @jagsan-cyber · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
描述
## Summary - Opt-in HIP builds accept wave32 `gfx1100`, `gfx1200`, and `gfx1201` on Linux or Windows. Windows setup uses an installed ROCm tree that contains the card's bitcode, preferring the ROCm 10 `rocm-sdk` over the HIP SDK when `ROCM_PATH` is unset. - The host compiler is ROCm clang. CMake rejects mixing MSVC with Clang HIP. The CUDA compatibility header is force-included with `/FI` for MSVC and `-include` otherwise. - A 3-shard Unsloth GGUF loads through the native pack: gate, up, and down may sit in different shards, Q8_0 expert and embedding rows have a GPU dot, and `output.weight` is accepted when `general.architecture` is only on the first shard. - On gfx1201, hipBLAS can return success and still leave `hipErrorInvalidValue` set. That sticky error is cleared after a successful GEMM. An F32 `ple_conv1d` is converted to FP16 at startup, because the convolution kernels read FP16. ## Test Windows 11, AMD Radeon RX 9070 (`gfx1201`, 16 GB), ROCm 10.0.0, Unsloth Qwen3.8-Flash-Next UD-Q3_K_XL. The engine probes the GPU and serves the model. Greedy short prompts return coherent answers (`2`, `7`, `東京`). A 4445-token prompt prefilled at about 238 tok/s and then decoded at about 23 tok/s with MTP. Context 131072 fits; the expert cache is then 2357 slots / 4.97 GiB. These rates are for this machine, not the published RTX 5070 or RX 7900 XTX numbers. A prebuilt `strata.exe` for this commit is on the fork release below. It still needs the ROCm 10 runtime on `PATH`. It is not part of this diff. https://github.com/jagsan-cyber/Strata/releases/tag/win-gfx1201-2026-09-30
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。