Pull requests / #919
#919 hip: Windows support for Radeon gfx1151 APUs (on #895)
closed · @storm-ace · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
描述
Windows support for the Radeon 8060S / 8050S (gfx1151), on top of #895. Details, measurements and the investigation are in #918. **Depends on #895:** this branch contains #895's two commits; only the last three commits (`95e7bed`, `76890d4`, `9bd3d50`) are new. Once #895 is merged, the diff here shrinks to those. [Diff of these commits alone](https://github.com/storm-ace/Strata/compare/598e9a2...windows-gfx1151). Closes #918. ## Changes - **setup.py:** the 8060S / 8050S by PCI id `0x1586` or name on Windows. Its registry figure (the carve-out) and HIP's (carve-out + shared RAM) are marked `shared_memory`, so neither adds to system RAM in the model choice. The default context is sized from the carve-out (16 GB, Task Manager's dedicated GPU memory), not from HIP's 43.8 GiB: 64K instead of 128K on the test PC (`76890d4`). `--low-ram auto` puts an APU in the low-RAM resident mode: its GPU cache is RAM too, so the usual full copy of the experts in RAM would hold most of them twice. On the test PC that full copy (IQ2_XS, 31.6 GiB) left no RAM to read the experts with and the first start failed with `short unbuffered read ... (error 1450)`. The estimate now counts the carve-out plus the RAM the cache borrows (it printed "~0%" before) (`9bd3d50`). - **src/program/generate.cpp, src/core/device.cu, include/strata/hip_compat/cuda_runtime.h:** #895's APU cap on the automatic expert cache also on Windows, plus what is left of the firmware carve-out (DXGI dedicated memory less local usage, `hip_compat::apu_dedicated_free()`), which Windows keeps outside system RAM. The arch check also covers a runtime that does not set `integrated`. - **tools/hip/build_windows.bat:** gfx1151 in the default archs. - **tools/hip/package_windows.py:** copies rocBLAS's gfx1151 kernels, which TheRock 10.2 ships as a folder. - **tests/hip/prefill_mmq_parity.cpp:** the output sentinel is set with `hipMemsetAsync` on the product's stream. The null-stream `hipMemset` landed after the MMQ kernel on the 8060S and left only the sentinel; the kernels were right. - **tools/test_setup_amd.py, docs/AMD_HIP.md:** a `test_strix_halo` case and the Windows notes with these measurements. ## Tested Ryzen AI Max+ PRO 395, Radeon 8060S, 64 GB (16 GB carve-out), Windows 11, AMD driver 32.0.22018.5, TheRock ROCm 10.2.0a20260930: - `tools\hip\build_windows.bat tests`: 280 targets build, the zip packages with gfx1151. - `strata-device --selftest`: OK. - HIP ctest: 60 of 65 pass, 2 skipped; the 5 failures are the known ones (`hip_handoff` on Windows, `ple_parity` / `expert_parity` / `pool_test` need model fixtures, `expert_cache_segmented_test` is CUDA-only). - `python tools/test_setup_*.py`: all pass. **Model run** (Qwen3.8-Flash-Next IQ2_XS, 64K context, `--kv int8`, MTP on, `--resident-experts`, reasoning off, temperature 0; single runs with other programs open, not a benchmark): ``` strata generate: integrated AMD GPU: 9.62 GiB of its carve-out free strata generate: expert cache auto: 37.07 GiB free, 700 MiB reserved (+143 MiB for the draft head) -> 24576 slots FileExpertSource: allocating 3.33 GiB pageable resident cache complement ``` | Request | Prompt | Prompt read | Output | Drafts accepted | | --- | ---: | ---: | ---: | ---: | | short question + one-liner | 42 tokens (35 reused) | 56.6 tok/s | 24 tokens, 35.4 tok/s | 17 of 20 | | palindrome function with asserts | 43 tokens | 57.3 tok/s | 102 tokens, 32.3 tok/s | 66 of 92 | | 900 lines of `setup.py` + a question about two constants | 17,116 tokens | 291.4 tok/s | 78 tokens, 32.6 tok/s | 55 of 65 | All answers were right, including both constants from the 17K-token prompt, which runs the MMQ prompt path end to end. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。