Pull requests / #919

#919 hip: Windows support for Radeon gfx1151 APUs (on #895)

closed · @storm-ace · 0 commentaires · Sur GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

Description

Windows support for the Radeon 8060S / 8050S (gfx1151), on top of #895. Details, measurements and the investigation are in #918.

**Depends on #895:** this branch contains #895's two commits; only the last three commits (`95e7bed`, `76890d4`, `9bd3d50`) are new. Once #895 is merged, the diff here shrinks to those. [Diff of these commits alone](https://github.com/storm-ace/Strata/compare/598e9a2...windows-gfx1151).

Closes #918.

## Changes

- **setup.py:** the 8060S / 8050S by PCI id `0x1586` or name on Windows. Its registry figure (the carve-out) and HIP's (carve-out + shared RAM) are marked `shared_memory`, so neither adds to system RAM in the model choice.
  The default context is sized from the carve-out (16 GB, Task Manager's dedicated GPU memory), not from HIP's 43.8 GiB: 64K instead of 128K on the test PC (`76890d4`).
  `--low-ram auto` puts an APU in the low-RAM resident mode: its GPU cache is RAM too, so the usual full copy of the experts in RAM would hold most of them twice. On the test PC that full copy (IQ2_XS, 31.6 GiB) left no RAM to read the experts with and the first start failed with `short unbuffered read ... (error 1450)`. The estimate now counts the carve-out plus the RAM the cache borrows (it printed "~0%" before) (`9bd3d50`).
- **src/program/generate.cpp, src/core/device.cu, include/strata/hip_compat/cuda_runtime.h:** #895's APU cap on the automatic expert cache also on Windows, plus what is left of the firmware carve-out (DXGI dedicated memory less local usage, `hip_compat::apu_dedicated_free()`), which Windows keeps outside system RAM. The arch check also covers a runtime that does not set `integrated`.
- **tools/hip/build_windows.bat:** gfx1151 in the default archs.
- **tools/hip/package_windows.py:** copies rocBLAS's gfx1151 kernels, which TheRock 10.2 ships as a folder.
- **tests/hip/prefill_mmq_parity.cpp:** the output sentinel is set with `hipMemsetAsync` on the product's stream. The null-stream `hipMemset` landed after the MMQ kernel on the 8060S and left only the sentinel; the kernels were right.
- **tools/test_setup_amd.py, docs/AMD_HIP.md:** a `test_strix_halo` case and the Windows notes with these measurements.

## Tested

Ryzen AI Max+ PRO 395, Radeon 8060S, 64 GB (16 GB carve-out), Windows 11, AMD driver 32.0.22018.5, TheRock ROCm 10.2.0a20260930:

- `tools\hip\build_windows.bat tests`: 280 targets build, the zip packages with gfx1151.
- `strata-device --selftest`: OK.
- HIP ctest: 60 of 65 pass, 2 skipped; the 5 failures are the known ones (`hip_handoff` on Windows, `ple_parity` / `expert_parity` / `pool_test` need model fixtures, `expert_cache_segmented_test` is CUDA-only).
- `python tools/test_setup_*.py`: all pass.

**Model run** (Qwen3.8-Flash-Next IQ2_XS, 64K context, `--kv int8`, MTP on, `--resident-experts`, reasoning off, temperature 0; single runs with other programs open, not a benchmark):

```
strata generate: integrated AMD GPU: 9.62 GiB of its carve-out free
strata generate: expert cache auto: 37.07 GiB free, 700 MiB reserved (+143 MiB for the draft head) -> 24576 slots
FileExpertSource: allocating 3.33 GiB pageable resident cache complement
```

| Request | Prompt | Prompt read | Output | Drafts accepted |
| --- | ---: | ---: | ---: | ---: |
| short question + one-liner | 42 tokens (35 reused) | 56.6 tok/s | 24 tokens, 35.4 tok/s | 17 of 20 |
| palindrome function with asserts | 43 tokens | 57.3 tok/s | 102 tokens, 32.3 tok/s | 66 of 92 |
| 900 lines of `setup.py` + a question about two constants | 17,116 tokens | 291.4 tok/s | 78 tokens, 32.6 tok/s | 55 of 65 |

All answers were right, including both constants from the 17K-token prompt, which runs the MMQ prompt path end to end.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Sur le site

Liens install, modèles, releases.