Issues / #918

#918 Windows: Radeon 8060S (gfx1151) builds, selftests and passes the HIP ctest on top of #895

open · @storm-ace · 1 comments · View on GitHub

Setup & installAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Description

On top of #895 (Linux gfx1151) I made the Windows side work on a Radeon 8060S, for the gfx1151 port mentioned in #612. The branch is ready if it helps; I can open it as a PR on #895 or on main once #895 is in, whichever you prefer.

**Branch:** https://github.com/storm-ace/Strata/tree/windows-gfx1151 (one commit on #895's head `598e9a2`: [diff](https://github.com/storm-ace/Strata/compare/598e9a2...windows-gfx1151))

## Test machine

Laptop, Ryzen AI Max+ PRO 395, Radeon 8060S (gfx1151, PCI `1002:1586`), 64 GB with a 16 GB firmware carve-out, Windows 11 Pro 26200, AMD driver 32.0.22018.5, TheRock ROCm 10.2.0a20260930 (the version `build_windows.bat` pins).

On Windows this APU's memory looks different from Linux: Windows reports 47.8 GB of RAM (the carve-out is outside it), the display registry 16 GB of VRAM, and HIP 43.8 GiB (carve-out plus shared RAM).

## What changes

- **setup:** the 8060S / 8050S by PCI id `0x1586` or by name. Both the registry figure and HIP's are marked `shared_memory`, so neither adds to system RAM in the model choice (the same rule #895 uses on Linux).
- **engine:** #895's APU cap on the automatic expert cache (`generate.cpp`, host available memory minus 4 GiB) now also runs on Windows, where `host_available_memory()` already has a `GlobalMemoryStatusEx` branch. On Windows it adds what is left of the carve-out (DXGI `DedicatedVideoMemory` less the process's local usage, new `hip_compat::apu_dedicated_free()` next to the WDDM budget code), since Windows keeps the carve-out outside system RAM. The arch check (`gfx1151`) also covers a runtime that does not set `integrated`.
- **build_windows.bat:** gfx1151 in the default `STRATA_HIP_ARCHS`. TheRock has a `rocm-sdk-device-gfx1151` wheel for Windows.
- **package_windows.py:** rocBLAS ships gfx1151's kernels as a folder (`rocblas/library/gfx1151/`, 150 files) instead of loose files; the packager now copies a folder with `copytree` (it failed with `PermissionError` on `copy2`).
- **tests/hip/prefill_mmq_parity.cpp:** see below.

## Results on that machine

- `tools\hip\build_windows.bat tests`: builds all 280 targets and packages the zip with gfx1151.
- `strata-device --selftest`: OK (gfx1151, wave32, 20 CUs, 43.8 GiB).
- HIP ctest: **60 of 65 pass**, 2 skipped. The 5 failures are the known ones: `hip_handoff` (Windows), `ple_parity` / `expert_parity` / `pool_test` (model fixtures), `expert_cache_segmented_test` (`--vram-elastic` is CUDA-only).
- `python tools/test_setup_*.py`: all pass, with a new `test_strix_halo`.
- **Not yet:** a model run on Windows.

## `hip_prefill_mmq_parity`: a test race, not a kernel bug

It first failed with `non-finite or unwritten MMQ output`. Every output value still held the `0xff` sentinel bits; HIP returned success for the launch. What I checked:

- The gfx1151 code object of `mul_mat_q<Q2_0, 16, false>` has the same instruction mix as gfx1100's (16 `v_wmma`, 8 global stores, 109 `ds_*`, no `s_trap`); only `s_delay_alu` hints differ (21 vs 24).
- The sentinel is written with `hipMemset` on the null stream, the kernel runs on a `hipStreamNonBlocking` stream, and those do not order against each other. With a `hipDeviceSynchronize()` after the memset all six products pass at 0.05-0.13% relative L2.
- The fix sets the sentinel with `hipMemsetAsync(..., stream)`. The test then passes without the extra sync. It does not depend on gfx1151: on other cards the race was just won.

The HIP runtime also logs `KMD failed to setup the trap handler` on this APU under Windows, so a device-side trap would end a kernel without an error there. That was not the cause here, but it is worth knowing when a Windows gfx1151 kernel "does nothing".

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.