Pull requests / #426
#426 Windows AMD: the HIP backend detects, builds and runs (RX 9070 XT / gfx1201)
closed · @RenZekta · 0 评论 · 在 GitHub 查看
Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
描述
## Summary
- `SETUP.bat` / `START-HERE.bat` stopped on any PC without an NVIDIA GPU ("no NVIDIA GPU found"), and `--backend hip`
refused Windows outright ("runs on Linux only"). The AMD backend now works on Windows the way it works on Linux:
setup detects the card, finds a HIP SDK, compiles the engine with it and starts the server.
- Verified end to end on real hardware — a Radeon RX 9070 XT (gfx1201, RDNA4): setup → engine build → server → a
live model session — plus the unit tests below.
- Every Windows branch is `if WIN`-guarded. Linux behaviour, paths and messages are unchanged, and the Linux tests
pass without edits.
- Tested on Windows 11 with RX 9070 XT and 32 GB RAM
## What changed
**`setup.py`** — the bulk of the change:
| Area | Change |
| --- | --- |
| Detection | `amd_gpus_win()` reads the HIP SDK's `bin\hipInfo.exe` — the runtime's own `device#`, `Name`, `totalGlobalMem`, `gcnArchName`, in the order `HIP_VISIBLE_DEVICES` indexes — and falls back to the display driver's registry (`HardwareInformation.qwMemorySize`; WMI's `AdapterRAM` stops at 4 GB) when no SDK is installed yet. The parsing is a pure `hip_info(text)` with tests; `amd_arch_from_name()` places a card from the driver name (RX 6600…6900 → gfx1032…gfx1030, RX 7600…9070, R9700 → gfx1102…gfx1201). |
| SDK discovery | `hip_sdk_root()`: `STRATA_HIP_ROOT` → `HIP_PATH` / `ROCM_PATH` → `%ProgramFiles%\AMD\ROCm\*` → a `therock-dist-*` unpacked next to the checkout → AMD's pip `rocm-sdk`. Valid only when `bin\hipcc.exe` and the ROCm clang exist. |
| `rocm_root()` | a Windows branch that returns the SDK instead of the per-family Python wheels, with an install hint when none is found. |
| `build_engine_hip()` | a Windows branch: the SDK's clang compiles **C, C++ and HIP**, `HIP_PATH`/`ROCM_PATH` are removed from the environment, every `-D` path is forward-slashed, `find_vcvars()` is required and the configure runs through a generated `build-hip.bat` (`vcvars64.bat` first). Missing VS C++ tools and a missing SDK are explicit failures with install instructions. |
| `build_vision_cpu()` | passes `find_vcvars()` + `build-vision.bat` on Windows (the new `--vision cpu` path otherwise raises `IsADirectoryError` writing to the repo root); the CUDA vision build already did this. |
| `get_prebuilt()` | returns `None` for a hip stamp — `int("gfx1201")` used to raise when a CUDA backend was requested on a PC whose `engine/` holds an AMD build. |
| `main()` | AMD is chosen by itself on a PC with no usable NVIDIA card; step 1 reports the card, the SDK and `amd_problem()`; `cfg["gpu"]` is written for `hip` on Windows (the server turns it into `HIP_VISIBLE_DEVICES`); `--gpus all` works; `--backend` help, the step-1 table and the rebuild hints use the OS's own launcher (`START-HERE.bat`). |
**Engine (2 lines)** — the runtime difference WDDM makes:
- `src/core/session.cpp`: the token-graph flush interval is `0` under `STRATA_USE_HIP && _WIN32`. On Windows a
mapped-pinned write becomes CPU-visible only once the HIP driver is entered; polling with the default 2 ms/layer
interval stalls the GPU instead of overlapping with it.
- `tests/hip/handoff.cpp`: the handoff test's two spin loops call `hipStreamQuery` under `_WIN32` for the same
reason — without it the test waits on writes the CPU cannot see yet.
**Docs, tests, other**
- `docs/AMD_HIP.md`: a new `## Windows (experimental)` section (requirements, detection, the configure rules, DLL/PATH
handling, limits), and Windows is no longer listed as out of scope.
- `README.md` and the `START-HERE.bat` header: AMD works on Windows, with a `START-HERE.bat --backend hip` example.
- `tools/test_setup_amd.py`: three new tests (`hip_info` parsing, `amd_arch_from_name` incl. RDNA2, an unsupported
iGPU) and `WIN=False` pinned in `GpuLists` so the Linux-path tests stay Linux-path tests.
- `.gitignore`: the generated `build-hip.bat`.
## Why these choices
1. **One compiler for all three languages.** CMake refuses Clang for HIP next to MSVC for C/C++ ("Use either Clang or
MSVC ... for all of C, C++, and/or HIP"), so the SDK's clang compiles all three, with `vcvars64.bat` supplying the
MSVC and Windows SDK headers.
2. **`HIP_PATH`/`ROCM_PATH` are removed for the configure.** clang adopts `HIP_PATH` as `--rocm-path` and then cannot
find this layout's device bitcode ("cannot find ROCm device library").
3. **Every `-D` path is forward-slashed.** CMake writes some of them verbatim, where a backslash is an invalid escape
(`Invalid character escape '\a'`).
4. **Detection by `hipInfo`, not WMI.** It is the runtime's own device ordering (exactly what `HIP_VISIBLE_DEVICES`
indexes) and it carries `gcnArchName`; the registry path covers the first run, before an SDK exists.
5. **No server change.** `serve/server.py` already prepends `cfg["lib_dirs"]` to `PATH` on Windows and sets
`HIP_VISIBLE_DEVICES` for `backend == "hip"` — the config's `gpu` is simply made to be the `hipInfo` number.
## Testing
Machine: Windows 11 Pro (build 28000) · Ryzen 9 5900X, 32 GB RAM · Radeon RX 9070 XT 16 GB (gfx1201, wave32) ·
Visual Studio 2022 (MSVC 14.44) · CMake 4.3.2 + Ninja · TheRock Windows dist 7.14.0a20260612 (hipBLASLt 1.4.0).
- `START-HERE.bat --check` → **exit 0**: `GPU 0: AMD Radeon RX 9070 XT, 16 GB VRAM - can be used`,
`[ok] HIP SDK: …` — no "no NVIDIA GPU found", no "runs on Linux only".
- Engine built through setup's own path: `build-hip.bat` → configure + 127 Ninja steps → `engine/strata.exe`, with
`BUILD.json` = `{"source": "local-hip", "backend": "hip", "version": "0.1.32", "archs": ["gfx1201"],
"lib_dirs": ["…\\therock-dist…\\bin"], "hip_root": …}`. A second run reports
`[ok] engine already built for this PC`.
- `build-hip\strata-device.exe --selftest` → `HIP arch gfx1201 wave32 (compiled for gfx1201)`, memory plan `FITS`,
`strata-device selftest OK`.
- `engine\strata-vision.exe` (the `--vision cpu` encoder) builds on Windows with the `build-vision.bat` fix.
- Unit tests: **129 tests in 11 modules pass**, including the new Windows ones.
- **End to end:** a full setup run and a live model session on this card with the `coder` family
(Qwen3.8-Flash-Next-Coder), verified on this machine.
## Limitations
- Images only through the CPU encoder (`--vision cpu`); no GPU image encoder, no calibration.
- The Monitor shows no GPU statistics on Windows (it reads them from Linux sysfs).
- No hipBLASLt tuning table for gfx1201 with hipBLASLt 1.4.0, so setup warns and prompts use plain hipBLAS (same
answers, a little slower); the `gfx1201-hipblaslt-100500.txt` table in `tools/hip` applies to hipBLASLt 1.5.0.
- A layer split across two GPU families is refused on Windows (the per-family wheel indexes have no Windows
equivalent wired up); cards of one family work, `--gpus all` included.
- Only this machine was tested. gfx1030 / gfx1100 / gfx1101 / gfx1200 on Windows are untested here.
## Notes
- `main` has moved to 0.1.33 since this branch; `git merge-tree upstream/main HEAD` reports no conflicts.
## Checklist
Linux behaviour untouched (the existing tests pass as-is; `GpuLists`/`KfdDetection` still run with `WIN=False`)
New parsing/detection logic covered by tests (`hip_info`, `amd_arch_from_name`)
Docs updated: `docs/AMD_HIP.md`, `README.md`, `START-HERE.bat`
Verified on hardware, end to end
Later will test with more RAM
AI assistance: Code and PR assisted by MiMo-V2.6-Flash站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。