Pull requests / #325
#325 AMD working on Windows (9070XT tested)
closed · @dvasdekis · 0 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Beschreibung
Windows HIP for gfx1200/gfx1201 (port of #247) + the GPU-selection fix
Windows HIP was explicitly out of scope in `docs/AMD_HIP.md`. This adds it, on top of
@main at 30ec18e. It carries @jagsan-cyber's #247 work (its one substantive commit,
`aea1ec1`, re-resolved onto current main), @araujoluks's #302, and two fixes that
neither had.
Kept deliberately opt-in, defaults unchanged, and nothing on the NVIDIA or Linux HIP
paths is altered.
## What this adds
## Context x quant: generation speed measured on an RX 9070 XT, 64GB DDR5
**Generation tok/s:**
| quant | 32K | 131K | 262K |
|---|---:|---:|---:|
| **IQ2_XS** | **27.5** | **24.0** | **23.2** |
| **IQ3_XXS** | **25.3** | **21.8** | **22.0** |
- **Windows detection and ROCm resolution.** `setup.py` finds the card through the
display-class registry (PCI device id, then driver name), and resolves a ROCm tree:
an installed ROCm 10 `rocm-sdk` or HIP SDK first, otherwise AMD's TheRock wheels
into `.venv`. `--backend hip` on `START-HERE.bat`, chosen by itself when there is no
usable NVIDIA card.
- **A Windows build.** CMake refuses to mix MSVC host flags with Clang HIP, so ROCm's
`clang++` compiles the host code too and the CUDA shim is force-included with `/FI`.
The `.cu` sources are unchanged; `hip_backend.cmake` still just sets `LANGUAGE HIP`.
- **gfx1200 as a target** alongside gfx1100/gfx1201.
- **The gfx1201 hipBLAS sticky error.** hipBLAS can return success and still leave
`hipErrorInvalidValue` set, which the next kernel's check would turn into an exit.
Cleared after a successful GEMM, **gated on `_WIN32`** — clearing it on Linux could
hide a real error there, so the Linux build keeps reporting it as before.
- **Multi-shard expert packs and Q8_0 embedding rows** (`iq_pack.py`), so a 3-shard
GGUF packs natively, and an F32 `ple_conv1d` is converted to FP16 at startup
because the convolution kernels read FP16.
- **#302's copy-dependency batching**, opt-in via `STRATA_HIP_COPY_BATCH` /
`STRATA_HIP_STAGE_BATCH` / `STRATA_HIP_STAGE_COPY_BATCH`, defaults unchanged at
1/1/1, guarded by `STRATA_USE_HIP && _WIN32`.
## The GPU-selection bug (#247 and #302 do not have this)
**`setup.py` detects AMD cards by display-class enumeration; `HIP_VISIBLE_DEVICES`
counts the devices the HIP runtime enumerates. The orders differ whenever an
integrated GPU is present** — the iGPU takes ordinal 0 and pushes the discrete card
to 1. setup then handed the engine the registry index, so on a PC with an iGPU the
engine was pointed at the iGPU.
This is not hypothetical; it is what happens on the test machine. `setup.py --check`
selected index 0 for the RX 9070 XT, while the runtime reports:
```
$ strata-device --list-devices
device 0: AMD Radeon(TM) Graphics (gfx1036) <- the iGPU
device 1: AMD Radeon RX 9070 XT (gfx1201) <- the card setup had chosen
```
The arch check then refuses the iGPU, so the failure lands *after* install, *after*
the ~58 GB download, and reads as "the Windows AMD port is broken" rather than "your
GPU numbering shifted by one". Any AMD PC with an iGPU hits it.
- `device_count()` and `strata-device --list-devices` enumerate without tripping the
arch check (which throws for a card the binary has no code for — exactly what a
listing needs to show).
- `setup.py:hip_ordinal()` resolves the card's HIP ordinal and records it as
`hip_ordinal` in the engine config, warning when it differs from the registry
index. `serve/server.py:hip_visible()` prefers it, falling back for configs
written before this (correct only without an iGPU).
It matches the card by the **architecture on the line after its `device N:` header**.
Matching on the name is wrong (the registry and runtime names differ), and matching
"does the arch string appear anywhere" is worse: the iGPU's arch-problem message
names *this build's* architectures, so `gfx1201 in line` matches the iGPU. That
mistake made the first version of this return 0.
## Also fixed
- **`hipHostGetDevicePointer` returns the host pointer on Windows**, not a device-side
alias — measured, both addresses print as `0x304000000`
(`tests/hip/mapped_alias`). A non-null alias is therefore *not* evidence of a usable
mapping, so `native_head.cpp` only takes the mapped-alias path when the pointer is
genuinely distinct and otherwise copies the token embedding into VRAM.
- `native_head.cpp`'s destructor freed `dev_owned_` and then `dev_`, but `load()` only
ever sets one, so a device-resident table was freed twice.
- `iq_pack.py`: the `tmp.replace()` completion mark (#173) is preserved. #247's hunk
silently drops it, which would let a torn `native_experts.txt` be read as a finished
pack.
## Not fixed, deliberately
`tests/hip/handoff` fails on gfx1201/Windows ("separate copy/ring timeout") for the
same mapped-alias reason — a `DeviceToDevice` copy into a host address publishes
nothing. It passes on Linux. The engine does not use that copy path on Windows (its
arena arrives via `cudaHostRegister`ed pinned memory), and fixing it properly means
giving the handoff a real device-side buffer, which is a change to the production path.
Recorded in `docs/AMD_HIP.md` rather than papered over.
## Test
Windows 11, **AMD Radeon RX 9070 XT** (gfx1201, 16 GB, 32 CUs), Ryzen 7 7800X3D,
63 GB RAM, ROCm 10.2.0a20260930 in `.venv`, engine 0.1.30, this branch.
- Build: `START-HERE.bat --backend hip` compiles cleanly, including the `FetchContent`ed
pinned llama.cpp/ggml as HIP. Both were open questions; both are fine.
- `strata-device --selftest` passes: arena alignment, the NaN poison path, over-allocation
refused.
- `ctest`: **42 of 46 pass, 1 skipped, 4 failed.** The failures are `hip_handoff`
(above) and `ple_parity` / `expert_parity` / `pool_test`, which all read
`pack/full/experts.bin`, a model fixture this checkout does not carry — the same
three are already documented Linux exclusions for the same reason.
- End to end, Coder IQ1_M at 32K context: 23.42 GiB in RAM, 5,047 expert slots /
9.59 GiB VRAM, 503 MiB free, **29.2 tok/s decode**, ~181 tok/s prefill. Correct
answers (`17*23` → 391, capital of France → Paris, valid `reverse_string`).
Numbers and method in `bench/results/2026-09-30-rx9070xt-windows/`.
### On #302's benchmark
#302 reports +62% PP / −38% TTFT. Measured here: **+4.6%** cold prefill (173 → 181
tok/s, 4 trials/arm, median, per-trial random salt so the prefix cache cannot serve
it, `STRATA_PREFILL_TIMING=1` both arms).
The counters show #302 doing exactly what it was written to do — dependency waits
collapse 8,270 → 1,034, matching `completion_batch=8` — but **"wait copy" is only
2.2–2.3% of GPU time in both arms**, so there was almost nothing there to win. #302's
figure came from Q2_0, where copy-wait was 2,746 ms of a 6,144 ms prefill. On a 16 GB
card holding 5,047 of 12,288 experts the run is bound by gdn (42% of GPU time), qsa
proj (13%), dequant (11%) and the gemms (11%). A user with less VRAM has more expert
traffic and will see more of #302's benefit. It is a correct change at the wrong
bottleneck for this configuration, which is why the opt-in defaults stay 1/1/1.
## The expert cache could take the VRAM allocated after it (fixed here)
Found by running the full model on a 16 GB card, not by reading the code. At
IQ3_XXS the engine sized a 6,124-slot expert cache (9.92 GiB), printed
"almost ready", and died:
```
strata serve: the penalty-history allocation failed
```
The penalty history is 0.1 MiB and under 128 KB was spare. `stage_room()` already
reserves for exactly this on the layer split - it holds back `kWindowMib` and
`kDrafterMib` because the verify windows, drafter and head are all allocated after
a stage's cache is sized. The single-GPU path reserved `vram_reserve_mib` alone,
so a profile-sized cache took the rest. The Coder survived only because its cache
is 0.33 GiB smaller.
- `kPostCacheMib` (192) beside `kWindowMib`/`kDrafterMib`, held back on the
single-GPU path too. 192 covers the measured 74.1 MiB of verify windows at
n_embd 2560, the penalty history, and the R4 hit-path scratch (which scales with
n_embd). `kDrafterMib` is deliberately not re-added - `mtp_bind` is this run's
own figure for those bytes.
- The profile-sizing pass re-read `cudaMemGetInfo` and capped against
`vram_reserve_mib` alone, handing the slack straight back. Same fix.
- The failure now names the cause and the fix rather than only reporting the fact.
At the **default** reserve, with no `--vram-reserve-mib`:
| | slots | cache | outcome |
|---|---:|---:|---|
| before | 6,124 | 9.92 GiB | crash on a 0.1 MiB allocation |
| after | 6,009 | 9.73 GiB | 662 MiB free, serves at 27.7-35.6 tok/s |
That is *more* cache than the `--vram-reserve-mib 2048` workaround managed
(5,303 slots), so this is not headroom traded for capacity. Coder unaffected:
4,946 slots, 693 MiB free, still serving.
This is a pre-existing bug on the NVIDIA path too - nothing here is HIP-specific -
so it is worth a look even if the rest of this PR is not taken.
## Credits
@jagsan-cyber for #247 (the Windows ROCm resolution, the MSVC `/FI` build, the sticky
-error, the multi-shard pack), @araujoluks for #302, and @Niko1221 for the `#247`
discussion that pointed out a Windows-HIP-only PR was welcome. The arch check,
`gfx1101`/`gfx1102` dp4a targets and the #173 pack fix in this branch are current
`main`'s, kept over #247's older copies of those hunks.
Mehr auf der Site
Links zu Install, Modellen, Releases.