Pull requests / #325

#325 AMD working on Windows (9070XT tested)

closed · @dvasdekis · 0 comments · View on GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Description

Windows HIP for gfx1200/gfx1201 (port of #247) + the GPU-selection fix

Windows HIP was explicitly out of scope in `docs/AMD_HIP.md`. This adds it, on top of
@main at 30ec18e. It carries @jagsan-cyber's #247 work (its one substantive commit,
`aea1ec1`, re-resolved onto current main), @araujoluks's #302, and two fixes that
neither had.

Kept deliberately opt-in, defaults unchanged, and nothing on the NVIDIA or Linux HIP
paths is altered.

## What this adds

## Context x quant: generation speed measured on an RX 9070 XT, 64GB DDR5

**Generation tok/s:**

| quant | 32K | 131K | 262K |
|---|---:|---:|---:|
| **IQ2_XS** | **27.5** | **24.0** | **23.2** |
| **IQ3_XXS** | **25.3** | **21.8** | **22.0** |

- **Windows detection and ROCm resolution.** `setup.py` finds the card through the
  display-class registry (PCI device id, then driver name), and resolves a ROCm tree:
  an installed ROCm 10 `rocm-sdk` or HIP SDK first, otherwise AMD's TheRock wheels
  into `.venv`. `--backend hip` on `START-HERE.bat`, chosen by itself when there is no
  usable NVIDIA card.
- **A Windows build.** CMake refuses to mix MSVC host flags with Clang HIP, so ROCm's
  `clang++` compiles the host code too and the CUDA shim is force-included with `/FI`.
  The `.cu` sources are unchanged; `hip_backend.cmake` still just sets `LANGUAGE HIP`.
- **gfx1200 as a target** alongside gfx1100/gfx1201.
- **The gfx1201 hipBLAS sticky error.** hipBLAS can return success and still leave
  `hipErrorInvalidValue` set, which the next kernel's check would turn into an exit.
  Cleared after a successful GEMM, **gated on `_WIN32`** — clearing it on Linux could
  hide a real error there, so the Linux build keeps reporting it as before.
- **Multi-shard expert packs and Q8_0 embedding rows** (`iq_pack.py`), so a 3-shard
  GGUF packs natively, and an F32 `ple_conv1d` is converted to FP16 at startup
  because the convolution kernels read FP16.
- **#302's copy-dependency batching**, opt-in via `STRATA_HIP_COPY_BATCH` /
  `STRATA_HIP_STAGE_BATCH` / `STRATA_HIP_STAGE_COPY_BATCH`, defaults unchanged at
  1/1/1, guarded by `STRATA_USE_HIP && _WIN32`.

## The GPU-selection bug (#247 and #302 do not have this)

**`setup.py` detects AMD cards by display-class enumeration; `HIP_VISIBLE_DEVICES`
counts the devices the HIP runtime enumerates. The orders differ whenever an
integrated GPU is present** — the iGPU takes ordinal 0 and pushes the discrete card
to 1. setup then handed the engine the registry index, so on a PC with an iGPU the
engine was pointed at the iGPU.

This is not hypothetical; it is what happens on the test machine. `setup.py --check`
selected index 0 for the RX 9070 XT, while the runtime reports:

```
$ strata-device --list-devices
device 0: AMD Radeon(TM) Graphics (gfx1036)   <- the iGPU
device 1: AMD Radeon RX 9070 XT (gfx1201)      <- the card setup had chosen
```

The arch check then refuses the iGPU, so the failure lands *after* install, *after*
the ~58 GB download, and reads as "the Windows AMD port is broken" rather than "your
GPU numbering shifted by one". Any AMD PC with an iGPU hits it.

- `device_count()` and `strata-device --list-devices` enumerate without tripping the
  arch check (which throws for a card the binary has no code for — exactly what a
  listing needs to show).
- `setup.py:hip_ordinal()` resolves the card's HIP ordinal and records it as
  `hip_ordinal` in the engine config, warning when it differs from the registry
  index. `serve/server.py:hip_visible()` prefers it, falling back for configs
  written before this (correct only without an iGPU).

It matches the card by the **architecture on the line after its `device N:` header**.
Matching on the name is wrong (the registry and runtime names differ), and matching
"does the arch string appear anywhere" is worse: the iGPU's arch-problem message
names *this build's* architectures, so `gfx1201 in line` matches the iGPU. That
mistake made the first version of this return 0.

## Also fixed

- **`hipHostGetDevicePointer` returns the host pointer on Windows**, not a device-side
  alias — measured, both addresses print as `0x304000000`
  (`tests/hip/mapped_alias`). A non-null alias is therefore *not* evidence of a usable
  mapping, so `native_head.cpp` only takes the mapped-alias path when the pointer is
  genuinely distinct and otherwise copies the token embedding into VRAM.
- `native_head.cpp`'s destructor freed `dev_owned_` and then `dev_`, but `load()` only
  ever sets one, so a device-resident table was freed twice.
- `iq_pack.py`: the `tmp.replace()` completion mark (#173) is preserved. #247's hunk
  silently drops it, which would let a torn `native_experts.txt` be read as a finished
  pack.

## Not fixed, deliberately

`tests/hip/handoff` fails on gfx1201/Windows ("separate copy/ring timeout") for the
same mapped-alias reason — a `DeviceToDevice` copy into a host address publishes
nothing. It passes on Linux. The engine does not use that copy path on Windows (its
arena arrives via `cudaHostRegister`ed pinned memory), and fixing it properly means
giving the handoff a real device-side buffer, which is a change to the production path.
Recorded in `docs/AMD_HIP.md` rather than papered over.

## Test

Windows 11, **AMD Radeon RX 9070 XT** (gfx1201, 16 GB, 32 CUs), Ryzen 7 7800X3D,
63 GB RAM, ROCm 10.2.0a20260930 in `.venv`, engine 0.1.30, this branch.

- Build: `START-HERE.bat --backend hip` compiles cleanly, including the `FetchContent`ed
  pinned llama.cpp/ggml as HIP. Both were open questions; both are fine.
- `strata-device --selftest` passes: arena alignment, the NaN poison path, over-allocation
  refused.
- `ctest`: **42 of 46 pass, 1 skipped, 4 failed.** The failures are `hip_handoff`
  (above) and `ple_parity` / `expert_parity` / `pool_test`, which all read
  `pack/full/experts.bin`, a model fixture this checkout does not carry — the same
  three are already documented Linux exclusions for the same reason.
- End to end, Coder IQ1_M at 32K context: 23.42 GiB in RAM, 5,047 expert slots /
  9.59 GiB VRAM, 503 MiB free, **29.2 tok/s decode**, ~181 tok/s prefill. Correct
  answers (`17*23` → 391, capital of France → Paris, valid `reverse_string`).
  Numbers and method in `bench/results/2026-09-30-rx9070xt-windows/`.

### On #302's benchmark

#302 reports +62% PP / −38% TTFT. Measured here: **+4.6%** cold prefill (173 → 181
tok/s, 4 trials/arm, median, per-trial random salt so the prefix cache cannot serve
it, `STRATA_PREFILL_TIMING=1` both arms).

The counters show #302 doing exactly what it was written to do — dependency waits
collapse 8,270 → 1,034, matching `completion_batch=8` — but **"wait copy" is only
2.2–2.3% of GPU time in both arms**, so there was almost nothing there to win. #302's
figure came from Q2_0, where copy-wait was 2,746 ms of a 6,144 ms prefill. On a 16 GB
card holding 5,047 of 12,288 experts the run is bound by gdn (42% of GPU time), qsa
proj (13%), dequant (11%) and the gemms (11%). A user with less VRAM has more expert
traffic and will see more of #302's benefit. It is a correct change at the wrong
bottleneck for this configuration, which is why the opt-in defaults stay 1/1/1.

## The expert cache could take the VRAM allocated after it (fixed here)

Found by running the full model on a 16 GB card, not by reading the code. At
IQ3_XXS the engine sized a 6,124-slot expert cache (9.92 GiB), printed
"almost ready", and died:

```
strata serve: the penalty-history allocation failed
```

The penalty history is 0.1 MiB and under 128 KB was spare. `stage_room()` already
reserves for exactly this on the layer split - it holds back `kWindowMib` and
`kDrafterMib` because the verify windows, drafter and head are all allocated after
a stage's cache is sized. The single-GPU path reserved `vram_reserve_mib` alone,
so a profile-sized cache took the rest. The Coder survived only because its cache
is 0.33 GiB smaller.

- `kPostCacheMib` (192) beside `kWindowMib`/`kDrafterMib`, held back on the
  single-GPU path too. 192 covers the measured 74.1 MiB of verify windows at
  n_embd 2560, the penalty history, and the R4 hit-path scratch (which scales with
  n_embd). `kDrafterMib` is deliberately not re-added - `mtp_bind` is this run's
  own figure for those bytes.
- The profile-sizing pass re-read `cudaMemGetInfo` and capped against
  `vram_reserve_mib` alone, handing the slack straight back. Same fix.
- The failure now names the cause and the fix rather than only reporting the fact.

At the **default** reserve, with no `--vram-reserve-mib`:

| | slots | cache | outcome |
|---|---:|---:|---|
| before | 6,124 | 9.92 GiB | crash on a 0.1 MiB allocation |
| after | 6,009 | 9.73 GiB | 662 MiB free, serves at 27.7-35.6 tok/s |

That is *more* cache than the `--vram-reserve-mib 2048` workaround managed
(5,303 slots), so this is not headroom traded for capacity. Coder unaffected:
4,946 slots, 693 MiB free, still serving.

This is a pre-existing bug on the NVIDIA path too - nothing here is HIP-specific -
so it is worth a look even if the rest of this PR is not taken.

## Credits

@jagsan-cyber for #247 (the Windows ROCm resolution, the MSVC `/FI` build, the sticky
-error, the multi-shard pack), @araujoluks for #302, and @Niko1221 for the `#247`
discussion that pointed out a Windows-HIP-only PR was welcome. The arch check,
`gfx1101`/`gfx1102` dp4a targets and the #173 pack fix in this branch are current
`main`'s, kept over #247's older copies of those hunks.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.