Pull requests / #256
#256 HIP: support RDNA4 gfx1200 (RX 9060 XT)
closed · @Efeisot · 0 コメント · GitHub で見る
BenchmarksSetup & installAMD / HIPModels & quantsDocumentation
本文
Continues the gfx1200 part of #176 on current main (0.1.29, `d6708a4`), following the review there and in #192. One commit (`a1128ec` on branch `hip-gfx1200`), +197/−16, seven files:
1. **The arch check, comparing against the compiled arch (as asked in #192).** `cmake/hip_backend.cmake` keeps a `STRATA_HIP_ARCHS` list (`gfx1100; gfx1200`), builds for exactly one arch from it, and exports it as `STRATA_HIP_ARCH`. `src/core/device.cu` compares the running card's `gcnArchName` against that macro (still requiring wave32) and fails at startup with a rebuild hint when they differ — so a gfx1100 build carried to a gfx1200 card no longer runs until an `invalid device function`. Demonstrated on the card: a gfx1100-compiled `strata-device` prints
`HIP backend was compiled for gfx1100 but device AMD Radeon RX 9060 XT reports gfx1200; rebuild with -DCMAKE_HIP_ARCHITECTURES=gfx1200 (each build runs on the arch it was compiled for)`
while the gfx1200 binary plans normally.
2. **setup.py:** `AMD_ARCHS = ("gfx1100", "gfx1200")`, `AMD_NAMES` gains the RX 9060 XT, and the two "not supported" messages list both cards. Setup already passes the detected arch as `CMAKE_HIP_ARCHITECTURES`, so no other change is needed there.
3. **The hipBLASLt table** `tools/hip/gfx1200-hipblaslt-100202.txt`: 67 rows (bf16 + f16) calibrated on the card with `tools/hip/tune_hipblaslt`, header `STRATA_HIPBLASLT_TUNING_V1 gfx1200 100202`. The engine refuses a table made for another arch or hipBLASLt version and falls back to plain hipBLAS, as before; setup finds it by its `{arch}-hipblaslt-*.txt` name.
4. **The doc** `docs/AMD_HIP_GFX1200.md`: deltas, hand build, tests, measured numbers.
5. **The launcher** under `tools/hip/gfx1200-run.sh` (per the #176 review — no root script), used for the measurements below.
Left out as requested: no `PR_BODY.md`, no `include/math_constants.h` (0.1.29 doesn't need it), no unroll pragma, nothing touching `prefill.cpp`.
## Testing
ROCm 7.2 in /opt/rocm (hipBLASLt 100202), RX 9060 XT 16 GiB (gfx1200, wave32) + Ryzen 9 7950X + 64 GiB RAM, pinned llama.cpp `3cf03257f` as the ggml source.
- ctest excluding `ple_parity` and `platform_memory_test` (as in docs/AMD_HIP.md): **32/32** on 0.1.29, including `hip_prefill_hipblaslt_gemm` against the new table.
- End-to-end (Qwen3.8-Flash-Next IQ1_M Coder, greedy, MTP `--spec 4 --spec-min-p 0.5`, the launcher's flags): prefill **540 tok/s** on a 2,374-token prompt, **748–753 tok/s** on a 65K prompt (`--kv int8 --kv-resident 65536 --prefill 16384`), **725 tok/s** on a 130K prompt — clean start to finish. Decode scales with MTP draft acceptance: 27.1 tok/s at 0.69 acceptance on this prompt, 31.0 at 1.00 on the code prompt from #176. For reference, llama.cpp HIP measured 20 decode / 450 prefill on this card.
- The cross-arch refusal above, from a gfx1100 build of this branch.
One note: in #176 I flagged that raising `attn_batch` (32→64) overflowed 0.1.27's borrowed prefill buffers at any chunk size. On 0.1.29 it fits again (and measures within a few percent of 32 on this card), so that headroom issue looks resolved upstream — nothing in this PR touches it.
🤖 Generated with Codebuff関連リンク
インストール・モデル・リリースへの站内リンク。