Pull requests / #1151

#1151 hip: the FP32 tiled QSA block scorer on HIP cards (opt-in STRATA_SELECT_SIMT=1, as on CUDA; a 128K prompt +7.4% on gfx1030)

open · @xjc10 · 0 コメント · GitHub で見る

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants

本文

Without matrix cores the prompt path scores the QSA blocks on the warp kernel: 2.7 ms per 256 queries at a 128K context on an RX 6900 XT, the one prompt term that grows with the context. The FP32 tiled scorer of 65ce329 (konijiwa110's #742) is plain FP32 arithmetic, so it runs on HIP cards as it is. This PR lets `STRATA_SELECT_SIMT=1` reach it on HIP, opt-in as on CUDA (another summation order than the warp kernel, so not bitwise, so off by default).

**The tiling is a template parameter.** Each score is still one fmaf chain over k = 0..127 in order and the same relu sum over the heads, so the tiling only decides which thread computes which score. On an RX 6900 XT, 14 tilings at 32K and 128K gave scores byte-identical to the original kernel, also with the card restricted to 60 and 72 of its 80 CUs (`HSA_CU_MASK`, which slows the original tiling 1.40x at 60 CUs, so it takes effect), and the fastest tiling was the same at every width. CUDA keeps the tiling it was written with (32/64/32/4); gfx103x runs 32/128/16/8; other HIP cards the original.

`qsa_select_bench`, one card, 256 queries (FP64 gate passed, selections identical to the warp kernel 256/256):

| | warp kernel | original tiling | 32/128/16/8 |
|---|---:|---:|---:|
| 128K | 2.69 ms | 0.695 ms | 0.587 ms |
| 32K | 0.67 ms | 0.186 ms | 0.169 ms |

End to end, 2x RX 6900 XT layer split, IQ3_S, 0.1.40.3 with this change, one cold server per arm, ABBA, same binary:

| prompt | warp kernel | `STRATA_SELECT_SIMT=1` |
|---|---:|---:|
| 128K | 1,735 / 1,727 tok/s | 1,861 / 1,858 tok/s (+7.3%, +7.5%) |
| 32K | 1,571 / 1,561 tok/s | 1,584 / 1,574 tok/s (+0.8%, +0.8%) |

Decode is not compared: the scores' summation order changes the text. Each arm's two starts gave the same text, except one 128K request between the two warp-kernel starts. No stalls. Per-request numbers and the tiling tables: `bench/results/2026-10-07-rdna2-qsa-select-simt`.

**What it replaces.** This PR first ran these scores as rocBLAS SGEMMs with the solution measured on the card at engine start, and later a per-device cache of those picks. The tiled scorer is faster (0.587 against 0.817 ms at 128K), needs no rocBLAS, no measurement and no cache, and gives the same bytes on every card of an architecture. Those commits are gone from the branch; #1167 (the prompt GEMMs' rocBLAS table) does not depend on this PR any more.

Tests: a HIP ctest `qsa_select_bench_fp32_tiled` (131K, `STRATA_SELECT_SIMT=1`). HIP ctest on the RX 6900 XT: 96 tests, 86 pass, 6 skip (WMMA / hipBLASLt / fused-kernel tests that do not apply to gfx1030), 4 fail for reasons outside this change (`expert_cache_segmented_test`: `--vram-elastic` is CUDA-only; `ple_parity`: no Q2_0 PLE file here; `platform_memory_test`: `ulimit -l`; `expert_multi_test`: no AVX-512). The CUDA path's kernel is now the template's 32/64/32/4 instance with the same arithmetic; I could not run it on a CUDA card here.

Developed with an AI coding assistant; every number above was measured on 2x RX 6900 XT (gfx1030, PCIe 4.0 x8 each) / Ryzen 5 5600X, ROCm 10.0.

関連リンク

インストール・モデル・リリースへの站内リンク。