Issues / #1277

#1277 HIP: fused IQ expert kernels via RDNA4 int8 WMMA (RX 9060 XT / gfx1200)

open · @slogomansdad · 0 コメント · GitHub で見る

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsWindows

本文

The fused int8 expert kernels (`STRATA_PF_FUSED`, `src/prefill/moe_fused*.cu`) are sm_80+ CUDA-only; on HIP,
IQ packs (IQ2_XS / IQ3_XXS) fall back to MMQ for the prompt path's expert compute. RDNA4 (gfx1200/gfx1201)
has the hardware for the fused approach — `v_wmma_f32_16x16x16_iu8` (int8 WMMA) — and the primitive is already
proven in this codebase: `STRATA_HIP_WMMA` runs the prompt attention on `v_wmma_f32_16x16x16_f16` at 7.2-7.5x the
portable kernel on gfx12 (PR #329's kernel).

0.1.40 shipped RDNA3.5 matrix-core GEMM kernels (#313) and the gfx1151 auto-switches, which is why this seems
like the natural next step in the same direction: a HIP build of the fused IQ expert path against RDNA4's int8
WMMA, with MMQ as the fallback it already is.

Expected gain (from the NVIDIA measurements in 0.1.36): IQ2_XS prompts +12% at 4K, +3% at 32K. On gfx1200 the
MMQ-vs-WMMA gap has never been measured — that's part of the request.

**Card and numbers for reference:** RX 9060 XT 16 GB (gfx1200), Windows 11, Ryzen 7 5700X (AVX2, 8C/16T),
32 GB DDR4, NVMe. Current state on 0.1.40.1, IQ3_XXS at 128K, tuned config (budget 15 GiB + KV streaming +
WMMA + calibrated hipBLASLt table): prefill 592 tok/s at 126K / ~508 at 13.5K, decode 22.3 tok/s at 126K /
28.6-30.3 short. GPU compute engine sits at 91-92% utilization during decode (sampled via GPU Engine
counters).

関連リンク

インストール・モデル・リリースへの站内リンク。