Issues / #1277
#1277 HIP: fused IQ expert kernels via RDNA4 int8 WMMA (RX 9060 XT / gfx1200)
open · @slogomansdad · 0 Kommentare · Auf GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsWindows
Beschreibung
The fused int8 expert kernels (`STRATA_PF_FUSED`, `src/prefill/moe_fused*.cu`) are sm_80+ CUDA-only; on HIP, IQ packs (IQ2_XS / IQ3_XXS) fall back to MMQ for the prompt path's expert compute. RDNA4 (gfx1200/gfx1201) has the hardware for the fused approach — `v_wmma_f32_16x16x16_iu8` (int8 WMMA) — and the primitive is already proven in this codebase: `STRATA_HIP_WMMA` runs the prompt attention on `v_wmma_f32_16x16x16_f16` at 7.2-7.5x the portable kernel on gfx12 (PR #329's kernel). 0.1.40 shipped RDNA3.5 matrix-core GEMM kernels (#313) and the gfx1151 auto-switches, which is why this seems like the natural next step in the same direction: a HIP build of the fused IQ expert path against RDNA4's int8 WMMA, with MMQ as the fallback it already is. Expected gain (from the NVIDIA measurements in 0.1.36): IQ2_XS prompts +12% at 4K, +3% at 32K. On gfx1200 the MMQ-vs-WMMA gap has never been measured — that's part of the request. **Card and numbers for reference:** RX 9060 XT 16 GB (gfx1200), Windows 11, Ryzen 7 5700X (AVX2, 8C/16T), 32 GB DDR4, NVMe. Current state on 0.1.40.1, IQ3_XXS at 128K, tuned config (budget 15 GiB + KV streaming + WMMA + calibrated hipBLASLt table): prefill 592 tok/s at 126K / ~508 at 13.5K, decode 22.3 tok/s at 126K / 28.6-30.3 short. GPU compute engine sits at 91-92% utilization during decode (sampled via GPU Engine counters).
Mehr auf der Site
Links zu Install, Modellen, Releases.