Pull requests / #1570
#1570 HIP gfx12: native int8-WMMA prompt experts on gfx1200/gfx1201 (opt-in STRATA_PF_FUSED=1, #1277)
open · @jkuepker · 0 评论 · 在 GitHub 查看
BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
描述
## Title Issue: Resolves #1277 ## Summary `STRATA_PF_FUSED=1` runs the native IQ packs' prompt experts on the int8-WMMA kernels in `src/prefill/moe_fused_iq.cu` (`native_w11_kernel`). Those kernels were gfx11-only: on RDNA4 the flag was accepted and MMQ stayed in use, with no banner. This PR makes them build and run on gfx1200 / gfx1201. The arithmetic is the same as on gfx11; the difference is gfx12's WMMA lane layout. The path stays **opt-in**. It rounds differently from MMQ, and `include/strata/core/arch_defaults.hpp` holds only bit-exact switches, so this PR does not touch `arch_defaults`. `STRATA_PF_FUSED_NATIVE=0` keeps MMQ with the flag set, as before. Measured on a Radeon AI PRO R9700 (gfx1201), Coder IQ1_M pack, MMQ (default) against `STRATA_PF_FUSED=1`: | Check | Result | |---|---| | Parity against the int8 model (`hip_prefill_fused_iq`: all 10 IQ gate/up + down pairs, plus Q4_K / Q5_K) | pass; rel RMS <= 9.1e-5, worst row <= 2.4e-3. Whole-layer MMQ vs fused at the real shape: rel RMS 1.3-1.9e-4, worst row 2.9e-3 to 8.3e-3 | | Kernels vs MMQ, 48 layers | 1.48x / 1.20x / 1.14x at 4K / 16K / 32K (fused 359.9 / 1,372 / 2,701 ms) | | Prompt time in the engine, server-reported (3 alternating starts per arm, 3 requests per size) | 4K +12% to +19% depending on the start (1,612 -> 1,353 ms here; see Rebased build), 16K 5,579 -> 5,147 ms (+8.4%), 32K 11,890 -> 10,962 ms (+8.5%). localeval prompt speed 2,547 -> 3,020 / 2,988 -> 3,235 / 2,851 -> 3,092 tok/s | | First-token KL against the FP16 prompt path (`STRATA_PREFILL_MMQ=0`), mean | 4K (n=48): MMQ 3.78e-4, fused 4.17e-4, ratio 1.10. 16K (n=8): 1.32e-4 vs 1.45e-4, ratio 1.10. 32K (n=8): 4.47e-3 vs 7.37e-4, ratio 0.16 | | Top-1 against the FP16 path | 48/48 at 4K for both, 8/8 at 16K and 32K | | Long-context retrieval (qcheck) | 4/4 (needles at 4,181 / 16,467 / 57,509 tokens, plus a code check) | | Determinism | MMQ and fused outputs each bit-identical across 3 repeats (the FP16 reference is not) | Read the quality rows with their limits: - The 4K KL ratio is a noisy point estimate: bootstrap 95% CI 0.75-2.28, median per-prompt ratio 0.77, fused lower on 25 of 48 prompts. Against the FP16 path this sample cannot tell fused from MMQ. On two prompts (q4k-044, q4k-045) fused deviates about 10x more than MMQ; on two others (q4k-015, q4k-027) it deviates about as much less. - Every argmax in all 207 dumps was the same token (the reasoning opener), so top-1 says little here. - Fused is closer to MMQ than MMQ is to the FP16 path (KL 0.56x / 0.56x / 0.36x at 4K / 16K / 32K). ## What changed - `src/prefill/moe_fused_iq.cu` - Compile gating: `STRATA_NAT_W12` for `__gfx1200__` / `__gfx1201__`, `STRATA_NAT_WMMA` for either family. The kernel calls the `_gfx12` spelling of the iu8 WMMA builtin (the un-suffixed one does not compile on gfx12). Any other target keeps the trapping body. - Kernel: a gfx12 fragment is 8 int8 per lane and the two lane halves hold different k (8*hi .. 8*hi+7), where gfx11 replicates them. A and B are loaded as `uint2` at the matching offsets; each 16-value k step is still one WMMA, so the per-16 scales of IQ2_S / IQ2_XS are untouched. A lane's 8 accumulator rows are consecutive, so a feature's gate and up sit in the same lane: SwiGLU needs no `__shfl_xor`, and the down tile is stored as two `float4` per lane. - Host side: `dev_info()` takes the gfx11 list or gfx12 and also checks that the loaded kernel is the real one (a gfx12-generic or SPIR-V build on a listed card gets the trap stub, which still passes `hipFuncGetAttributes` and the occupancy query). `native_supported()` reads the `STRATA_PF_FUSED=1` opt-in itself, because `fused::available()` / `requested()` belong to the Q2_0 pack's `expert_w11_kernel`, which is not ported and stays false on gfx12. Off gfx12 the result is unchanged. - Grid factor per format: 3 blocks per WGP for IQ2_S / IQ2_XS gate/up on gfx12 (LDS and VGPRs), 4 for the rest, from a small table because the occupancy API models 64 KB of LDS per multiprocessor. `STRATA_PF_OCC_W=N` overrides it. gfx11 keeps 4. - `experts_native()` refuses a `dm` that is not 16-byte aligned on gfx12 (the `float4` stores); the prompt path's buffers always are. - `include/strata/kernels/gfx_arch.hpp`: `gfx_arch_is_gfx12_wmma()`, an exact gfx1200 / gfx1201 match (never a prefix), like the gfx11 helper. - `src/prefill/moe_fused.cu`: the device check compares the gfx11 list (`gfx_arch_is_gfx11_wmma`) instead of the `"gfx11"` prefix. Same result on every listed part (gfx1100 / 1101 / 1102 / 1150 / 1151); a gfx11 part outside the list no longer matches by prefix. `moe_fused_iq.cu` does the same in `dev_info()`. - `src/prefill/prefill.cpp` (`fused_ring()`): a native pack asks `native_supported()` alone, where it used to ask `fused::enabled()` first. On gfx12 `enabled()` is false while the native layers do run fused, so without this they would keep MMQ's buffer layout and ring. For the Q2_0 pack the check is still `fused::enabled()`. - `include/strata/prefill/moe_fused.hpp`, `moe_fused_iq.hpp`: comments only (the gfx12 case, the 16-byte alignment). - `tests/cuda/prefill_fused_iq_test.cpp`: runs on gfx12 (the skip gate also asks `native_supported()`), adds the other four pairs of the Coder IQ1_M pack, `S20_E` sets the expert count (256 for the Coder pack), and the real-shape check of fused against MMQ is now rel RMS <= 1e-3 and worst row <= 2e-2, where it was rel RMS <= 0.05. - `docs/AMD_HIP.md`: a bullet in the RDNA4 speed switches (what the flag does on gfx12, the start-up line, the numbers above, what was not run). `docs/DETAILS.md`: the `STRATA_PF_FUSED=1` sentence says where the native kernels run on HIP. - `CMakeLists.txt`, `include/strata/kernels/gfx_arch.hpp`: comments that said these kernels are gfx11 only now include gfx12 (comments only). ## Extra Notes ### How it was tested - Hardware and toolchain: Radeon AI PRO R9700 (gfx1201), Linux, ROCm 7.14 toolchain, Coder IQ1_M pack. HIP build with `-DCMAKE_HIP_ARCHITECTURES=gfx1201`. - In-engine A/B: one frozen binary for both arms; the two server configs differ only in `STRATA_PF_FUSED=1` and the log path. 3 alternating starts per arm (A B A B A B). Speed from localeval at 4,096 / 16,384 / 32,768 tokens, 3 requests per size, 128 generated tokens, no `STRATA_PREFILL_TIMING`. 54 measured requests, every one with 0 reused tokens. The "prompt experts on the fused int8 kernels" line appeared in every fused start and in no MMQ start. The gain in each start pair was +15.2 / +19.1 / +19.3% (4K), +8.2 / +8.3 / +8.5% (16K) and +8.7 / +8.5 / +8.6% (32K); the spread between starts of one arm was 0.1-3.7%. - Part of the gain is layout, not only the kernels: the fused layout borrows 5,000 expert-cache slots instead of 6,490 and, at 4K, streams fewer RAM blobs (1,644 vs 1,986). That probably explains why 16K also gains in phases outside the MoE path. - Quality: first-token logits per request from `--serve` (`STRATA_DUMP_FIRST_LOGITS`; the run used a local 9-line serve-loop hook that overwrote one file per request; it is not part of this PR, and upstream has since added the same capability in 421e163, which writes `<path>.<n>`), KL computed in float64 from the raw vocab rows. The FP16 reference arm ran the same 66 request bodies as the other two arms (48 x 4K, 8 x 16K, 8 x 32K, plus a determinism canary), all with 0 reused tokens. - `hip_prefill_fused_iq` passes on the R9700; `hip_prefill_fused_moe` exits 77 (skip), because the Q2_0 pack's kernels are not ported to gfx12. ### Rebased build The numbers above were taken on the development branch before it was rebased onto fb58e0d. The kernel, header and test files in this PR are byte-identical to what was measured; `prefill.cpp` (5 lines) now sits on upstream's newer version. The rebased port commit was built for gfx1201 and run again on the R9700 (the second commit changes only docs and comments): - `hip_prefill_fused_iq` passes, with output byte-identical to the pre-rebase logs (all 10 IQ pairs plus Q4_K / Q5_K, and the same output hash for the IQ2_S / Q2_0 real-shape layer). `hip_prefill_fused_moe` exits 77 (skip) and `arch_defaults_test` passes. `ctest -R "hip_prefill_fused|arch_defaults"`: 3 of 3 pass or skip. `hip_prefill_fused_iq` takes 54 s against its 600 s ctest timeout. - In the engine: 3 alternating starts per arm, localeval at 4,096 / 16,384 tokens, 3 requests per size, every request with 0 reused tokens. Median prompt time per start: - 4K: MMQ 1,553-1,563 ms, fused 1,349-1,393 ms. Per start pair that is +15.8 / +11.9 / +15.1%; one fused start ran in a slower mode (1,393 ms). - 16K: MMQ 5,551-5,557 ms, fused 5,128-5,129 ms, +8.2 / +8.3 / +8.3%. - MMQ at 4K is about 3% faster than before the rebase, so the 4K gain is somewhat smaller on current main: +12% to +16% here, against +15% to +19% before. - qcheck: 4/4 in both arms, with the same answer hashes. ### For reviewers - The test's real-shape gate (fused against MMQ) went from rel RMS <= 0.05 to rel RMS <= 1e-3 and worst row <= 2e-2. The 4 added pairs also run on every backend. Both were measured only on gfx1201 (rel RMS 1.3-1.9e-4, worst row 2.9e-3 to 8.3e-3). If `prefill_fused_iq_test` fails at the new bound on sm_80+ or gfx1151, loosen the bound; the kernels are not the cause. - On gfx11 the changes are meant to change nothing. The places to check on a gfx1151: - `dev_info()` in `src/prefill/moe_fused_iq.cu`, around line 1038: the exact gfx11 list instead of the prefix, plus the LDS check; - the same list check in `src/prefill/moe_fused.cu`, around line 611. `hip_prefill_fused_iq` should still pass there, and `STRATA_PF_FUSED=1` should still print the banner. ### Not tested - gfx1200 hardware (RX 9060 / 9060 XT). It is compiled from the same code and the same gating; no gfx1200 card was available. - Windows. - gfx11 and CUDA hardware. `moe_fused.cu`, `moe_fused_iq.cu` and `prefill.cpp` are shared with those targets; the changes there are written to leave them as they were (see "What changed"), but nothing was run on them for this PR. - K-quant packs (`STRATA_PF_FUSED_KQ=1`): the Q4_K / Q5_K pairs pass the parity test on gfx12, but there is no end-to-end run. - The Q2_0 pack's `expert_w11_kernel` (`moe_fused.cu`) is not ported; a Q2_0 pack on gfx12 keeps MMQ. - Quality beyond first-token logits and the 4 retrieval checks (no perplexity or answer-quality benchmark). ### Related - `bench/results/2026-10-03-community-r9700-windows` found "no observable effect" of `STRATA_PF_FUSED=1` on an R9700 and no banner. That fits the old behaviour: the flag did nothing on gfx12 before this change. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。