Pull requests / #1067
#1067 hip/gfx906: `__dp4a` falls back to a 4-iteration loop although gfx906 has v_dot4_i32_i8
closed · @xxDoman · 0 commentaires · Sur GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
Description
## What `dp4a` in `include/strata/hip_compat/intrinsics.hpp` has a native signed-dot4 path for RDNA3/RDNA4 (`__builtin_amdgcn_sudot4`) and for RDNA2 gfx103x (`__builtin_amdgcn_sdot4`), and a portable 4-iteration loop for every other target. The gfx906 build (`STRATA_HIP_GFX906`) is one of the "other targets", so every `__dp4a` in the decode path is 4 iterations instead of one instruction. gfx906 (CDNA1 / GCN5, Instinct MI50/MI60, Radeon VII) has `v_dot4_i32_i8`, the same instruction the gfx103x branch already uses: ``` $ llvm-mc -arch=amdgcn -mcpu=gfx906 -show-encoding <<< 'v_dot4_i32_i8 v1, v2, v3, v4' v_dot4_i32_i8 v1, v2, v3, v4 // dot1-insts ``` ## Fix One line: add `__gfx906__` to the gfx103x condition. Same signed x signed byte products, accumulated modulo 2^32, no clamp — exactly the CUDA `__dp4a` contract, and bit-identical to what the gfx103x path produces (no new instruction semantics are introduced). ## Note on ordering This is independent of the build fix in #1066: on the current tree `fused_gr.cu` does not even compile, but this change is orthogonal to it. Either order works. ## Measurement (2x gfx906, not a repo benchmark) Instinct MI50 32 GB master + Radeon VII 16 GB helper, IQ3_XXS, ctx 32000, 11552-token prompt, 128 decode, MI50 at 200 W, 4 runs: | tree | prefill | decode | |---|---|---| | 0.1.39 + the same dot4 line | 379-380 | 38.6-38.9 | | **0.1.40 + this line** | 291-299 | **39.5-40.7** | The decode difference is the point (~+3-4%); the value itself is one machine, so I would not quote it as a benchmark. An exactness check against the portable path (the tree's `tests` dp4a parity, or a small local one) is the right gate — happy to run and post it here if you want it before merging.
Sur le site
Liens install, modèles, releases.