Pull requests / #1084
#1084 hip/gfx906: `__dp4a` falls back to a 4-iteration loop although gfx906 has v_dot4_i32_i8
open · @xxDoman · 0 评论 · 在 GitHub 查看
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
描述
Re-creation of #1067 on the new main (the history rewrite closed it; the file is byte-identical to v0.1.40, so I cherry-picked my commit onto `main`). ## What `dp4a` in `include/strata/hip_compat/intrinsics.hpp` has a native signed-dot4 path for RDNA3/RDNA4 (`__builtin_amdgcn_sudot4`) and gfx103x (`__builtin_amdgcn_sdot4`), and a portable 4-iteration loop for every other target. The gfx906 build (`STRATA_HIP_GFX906`) is one of those, so every `__dp4a` in the decode path is 4 iterations instead of one instruction. gfx906 (CDNA1 / GCN5, Instinct MI50/MI60, Radeon VII) has `v_dot4_i32_i8`, the same instruction the gfx103x branch uses: ``` $ llvm-mc -arch=amdgcn -mcpu=gfx906 -show-encoding <<< 'v_dot4_i32_i8 v1, v2, v3, v4' v_dot4_i32_i8 v1, v2, v3, v4 // dot1-insts ``` ## Fix One line: add `__gfx906__` to the gfx103x condition. Same signed x signed byte products, accumulated modulo 2^32, no clamp — exactly the CUDA `__dp4a` contract, bit-identical to what the gfx103x path produces. ## Ordering Independent of the build fix (#1083): on the current tree `fused_gr.cu` does not even compile, but this change is orthogonal. Either order works. ## Measurement (2x gfx906, not a repo benchmark) Instinct MI50 32 GB master + Radeon VII 16 GB helper, IQ3_XXS, ctx 32000, 11552-token prompt, 128 decode, MI50 at 200 W, 4 runs: | tree | prefill | decode | |---|---|---| | without this line | ~294 | 38.6-38.9 | | **with this line** | ~294-381 | **39.3-40.2** | The decode difference is the point (~+2-4%); the value itself is one machine, so I would not quote it as a benchmark. An exactness check against the portable path is the right gate — happy to run and post it here if you want it before merging.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。