Pull requests / #1067

#1067 hip/gfx906: `__dp4a` falls back to a 4-iteration loop although gfx906 has v_dot4_i32_i8

closed · @xxDoman · 0 评论 · 在 GitHub 查看

BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants

描述

## What

`dp4a` in `include/strata/hip_compat/intrinsics.hpp` has a native signed-dot4 path for RDNA3/RDNA4 (`__builtin_amdgcn_sudot4`) and for RDNA2 gfx103x (`__builtin_amdgcn_sdot4`), and a portable 4-iteration loop for every other target. The gfx906 build (`STRATA_HIP_GFX906`) is one of the "other targets", so every `__dp4a` in the decode path is 4 iterations instead of one instruction.

gfx906 (CDNA1 / GCN5, Instinct MI50/MI60, Radeon VII) has `v_dot4_i32_i8`, the same instruction the gfx103x branch already uses:

```
$ llvm-mc -arch=amdgcn -mcpu=gfx906 -show-encoding <<< 'v_dot4_i32_i8 v1, v2, v3, v4'
v_dot4_i32_i8 v1, v2, v3, v4        // dot1-insts
```

## Fix

One line: add `__gfx906__` to the gfx103x condition. Same signed x signed byte products, accumulated modulo 2^32, no clamp — exactly the CUDA `__dp4a` contract, and bit-identical to what the gfx103x path produces (no new instruction semantics are introduced).

## Note on ordering

This is independent of the build fix in #1066: on the current tree `fused_gr.cu` does not even compile, but this change is orthogonal to it. Either order works.

## Measurement (2x gfx906, not a repo benchmark)

Instinct MI50 32 GB master + Radeon VII 16 GB helper, IQ3_XXS, ctx 32000, 11552-token prompt, 128 decode, MI50 at 200 W, 4 runs:

| tree | prefill | decode |
|---|---|---|
| 0.1.39 + the same dot4 line | 379-380 | 38.6-38.9 |
| **0.1.40 + this line** | 291-299 | **39.5-40.7** |

The decode difference is the point (~+3-4%); the value itself is one machine, so I would not quote it as a benchmark. An exactness check against the portable path (the tree's `tests` dp4a parity, or a small local one) is the right gate — happy to run and post it here if you want it before merging.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。