Pull requests / #1084

#1084 hip/gfx906: `__dp4a` falls back to a 4-iteration loop although gfx906 has v_dot4_i32_i8

open · @xxDoman · 0 commentaires · Sur GitHub

BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants

Description

Re-creation of #1067 on the new main (the history rewrite closed it; the file is byte-identical to v0.1.40, so I cherry-picked my commit onto `main`).

## What

`dp4a` in `include/strata/hip_compat/intrinsics.hpp` has a native signed-dot4 path for RDNA3/RDNA4 (`__builtin_amdgcn_sudot4`) and gfx103x (`__builtin_amdgcn_sdot4`), and a portable 4-iteration loop for every other target. The gfx906 build (`STRATA_HIP_GFX906`) is one of those, so every `__dp4a` in the decode path is 4 iterations instead of one instruction.

gfx906 (CDNA1 / GCN5, Instinct MI50/MI60, Radeon VII) has `v_dot4_i32_i8`, the same instruction the gfx103x branch uses:

```
$ llvm-mc -arch=amdgcn -mcpu=gfx906 -show-encoding <<< 'v_dot4_i32_i8 v1, v2, v3, v4'
v_dot4_i32_i8 v1, v2, v3, v4        // dot1-insts
```

## Fix

One line: add `__gfx906__` to the gfx103x condition. Same signed x signed byte products, accumulated modulo 2^32, no clamp — exactly the CUDA `__dp4a` contract, bit-identical to what the gfx103x path produces.

## Ordering

Independent of the build fix (#1083): on the current tree `fused_gr.cu` does not even compile, but this change is orthogonal. Either order works.

## Measurement (2x gfx906, not a repo benchmark)

Instinct MI50 32 GB master + Radeon VII 16 GB helper, IQ3_XXS, ctx 32000, 11552-token prompt, 128 decode, MI50 at 200 W, 4 runs:

| tree | prefill | decode |
|---|---|---|
| without this line | ~294 | 38.6-38.9 |
| **with this line** | ~294-381 | **39.3-40.2** |

The decode difference is the point (~+2-4%); the value itself is one machine, so I would not quote it as a benchmark. An exactness check against the portable path is the right gate — happy to run and post it here if you want it before merging.

Sur le site

Liens install, modèles, releases.