Issues / #1180
#1180 STRATA_PF_FUSED on gfx1100: missing __syncthreads() in native_w11_kernel + waves_per_eu(8) wrong on clang 22 / ROCm 7.10
open · @zhhhn · 2 Kommentare · Auf GitHub
Server & APIAMD / HIPNVIDIA / CUDAModels & quants
Beschreibung
# `STRATA_PF_FUSED=1` on gfx1100: a missing `__syncthreads()` in `native_w11_kernel`, and `waves_per_eu(8)` computes wrong on ROCm 7.10 / clang 22
## Problem
With `STRATA_PF_FUSED=1` (`+ STRATA_PF_GEMM=1`), some prompts come back as nothing but `!` — the model confidently
predicts `!`, MTP drafts agree (2 of 2 accepted), the engine logs no error and the prefill speed looks normal.
The same prompts are correct with `STRATA_PF_FUSED_NATIVE=0` (MMQ), so the native gfx11 fused expert kernels
(`src/prefill/moe_fused_iq.cu`) are the culprit.
There are **two independent defects**; both must be fixed:
| # | Defect | Nature | Fix |
| --- | --- | --- | --- |
| A | the stage loop's `put(raw, (s + 1) & 1)` (the double-buffered weight LDS write) has **no barrier after it** | intermittent — 45–47 % of runs here | add one `__syncthreads()` |
| B | the `waves_per_eu(8)` occupancy variant (`LB=8`) computes **wrong results on this toolchain** | deterministic — 100 % of runs here | `LB=8` → `LB=1` (or gate it on the toolchain) |
They are unrelated: `LB=8` **plus** the barrier is still 0/20, while `LB=1` plus the barrier is 20/20.
## Reproduction
The repository's own test is the oracle, and it fails in **seconds**:
```sh
cmake -S . -B build-hip -G Ninja -DSTRATA_BUILD_TESTS=ON -DSTRATA_ENABLE_HIP=ON \
-DCMAKE_HIP_ARCHITECTURES=gfx1100 -DSTRATA_PREFILL_MMQ=ON \
-DCMAKE_HIP_COMPILER=$ROCM/llvm/bin/clang++ -DCMAKE_HIP_COMPILER_ROCM_ROOT=$ROCM \
-DCMAKE_HIP_PLATFORM=amd -DCMAKE_PREFIX_PATH="$ROCM;$LIBS" \
"-DCMAKE_HIP_FLAGS=--rocm-path=$ROCM --rocm-device-lib-path=$ROCM/lib/llvm/amdgcn/bitcode"
ninja -C build-hip hip_prefill_fused_iq
./build-hip/hip_prefill_fused_iq --no-timing # part 1, ~5-10 s per run
./build-hip/hip_prefill_fused_iq --no-timing --only=IQ3_XXS
```
`--no-timing` keeps it to part 1 (256 tokens, 64 experts) — that is where it fails; `--chunks=` only affects the
timed part 2.
```
prefill fused MoE (native formats) parity failed: IQ3_XXS / IQ4_NL: the fused path differs from its own arithmetic's model
fused vs the int8-rounded model (same activation / H rounding, double sums): rel RMS 6.544e-02 worst 7.952e-01
```
## Cause A: the missing barrier
In `native_w11_kernel` (src/prefill/moe_fused_iq.cu, ~line 816):
```cpp
for (int s = 0; s < NS; ++s) {
__syncthreads(); // stage s's weights are in buffer s & 1
...
if (on) { ... WMMA reading wt[s & 1] / ws[s & 1] ... }
if (s + 1 < NS) {
put(raw, (s + 1) & 1); // the other buffer: its readers passed this barrier
// <-- no barrier here
if (s + 2 < NS) load_unit<WT>(unit(s + 2), sub(s + 2), raw);
...
}
}
```
The comment says the readers of that buffer already passed the loop-top barrier, but the write still races with them.
Interleaved rounds (control and patched binaries alternated in the same session, 15 rounds each, `--no-timing`):
| variant | pass / fail |
| --- | --- |
| unpatched | 8 / **7** (47 % fail) |
| **only the barrier after the `put`** | **15 / 0** |
| only a barrier at the end of the item loop | 11 / 4 |
Re-run alone, the barrier is **20/20** (if the true failure rate were 45 %, 20/20 has probability ≈ 6e-6).
## Cause B: `waves_per_eu(8)` is wrong on clang 22 / ROCm 7.10
`native_w11_kernel<WT, GU, ALIAS=true, LB=8>` caps VGPRs via
`__attribute__((amdgpu_waves_per_eu(LB > 0 ? LB : 1)))`. The source comment justifies it with measured spill counts
("gate/up IQ2_* / IQ3_* / IQ4_XS 0-5 regs") — that measurement is **toolchain-dependent**:
| build | `--only=IQ3_XXS` | default 4 pairs |
| --- | --- | --- |
| `LB=8` (as shipped) | **0/5** (rel RMS 6.5e-02) | 0/12 |
| `LB=1` (ALIAS kept) | **5/5** | 2/5 → **20/20** with the barrier |
| `-DSTRATA_W_NO_OCC` (`LB=0`, no ALIAS) | 5/5 | 8/12 |
| `LB=8` + the barrier | **0/10** | **0/20** |
So the LDS alias (`hs` on the weight buffers) is fine — only the VGPR cap breaks. `STRATA_PF_OCC=1/4` (grid size) and
disabling ASLR do not change the failure rate, and the test's own synchronisation is complete (stream sync +
synchronous `cudaMemcpy`; the CPU model uses a fixed seed and does not read the GPU).
## Fix
```diff
--- a/src/prefill/moe_fused_iq.cu
+++ b/src/prefill/moe_fused_iq.cu
@@ -816,6 +816,7 @@
if (s + 1 < NS) {
put(raw, (s + 1) & 1); // the other buffer: its readers passed this barrier
+ __syncthreads();
if (s + 2 < NS) load_unit<WT>(unit(s + 2), sub(s + 2), raw);
if (on) {
@@ -1067,7 +1068,7 @@
-#define STRATA_NW_GU(T) do { if (occ_gu) native_w11_kernel<T, true, true, 8><<<...
+#define STRATA_NW_GU(T) do { if (occ_gu) native_w11_kernel<T, true, true, 1><<<...
@@ -1079,7 +1080,7 @@
-#define STRATA_NW_D(T) do { if (occ_d) native_w11_kernel<T, false, true, 8><<<...
+#define STRATA_NW_D(T) do { if (occ_d) native_w11_kernel<T, false, true, 1><<<...
```
For B I would not hard-code `1`: the "does not spill under 192 VGPRs" assumption is what needs to be verified per
compiler (or the variant simply disabled for AMD clang 22 / ROCm 7.10). The extra barrier costs one sync per stage
(40 per work item) and is not measurable in the run time.
## After the fix
Engine-level, with `STRATA_PF_FUSED=1 STRATA_PF_GEMM=1 STRATA_PF_SWITCH_MIN_T=4096`:
| correctness | result |
| --- | --- |
| 175,140-token prompt (2/2 garbage before) | **5 consecutive correct runs** |
| 6,850 / 25,081 / 48,978 / 74,138 / 97,488 / 146,049 tokens | all correct |
| quality (`17*23`) / tool call | `391` / `finish_reason=tool_calls` |
| prompt lend region resident in RAM / errors | 3333/3333 (100 %) / 0 |
| interleaved A/B, n=4 | 4.2K prefill | 8.8K prefill | decode |
| --- | --- | --- | --- |
| 0.1.40 baseline | 548.5 | 739.5 | 88.5 |
| **`PF_FUSED` + `PF_GEMM` (patched)** | **874.5 (+59.4 %)** | **991.9 (+34.1 %)** | 89.9 (+1.6 %) |
## Environment
- RX 7900 XTX (Navi 31, **gfx1100**, RDNA3), 24 GB VRAM, 50 GB RAM, ROCm 7.10 (TheRock wheels), **AMD clang 22**
- engine v0.1.40 built from source (`-DSTRATA_ENABLE_HIP=ON -DCMAKE_HIP_ARCHITECTURES=gfx1100 -DSTRATA_PREFILL_MMQ=ON`).
v0.1.40.1 changes no engine file (verified with the compare API and `diff -rq` on `src/` and `include/`), so this
applies to 0.1.40.1 unchanged.
- model: an IQ3_XXS native pack whose 48 layers mix `IQ2_XXS / IQ2_XS / IQ3_XXS / IQ3_S / IQ2_S` gate-up with
`Q2_0 / IQ4_NL` down; 512K context, int8 KV, `--resident-experts`, `--prefill auto:16384`Mehr auf der Site
Links zu Install, Modellen, Releases.