Pull requests / #1318

#1318 docs(AMD_HIP): my R9700 numbers for hipBLASLt 1.4.1 and STRATA_HIP_WMMA

closed · @l33tm4st3r · 0 comentários · No GitHub

BenchmarksSetup & installAMD / HIPModels & quantsDocumentationLinux

Descrição

I set up Strata on my R9700 and measured the two AMD switches setup does not decide on its own. Only docs/AMD_HIP.md changes.

**My machine:** Radeon AI PRO R9700 32 GB (gfx1201, 1002:7551) on a Ryzen 9 9950X3D, 64 GB RAM (62 GiB usable), CachyOS, kernel 7.3.0-rc6-1-cachyos-rc with the kernel's amdgpu driver and linux-firmware-amdgpu 20260916, system ROCm from the rocm-gfx120x-bin 10.1.0 package (hipBLASLt 1.4.1, HIP 7.16.26385, clang 24.0.0), Python 3.14.7 in setup's own venv, engine 0.1.40 compiled here for gfx1201, model files on a 4 TB NVMe (btrfs).

**My settings:** Qwen3.8-Flash-Next IQ3_S with `--max-context 431072` - I run this model over 400K tokens of context - with `--kv int8`, `--spec 4 --spec-min-p 0.5` on the MTP draft, `--vram-reserve-mib 700` and `--expert-cache auto`, which lands on 10244 expert slots (19.44 GiB of VRAM) and 8192-token prompt chunks. KV streaming is off (IQ3_S needs about 62 GB of RAM and I have 62), so the KV cache stays in VRAM.

**What I measured, and what I did not:** every prompt number below is a fresh prompt of 4,210 or 8,830 tokens, three matched trials, GPU idle, with the first fresh prompt after each start left out because it is cold (949 -> 1,423 and 1,538 -> 1,762 there). 8,830 tokens is the longest prompt I measured. 431072 is the context the model runs at on this PC, not the length of those prompts. The hipBLASLt A/B ran while the config still said 131072 ctx; the WMMA A/B ran at 431072 ctx. I have no prompt-length number at 100K or more from this machine; the 16K and 262K numbers in docs/AMD_HIP.md are the maintainers'.

**hipBLASLt 1.4.1:** I calibrated `gfx1201-hipblaslt-100401.txt` with `tune_hipblaslt` over the 32 geometries of `gfx1201-hipblaslt-100500.txt`. It gives 1.15x geometric mean per GEMM over hipBLAS (1.55-1.68x at N=512/640 K=2560, 0.99-1.13x at N=10240/12288 K=2560) but only +1% end to end: medians 1,606 -> 1,613 tok/s over the three matched trials (a 1.008x geometric mean), decode 87.1 -> 85.1 tok/s. That matches the 0-3% the R9700 paragraph in docs/AMD_HIP.md already reports for 1.4.1, so I left the table out of the repo and recorded only the numbers.

I dropped two rows from that table, bf16 N=10240 K=320 ldy=10240 at T=4096 and T=8192: hipBLAS won T=8192 in five separate runs (0.609-0.619 ms against 0.648-0.677 ms for the best of 16 candidates, four different winning solution ids), and dropping only the T=8192 row would send that call to the T=4096 row rather than back to hipBLAS, because `closest()` takes the nearest T bucket.

**STRATA_HIP_WMMA=1** is what pays on this card: prompt medians 1,571 -> 2,064 tok/s, a 1.27x geometric mean over the three matched trials: +31.4% and +31.7% at 8,830 tokens (1,571 -> 2,064 and 1,568 -> 2,065) and +17.1% at 4,210 (1,699 -> 1,990), decode unchanged (89.8 -> 86.9 tok/s). `hip_prompt_attn_wmma` passes here. `STRATA_SELECT_WMMA=1` on top changed nothing at 4,210 and 8,830 tokens (2,026 tok/s).

Two calibration traps I ran into are written down too: the first case of a fresh `tune_hipblaslt` process times its hipBLAS baseline while the clocks are still ramping (the same shape read 0.31-0.63 ms across five runs while the case right after it held 0.606-0.619 ms), and the first fresh prompt after a model load is not a measurement (949 vs 1,423 tok/s in the two arms of my A/B, while the three trials after it differed by 0.4-1.5%).

No site

Links install, modelos, releases.