Issues / #979
#979 HIP: a gfx1201 hipBLASLt table for 1.2.1 (100201, ROCm 7.2.0) makes the prompt 6-17% slower than no table
closed · @aswin-dot-R · 1 comentarios · En GitHub
BenchmarksSetup & installAMD / HIPModels & quantsLinux
Descripción
### Summary A hipBLASLt table calibrated for **gfx1201 + hipBLASLt 1.2.1 (`100201`, the packaged ROCm 7.2.0)** with `tools/hip/tune_hipblaslt` loads cleanly (`fallbacks=0`) but makes the prompt **slower** than running without a table. Reporting it so nobody ships one for this combination expecting setup's "+40-60% prompt speed" (measured on the 7900 XTX); on RDNA4 + 1.2.1 the untuned path seems to be the better default already. ### Setup - AI PRO R9700 32 GB + RX 9070 XT 16 GB (both gfx1201), layer split `[1, 0]`, Qwen3.8-Flash-Next IQ3_S, `--max-context 262144 --kv int8`, MTP drafts - engine v0.1.39 (setup's flags, `STRATA_SH_STREAM=0` per #816), Linux 6.17, ROCm 7.2.0 (hipBLASLt 1.2.1) - table: the 32 cases of the shipped `gfx1201-hipblaslt-100500.txt`, calibrated on the R9700 as in AMD_HIP.md's "Tuning table" -> `STRATA_HIPBLASLT_TUNING_V1 gfx1201 100201`, 32 rows ### Result (fresh server per run, engine timings, `STRATA_HIPBLASLT_VERBOSE=1`) | | no table | with the 100201 table | | --- | ---: | ---: | | prompt, 40,920 tokens | **1,186 tok/s** | 980 tok/s (-17%) | | prompt, 199,312 fresh tokens of a 232K prompt | **1,218 tok/s** | 1,148 tok/s (-6%) | | decode, 1200 tokens | 79.4 | 79.6 | | `prefill gemm: hipBLASLt summary` | - | `launches=14670 fallbacks=0 unique_fallback_shapes=0` | | 232K needle | correct | correct | One run per variant; the 41K difference is far outside the run-to-run spread I see on this machine (~2-3%). I can share the table file or rerun with other `--tokens` if a different calibration is worth trying. ### Suggestion Maybe note in AMD_HIP.md's "Tuning table" section that on gfx1201 with hipBLASLt 1.2.1 a table measured slower than no table, so "compare before keeping it" is not a formality there.
En el sitio
Enlaces a install, modelos, releases.