Pull requests / #1167

#1167 hip: a rocBLAS solution table for the FP16 prompt GEMMs on gfx103x (the GDN projection 7x below 1,152 tokens; short prompts -10 to -14%)

open · @xjc10 · 0 comentarios · En GitHub

BenchmarksMulti-GPUAMD / HIPModels & quants

Descripción

Re-filed from #981 (closed with the history rewrite) as the one commit on top of #835, which is in 0.1.40.1. The table serves the `STRATA_HIP_PROMPT_F16=1` route; with it off nothing runs through it. In `f16_inplace` the table's kernel is tried first, `cublasGemmEx` otherwise, and the review's `STRATA_DBG_NAN` count follows either. The tool (`tools/hip/tune_rocblas`), the table format and the engine route are described in #981; the table file is the one measured there (`tools/hip/gfx1030-rocblas-5.6.0.8d1ae90e.txt`).



**Measured on 0.1.40.1** (`82f46a8` + this commit; RX 6900 XT, one card, ROCm 10.0.0 / rocBLAS 5.6.0.8d1ae90e; `STRATA_HIP_PROMPT_F16=1 STRATA_SH_STREAM=1`, `--kv int8 --kv-resident 32768 --max-context 131072 --spec 4`; the #884 GFXOFF workaround in place, which does not change speed):

- `hip_prefill_rocblas_gemm` (ctest, `STRATA_ROCBLAS_TUNING=tools/hip`): PASS - `rocBLAS tuning enabled (112 rows, 76 with a solution, gfx1030, rocBLAS 5.6.0.8d1ae90e)`, 4 cases within FP16 rounding of hipBLASEx, `tuned_launches=4 fallbacks=0`. On 0.1.40.1 the test needs `STRATA_HIP_PROMPT_F16=1` as well, since the FP16 route is opt-in now; without it the test skips.
- A/B: table on and off alternated, four cold server starts each (eight in all), each start reading three different text slices per length (no prompt-cache hits), `max_tokens` 8, temperature 0: 12 pairs per length, the same slice in each pair. Prompt read speed from the engine's own line:

| prompt tokens | pairs | table off, median (range) | table on, median (range) | gain per pair, median (min..max) |
| --- | ---: | ---: | ---: | ---: |
| 435-540 | 12 | 60 tok/s (56-65) | 60 (56-65) | -0.2% (-0.5..+0.2) - below `--short-read 640` the prompt is read token by token, no prompt GEMMs |
| 833-1,084 | 12 | 261 (233-281) | **292 (264-321)** | **+13.3%** (-1.9..+15.4; the -1.9 is the 1,084-token slice, next to the 1,152 bucket where rocBLAS's own choice is the fast kernel) |
| 3,065-4,376 | 12 | 731 (670-752) | **813 (743-839)** | **+11.1%** (+8.8..+12.7) |
| 6,122-8,332 | 12 | 901 (770-928) | 927 (797-956) | +3.0% (+2.6..+3.8) |

Decode is not touched (the table serves the prompt GEMMs only). The engine logs `rocBLAS tuning enabled (...)` when the table is read, so an arm's state is visible in its log.

Developed with an AI coding assistant; every number above was measured on 2x RX 6900 XT (gfx1030, PCIe 4.0 x8 each) / Ryzen 5 5600X, ROCm 10.0.

En el sitio

Enlaces a install, modelos, releases.