Pull requests / #1149
#1149 HIP gfx103x FP16 prompt (#835 follow-up): STRATA_F16_RANGE measures what reaches FP16's range, and hip_prefill_gemm covers the FP16-io route
open · @xjc10 · 0 Kommentare · Auf GitHub
BenchmarksMulti-GPUAMD / HIPDocumentation
Beschreibung
Two follow-ups to #835 that were on its branch when `main` was rewritten, re-filed on 0.1.40.1 with the review changes (opt-in `STRATA_HIP_PROMPT_F16=1`, the mode decided once per stage) kept. - `STRATA_F16_RANGE=1` records, per device, the largest |value| and the count beyond 65504 (or not finite) at the three places a value enters or leaves FP16 - the activation images (`hf_sat`), the BF16 weights converted to FP16 (`bf16_to_f16_rows`) and the FP16 GEMM outputs as they are widened (`widen_rows_f16`) - and prints one line per prompt. It sits beside the review's `STRATA_DBG_NAN` count in `widen_rows_f16`. Off, the cost is one flag read per value; on, a 34.4K-token prompt read at 889.8 tok/s against 888.5 without it. Measured on one RX 6900 XT over 466K tokens in nine prompts (three corpus slices, two prompts built to push activations - delimiter floods, repeated tokens, base64, mixed scripts - and four batches of real API requests): FP16 GEMM outputs peak at 410 of 65504, nothing beyond the range or non-finite. The numbers are in `bench/results/2026-10-04-rdna2-fp16-prompt/README.md`, which also gains the same-card comparison against the FP32 SGEMM route of #1006. - `hip_prefill_gemm` runs its shapes a second time with `set_f16_io(true)` on any HIP card: X as the FP16 image for the BF16 products, a scratch of 64 rows so a 96-row weight converts in two slices, `ldy > N`, an odd N, T = 1 and a `beta = 1` f16() call (which keeps the FP32-out GEMM); the FP16-out results are checked at FP16's rounding, the row padding and the guards around Y must be untouched. Its BF16 inputs are no longer all zero. HIP ctest on the RX 6900 XT at 0.1.40.1 (`82f46a8`), with this change and #1148 (the `hip_prefill_wmma_gemm_parity` skip) applied: 19 tests, 13 pass, 6 skip (the WMMA / hipBLASLt / fused-kernel tests that do not apply to gfx1030), 0 fail. Original branch: #835's `hip-gfx103x-fp16-prompt` after the merge. Developed with an AI coding assistant; every number above was measured on 2x RX 6900 XT (gfx1030, PCIe 4.0 x8 each) / Ryzen 5 5600X, ROCm 10.0.
Mehr auf der Site
Links zu Install, Modellen, Releases.