Pull requests / #1289

#1289 hip: the gfx1100 100401 table covers the small-T expert GEMM (prompt +31.9%)

closed · @leon-strong · 0 Kommentare · Auf GitHub

BenchmarksAMD / HIPLinux

Beschreibung

Extends the table PR #766 shipped as `7984688`, same card (RX 7900 XTX), same packaged ROCm 10.0.0, same version id 100401. That calibration covered "the 26 dense GEMM geometries of the shipped gfx1100 table", all at T=4096/8192.

A prompt reaches a different geometry, once per expert group at whatever T the group has, so it never matches a row:

```
STRATA_HIPBLASLT_VERBOSE=1, GSQ-RCO IQ2_XS, 192000 ctx, 3151-token prompt
  hipBLASLt tuning enabled (26 rows, gfx1100, version 100401)
  642 fallbacks, all f16 N=1280 K=2560 ldy=1280, over 642 distinct T values, 1..1873
```

That is the small-T expert GEMM, and it is nearly every launch in the prompt. `TuningTable::closest()` matches on `(n, k, ldy)` and takes the nearest `t_bucket`, so these rows are a spread of buckets over the small-T range rather than one row per T. `bf16 10240 2560 10240` (the hyper-connection projection, 1 fallback) is covered the same way.

**The change** — 17 rows added, 26 -> 43. No existing row is changed, including `bf16 4 10240 4 8192`, where a re-run on this box picked a different solution id (1036 against the shipped 1035); the shipped value is kept and the shape is not touched. Calibrated with `tune_hipblaslt` built from this tree against the engine's own 32 MiB workspace (`gemm.cu:441`). All 34 candidates emitted, i.e. none failed the tuner's accuracy gate (rel_l2 <= 1e-4, max_abs <= 1e-2); the 17 that are in are the two geometries the verbose run showed falling back.

**Measured**, same binary and the same config with the table as the only difference, 3 interleaved pairs, 3151-token prompt:

| pair | 26 rows | 43 rows | |
|---|---:|---:|---:|
| 1 | 691.6 tok/s | 940.7 tok/s | +36.0% |
| 2 | 712.3 tok/s | 943.2 tok/s | +32.4% |
| 3 | 727.7 tok/s | 928.7 tok/s | +27.6% |
| mean | 710.5 tok/s | 937.5 tok/s | **+31.9%** |

Decode unchanged (58.6 -> 60.6 tok/s, inside this box's 2.3% run drift). `STRATA_HIPBLASLT_VERBOSE=1` reports `fallbacks=0` with the 43 rows against 642 with the 26.

`hip_prefill_hipblaslt_gemm` passes (4/4 cases, exit 0). Note it smoke-tests only 2 of the table's rows (`N=48`, `N=512`), so it exercises the load and resolve path, not these rows' arithmetic; the per-solution accuracy above is the tuner's own gate.

Machine: Ryzen 7 9800X3D, RX 7900 XTX, WSL2.

Mehr auf der Site

Links zu Install, Modellen, Releases.