Pull requests / #1298
#1298 hip: the gfx110x fp16 (rocWMMA) qsa-select scorer arm (STRATA_SELECT_WMMA)
open · @xyzzing · 0 Kommentare · Auf GitHub
Beschreibung
Per the offer in the #944 thread: the gfx110x **fp16 (rocWMMA)** qsa-select scorer arm, as the gfx110x branch of the existing `STRATA_SELECT_WMMA` opt-in, alongside the bf16 hand-rolled kernel that 0.1.40 ships. **What it is.** One commit: the fp16 rocWMMA scorer kernel for gfx110x (`rocwmma.hpp`, 16x16x16, fp16 operands / fp32 accumulate) behind the same `STRATA_SELECT_WMMA` env, plus its parity harness `src/kernels/qsa_select_wmma_parity.cpp` (fp64 oracle + the bitwise warp host model, pre-declared bounds, negative controls). The bf16 kernel (gfx12 + the other gfx11 layouts) and the warp kernel stay the fallbacks everywhere else — the default path (env absent) is byte-identical to main; on cards without the rocWMMA headers the arm compiles out and the host entry refuses. **Why fp16 here.** Paired A/B on gfx1100 (2026-10-06, same engine build, arms differing only in the select kernel): the fp16 arm's qsa-select phase ran **1.64x faster** than the shipped bf16 arm at 128K (measured on the fork's 0.1.40 refresh line; the fork carries this exact kernel in production). The re-land on this base measured parity-neutral overall vs our previous production reads (x1.0009 on the canonical stream tail). **Gates on this base.** Full HIP build green; `qsa_select_wmma_parity` all cases PASS (worst fp64-oracle bound utilisation 11.0%, one-hot/coordinate-coded adversarial case bitwise vs its host model, 0 differing selection ids, negative control fails as designed). *Prepared with an AI engineering agent (GLM-5.3 / Z.ai) under human direction; every number is from our own recorded runs.*
Mehr auf der Site
Links zu Install, Modellen, Releases.