Pull requests / #312

#312 QSA select: ROCWMMA block-scores arm (opt-in STRATA_SELECT_WMMA=1, gfx1100)

closed · @xyzzing · 0 评论 · 在 GitHub 查看

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quants

描述

Rework of #228 per your note — this PR is **only the scorer, its parity test, the bench and the CMake detection**. No `.agentic/`, `AGENTS.md`, `CLAUDE.md`, CI workflows, `serve/telemetry.py`, `direct_file.*`, or `math_constants.h`. (The rates-table row and the verify-profiler fix that were bundled in #228 are deliberately left out.)

## What it is
The CUDA TF32 scorer can't run on gfx1100 (no `mma.sync`). This adds `qsa_block_scores_wmma`: a 16×16×16 ROCWMMA GEMM view of the block scores (FP16 operands, FP32 accumulation), reusing `block_scores_tail_kernel` unchanged so the tail `n_bid` stays bitwise the warp kernel's. Keys/queries convert to FP16 per call into an internal scratch.

## Opt-in, default off (as requested)
- `STRATA_SELECT_WMMA=1` → try the ROCWMMA arm, else fall through
- `STRATA_SELECT_OLD=1` → force the warp kernel
- default (unset) → the existing tensor-core (3xTF32) arm, then warp — **unchanged behavior**

The prompt path prints the arm that actually ran (`strata select: prompt scorer arm = ...`), so a quiet fallback can't masquerade as the fast path. Without rocWMMA headers the arm compiles out and the entry refuses; a CUDA build is untouched.

## Commits
1. **The scorer + CMake detection + prompt-path dispatch** (`qsa_select.cu`/`.hpp`, `prefill.cpp`, `CMakeLists.txt`).
2. **The WMMA parity harness** — two independent references against the pre-derived fp16 bound.
3. **The bench + adversarial parity fixtures** (`--selftest`).

## gfx1100 numbers (RX 7900 XTX, IQ3_S — my card carries these, per your note)
- **Kernel**: 10.5× vs the warp scorer at 128K (`nq=256`), median-of-9.
- **Parity**: 5/5 fixtures within the pre-derived bound `B(qi,b) = 2^-9.5·S + 8·2^-24·|score|` (PORTING.md §18) over one-hot + coordinate-coded rows.
- **End-to-end** (paired, greedy, n=10 each): 1K **+0.18%**, 4K **+0.93%**, 32K **−4.54%**, 64K **+3.77%**, 128K **+8.7%** — consistent with your read (+8.7% at 128K, −4.5..+3.8% elsewhere), and exactly why the arm stays opt-in rather than default.

Because the gain is context-dependent (clear at 128K, mixed below), the arm is opt-in and the default path is byte-for-byte unchanged. The parity harness and bench are included so you can rebuild and re-measure on any gfx1100. Numbers are from my campaign build (0.1.29-based engine of record); the scorer is additive and the prompt-path hook is a localized call-site edit that merges cleanly onto 0.1.30.

Built on the multi-query scorer (#187). Refs: PORTING.md §18 (parity bound).

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。