Pull requests / #1418

#1418 cuda: opt-in MMVQ activation reuse across output rows

open · @imanu86 · 0 comentários · No GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAWindows

Descrição

This adds an opt-in exact multi-column MMVQ kernel for Q4_K, Q5_K, Q6_K and IQ4_XS, plus `b6_mmvq_bench`. Each column's Q8_1 activation fragment is loaded once per weight block and reused across multiple output rows. The Q6_K scales remain in registers. The intended arithmetic order matches the existing exact layout.

`STRATA_B6_MMVQ_ROWS=1` selects two output rows per block; `2`, `4`, and `8` select those row counts explicitly. Unset or `0` retains the existing kernel. The new path requires the exact layout, `STRATA_TSUM` off, and a supported format; other calls fall back. The setter/getter are for tools and must be used before graph capture. Single-column calls retain their existing path.

Relationship to current upstream: this branch starts at `82f46a8`. Upstream `d5ea713` now also contains the separate interleaved MMVQ path (`MMVQ_IL`), whose selection is restricted to supported shapes, 2-4 columns, and sm_80 or newer. This PR has not been rebased or measured together with that addition. Its potential relevance includes Turing and calls that use the existing native layout; it should not be described as a measured improvement over the new interleaved path.

Historical evidence, with limits:

- Local fork microbenchmark logs from October 6-7 cover RTX 2080 Ti sm_75 and RTX 3060 sm_86, 15 model-shaped cases, widths 1, 2, 3, 4, 5 and 8. They end with zero failed comparisons for the tested modes. These synthetic checks compare the new and existing exact kernels, and both against individual single-column calls. They are historical fork results, not a fresh test of this PR head.
- In the historical two-row run, for example, GDN qkv Q6_K at four columns was 0.0848 to 0.0654 ms on the 2080 Ti and 0.1372 to 0.1044 ms on the 3060. CUDA-event microbenchmarks do not establish an end-to-end speedup.
- The fork's deterministic full-model runs reported a modest improvement, but forced IDs/windows do not control all cache and speculative work. No causal end-to-end percentage, general quality equivalence, or improvement on latest upstream is claimed here.
- Some wider cases regress; this does not change the default. HIP and other GPU architectures have not been runtime-qualified.

Publication validation:

- Head `8cae814e6e819736e47c95f3b5e8b056c7528c0f`, one commit over `82f46a8c8f475f001ad76d92f58f4a4f8ffb0253`; `git diff --check 82f46a8..HEAD` passes.
- No fresh build or GPU test of this exact branch was performed for publication. Compilation of this exact branch, synthetic parity on both cards, and an integration comparison against current upstream remain pending.

Reproduction after building the CUDA branch:

```text
cmake --build build --target b6_mmvq_bench
build/b6_mmvq_bench 500 -1 1
build/b6_mmvq_bench 500 0 4
```

The first benchmark selects all visible devices and two rows per block; the second selects device 0 and four rows per block. An exit code of zero covers only the benchmark's synthetic case matrix, not full-model equivalence.

No site

Links install, modelos, releases.