Pull requests / #1368
#1368 prefill: quantize each token's MoE input once and scatter it to its k rows (the same bytes)
open · @sergqwer · 0 コメント · GitHub で見る
BenchmarksAMD / HIPNVIDIA / CUDAModels & quants
本文
The prompt path quantizes a layer's MoE input to q8_1 in expert order: `mmq::quantize(m.mixed, m.src_dev, ...)` runs over T x K rows (K = 10), so every token's row of `mixed` is read and quantized ten times, once per expert. llama.cpp's `quantize_scatter_mmq_q8_1_cuda` (in the pinned `quantize.cu`, which strata_mmq already compiles) quantizes each token once and writes the block to all of the token's rows through the inverse map. Prefill already builds that map as `slot_dev`: token t's k-th row is `slot[t * K + k]`, the inverse of `src_dev`. This adds `mmq::quantize_scatter(x, slot, src, ...)` and uses it at that call site. The q8_1 bytes are the same, so it is the default. `STRATA_QUANT_GATHER=1` keeps the per-row gather. The peer and down-projection quantizations are unchanged. It adds no device code, so the CUDA and HIP builds compile the same kernels as before; the SYCL port has its own `prefill.cpp` and is untouched. ### Identity - **`prefill_mmq_scatter_test` (new ctest).** It compares scatter against the gather on the whole output buffer, for every q8_1 layout MMQ uses: D4 (Q2_0, Q8_0, IQ2_XS), DS4 (Q4_1) and D2S6 (Q2_K). T = 1, 2, 7, 600 and 2099, routing built as Prefill builds it (10 of 512 experts), one all-zero token. All 25 cases: 0 differing bytes. - **End to end against main `e8ca9afd`.** Tested with IQ2_XS at 600 tokens, 2K and 32K, and Q2_0 at 2K, all with a fixed cache. The first-token logits are identical bytes and the 32 greedy tokens are the same. - `prefill_fused_moe_test`, `prefill_fused_iq_test` and `quantize_act_parity` pass. ### Measured RTX 5090 (96 MB L2) + Ryzen 9 9950X3D. **The quantizer alone** (`prefill_mmq_scatter_test --bench`, median of 20): | tokens (x 10 rows) | IQ2_XS gather / scatter | Q2_0 gather / scatter | |---|---|---| | 2,048 | 0.053 / 0.027 ms | 0.054 / 0.027 ms | | 8,192 | 0.218 / 0.223 ms | 0.219 / 0.224 ms | | 32,768 | 1.690 / 1.207 ms | 1.746 / 0.912 ms | Both versions write the same T x K q8_1 rows. The gather also reads each input row K times. When the chunk's input fits in L2, as with an 8K chunk (84 MB) on this card, those re-reads are cheap and the two are even. When it does not, the gather goes back to DRAM for them. **In the engine** (IQ2_XS, a 32K prompt, nsys, the quantizer kernels' total): | | quantization | |---|---| | default chunk (8192) | 60.1 ms -> 59.4 ms (even) | | `--prefill auto:32768` | 77.5 ms -> 58.1 ms (-25%) | That is ~0.4% of a 4.6 s prompt. The whole-prompt times (3 interleaved pairs at 600, 2K and 32K) were within the run-to-run noise. On this card the gain only shows with large chunks. On cards with less L2 even an 8K chunk's input does not fit: a 5070 has 48 MB, a 3090 6 MB. I could not measure those cards here. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
関連リンク
インストール・モデル・リリースへの站内リンク。