Pull requests / #1125
#1125 cpu experts: an AVX2 Q8_K activation quantizer (byte-identical to ggml's)
open · @Hardin22 · 0 コメント · GitHub で見る
BenchmarksNVIDIA / CUDAModels & quantsWindowsLinux
本文
This replaces #851, which GitHub closed when `main` was force-pushed. Same commit, cherry-picked onto the new `main` (0.1.40.1); the only conflicts were next to 0.1.40's new `iq256_*_v` declarations and the `iq_avx2_parity` test entry, and both are kept. The i-quant expert rows the CPU computes take their activations as Q8_K: - the token's hidden state for gate/up (`native_quant_act`, once per token and layer, on the host before the pool starts); - each expert's FFN output for down (`native_quant_h`, in the pool). On x86, ggml-cpu's `quantize_row_q8_K` is the scalar reference (`arch/x86/quants.c` calls `quantize_row_q8_K_ref`). This adds `q8k_quant_avx2`, which does the same arithmetic in AVX2 and writes the same bytes: - the same max, where the first value of the largest magnitude keeps its sign; - the same iscale; - per value, one multiply and the same 1.5*2^23 rounding add; - the same clamp and int8 store, and the same bsums; - zero blocks left as the reference leaves them. - **Gating**: only behind `cpu_avx2_ok()`, so an AVX-only CPU or `STRATA_FORCE_ISA=avx` keeps ggml's. `STRATA_NO_Q8K_AVX2=1` keeps ggml's on any CPU. - **Test**: `q8k_quant_parity` (CTest, CPU only) compares three things against `quantize_row_q8_K_ref` with `memcmp` over whole rows: - the AVX2 copy; - ggml-cpu's `from_float`; - `native_quant_act` / `native_quant_h` with an IQ3_S expert's format. It covers 256, 512, 2560 and 6144 values, and random, scaled, sparse, zero, tied-maximum, extreme (±3e38), denormal, grid-tie, NaN and inf rows. A second entry runs it again with `STRATA_FORCE_ISA=avx`. `--bench` times both quantizers. ## Measured (i9-14900KF, Windows, MSVC) - **One row of 2560 values**: 2141 ns → 891 ns (2.4x). - **End to end** (measured on 0.1.39): Swift 1.5 IQ3_XXS at 160K, greedy, 4 prompts × 3 rounds, decode tok/s. - **RTX 4060 Ti + RTX 5080 layer split** (with the resident RAM mode from #848), two interleaved A/B pairs: 66.0 / 67.6 → 68.5 / 69.0 (**+2.9%**). The host's `actq` per window drops from 0.35 to 0.15 ms of a 36.5 ms window. Code and Italian prose were faster in both pairs, English prose in one pair (equal in the other). - **RTX 5080 alone**: `actq` 0.50 → 0.20 ms, but the window is 70-150 ms there, so the speed change is within the noise. So it matters most where the window is short and the host is on its critical path (layer splits, mostly-resident models), and costs nothing elsewhere. ## Notes - Built and tested with MSVC only. On GCC the multiply's result goes through an empty `asm` so it isn't contracted into an FMA (GCC's default `-ffp-contract=fast`; the reference rounds the product first). The parity test is what would catch it on Linux. - Independent of the layer split: it applies on main by itself.
関連リンク
インストール・モデル・リリースへの站内リンク。