Pull requests / #415
#415 IQ4_XS expert rows on the AVX-2 path: the multi-token kernel this format never had
closed · @pipeob0 · 0 comments · View on GitHub
NVIDIA / CUDAModels & quantsWindows
Description
## What IQ4_XS (GGML type 23) was missing from `iq256_supported`, so on an **AVX-2-only CPU** its gate/up expert rows fell back to ggml's per-token `vec_dot` while every other i-quant format got the multi-token kernel that decodes the weights once per verify window. This adds `Fmt32<23>` and widens the two gates that were keeping the format out. The machine is a Ryzen 7 5700X3D (Zen 3, AVX2, no AVX-512), Windows, engine 0.1.32 built from source. It is the CPU class the AVX-2 path exists for, and IQ4_XS is one of the formats people actually run on it — including the low-RAM path 0.1.31 added, which makes a 61 GiB-expert model reachable on a 64 GB PC. ## How the kernel works `Fmt32<23>` reuses the 16-value codebook `iq4nl_rows` already has. Each of the eight 32-value sub-blocks carries its own 6-bit scale — two nibbles of `scales_l[p]` plus two bits of `scales_h` — used **signed** as `(ls - 32)`, so the scale folds into the `int16` operand of `madd_epi16` instead of becoming a float multiply per sub-block. The codebook sign is carried the way `iq4nl_rows` does it: `|w|` as the unsigned `maddubs` operand, `w`'s sign applied to the activation with `vpsignb`. The arithmetic is ggml's `ggml_vec_dot_iq4_xs_q8_K`; only the order of the float additions differs. ## The gate is the part worth reviewing `native_expert.cpp` entered the multi-token path with `iq512_supported(f.gu_type)` alone. Widening that to `iq512_supported || iq256_supported` without also guarding each call is a **correctness trap, not just a missed optimisation**: IQ4_XS has an AVX-2 kernel and no AVX-512 one, so on an AVX-512 CPU the widened gate would enter `iq512_gu_rows`, hit a `switch` with no `case 23`, fall through it, return, and leave `ff` **unwritten** — garbage output rather than a fallback to ggml-cpu. Each call is therefore guarded by the formats its own kernel implements. `native_expert_parity.cpp` had the same gate, which is why the format was never tested on any CPU: with `iq512_supported` alone the tool skips IQ4_XS on AVX-512 machines and never reaches the AVX-2 check on AVX2 ones. ## Evidence Built from this branch (92/92 targets, no errors) and run against ggml's own `vec_dot` on the same Q8_K activations, real tensors from `orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF` (IQ4_XS), layers 0-9: ``` iq4_xs avx2 gate rows vs ggml vec_dot: rel 2.90e-08 ... 3.32e-08 (ten layers) native_expert_parity: 0 failures ``` Same order as the existing IQ3_S kernel on the same machine (median 3.10e-08, max 3.41e-08), i.e. float-addition order only. **Negative control**, so the parity gate is known to bite: changing the sub-block scale from `(ls - 32)` to `ls` gives ``` iq4_xs avx2 gate rows vs ggml vec_dot: rel 1.48e+00 iq4_xs avx2 gate rows MISMATCH (rel 1.48e+00 > 1e-5) native_expert_parity: 6 failures ``` Throughput, one thread, weights cache-resident (the same measurement the tool prints): | | 1 token | 3 tokens | |---|---|---| | IQ4_XS, this kernel | 226 us vs ggml 157 us (**0.70x**) | 308 us vs ggml 3x157 us (**1.53x**) | | IQ3_S, existing kernel | 364 us vs 244 us (0.67x) | 460 us vs 3x244 us (1.58x) | The signature is the existing one: slower for a single token, ~1.5x from two tokens up, which is what `STRATA_IQ_MT_MIN` (default 2) already encodes. ## End-to-end in the engine A/B with `STRATA_NO_IQ256=1`, which drops the AVX-2 multi-token path. On an IQ4_XS pack that path is exactly the gate/up rows: the `down` rows are IQ4_NL and have their own switch (`STRATA_NO_IQ4NL`), so they stayed on in both arms. IQ4_XS over `experts.bin` in mmap, `--kv-resident 20480`, `STRATA_IQ_PREFETCH=16384`, `--spec 4`, greedy, 300 tokens per run, three warmup passes and four interleaved rounds with the arm order rotated per round. Compared as **total time for the run** (rounds × ms/round): both arms generate ~300 tokens, and the draft acceptance moves tokens-per-round (2.63 vs 2.68) without moving the time, which is what makes ms/round alone misleading here. | round | with this kernel | ggml `vec_dot` | delta | |---:|---:|---:|---:| | 1 | 7617 ms | 7824 ms | **−2.6%** | | 2 | 7818 ms | 8116 ms | **−3.7%** | | 3 | 7645 ms | 10456 ms | −26.9% (the `vec_dot` arm collapsed: 127 draft rounds, 20.0 GB/s) | | 4 | 7829 ms | 7887 ms | −0.7% | **The kernel wins 4 of 4 pairs**; median −3.2%, −2.6% with the outlier round dropped. The gate/up phase is **−7.6%** (median), and the sustained row-read rate is 28.6 GB/s against 25.6. Small, and the reason is the machine: both arms sit at the DDR4-2666 wall (~28 GB/s), so this is not an ALU win. What it buys is that a verify window walks each expert row **once** instead of once per token in the window. ## Caveats - **No AVX-512 kernel for IQ4_XS**: `iq512_supported` is unchanged, so an AVX-512 CPU keeps using ggml-cpu for this format. Adding the AVX-512 variant is a separate change. - **IQ1_M still stays on ggml-cpu.** - The `down` rows of an IQ4_XS pack are IQ4_NL (type 20) and already have their AVX-2 multi-token kernel; this patch only covers gate/up. - The end-to-end number is one AVX2 machine, RAM-bound. On a machine with faster memory the gate/up share should matter more; on one with AVX-512 this format still does not reach a multi-token kernel at all.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.