Pull requests / #415

#415 IQ4_XS expert rows on the AVX-2 path: the multi-token kernel this format never had

closed · @pipeob0 · 0 评论 · 在 GitHub 查看

NVIDIA / CUDAModels & quantsWindows

描述

## What

IQ4_XS (GGML type 23) was missing from `iq256_supported`, so on an **AVX-2-only CPU** its gate/up
expert rows fell back to ggml's per-token `vec_dot` while every other i-quant format got the
multi-token kernel that decodes the weights once per verify window. This adds `Fmt32<23>` and
widens the two gates that were keeping the format out.

The machine is a Ryzen 7 5700X3D (Zen 3, AVX2, no AVX-512), Windows, engine 0.1.32 built from
source. It is the CPU class the AVX-2 path exists for, and IQ4_XS is one of the formats people
actually run on it — including the low-RAM path 0.1.31 added, which makes a 61 GiB-expert model
reachable on a 64 GB PC.

## How the kernel works

`Fmt32<23>` reuses the 16-value codebook `iq4nl_rows` already has. Each of the eight 32-value
sub-blocks carries its own 6-bit scale — two nibbles of `scales_l[p]` plus two bits of `scales_h` —
used **signed** as `(ls - 32)`, so the scale folds into the `int16` operand of `madd_epi16` instead
of becoming a float multiply per sub-block. The codebook sign is carried the way `iq4nl_rows` does
it: `|w|` as the unsigned `maddubs` operand, `w`'s sign applied to the activation with `vpsignb`.
The arithmetic is ggml's `ggml_vec_dot_iq4_xs_q8_K`; only the order of the float additions differs.

## The gate is the part worth reviewing

`native_expert.cpp` entered the multi-token path with `iq512_supported(f.gu_type)` alone. Widening
that to `iq512_supported || iq256_supported` without also guarding each call is a **correctness
trap, not just a missed optimisation**: IQ4_XS has an AVX-2 kernel and no AVX-512 one, so on an
AVX-512 CPU the widened gate would enter `iq512_gu_rows`, hit a `switch` with no `case 23`, fall
through it, return, and leave `ff` **unwritten** — garbage output rather than a fallback to
ggml-cpu. Each call is therefore guarded by the formats its own kernel implements.

`native_expert_parity.cpp` had the same gate, which is why the format was never tested on any CPU:
with `iq512_supported` alone the tool skips IQ4_XS on AVX-512 machines and never reaches the AVX-2
check on AVX2 ones.

## Evidence

Built from this branch (92/92 targets, no errors) and run against ggml's own `vec_dot` on the same
Q8_K activations, real tensors from `orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF` (IQ4_XS),
layers 0-9:

```
iq4_xs avx2 gate rows vs ggml vec_dot: rel 2.90e-08 ... 3.32e-08   (ten layers)
native_expert_parity: 0 failures
```

Same order as the existing IQ3_S kernel on the same machine (median 3.10e-08, max 3.41e-08), i.e.
float-addition order only.

**Negative control**, so the parity gate is known to bite: changing the sub-block scale from
`(ls - 32)` to `ls` gives

```
iq4_xs avx2 gate rows vs ggml vec_dot: rel 1.48e+00
iq4_xs avx2 gate rows MISMATCH (rel 1.48e+00 > 1e-5)
native_expert_parity: 6 failures
```

Throughput, one thread, weights cache-resident (the same measurement the tool prints):

| | 1 token | 3 tokens |
|---|---|---|
| IQ4_XS, this kernel | 226 us vs ggml 157 us (**0.70x**) | 308 us vs ggml 3x157 us (**1.53x**) |
| IQ3_S, existing kernel | 364 us vs 244 us (0.67x) | 460 us vs 3x244 us (1.58x) |

The signature is the existing one: slower for a single token, ~1.5x from two tokens up, which is
what `STRATA_IQ_MT_MIN` (default 2) already encodes.

## End-to-end in the engine

A/B with `STRATA_NO_IQ256=1`, which drops the AVX-2 multi-token path. On an IQ4_XS pack that path is
exactly the gate/up rows: the `down` rows are IQ4_NL and have their own switch (`STRATA_NO_IQ4NL`),
so they stayed on in both arms. IQ4_XS over `experts.bin` in mmap, `--kv-resident 20480`,
`STRATA_IQ_PREFETCH=16384`, `--spec 4`, greedy, 300 tokens per run, three warmup passes and four
interleaved rounds with the arm order rotated per round.

Compared as **total time for the run** (rounds × ms/round): both arms generate ~300 tokens, and the
draft acceptance moves tokens-per-round (2.63 vs 2.68) without moving the time, which is what makes
ms/round alone misleading here.

| round | with this kernel | ggml `vec_dot` | delta |
|---:|---:|---:|---:|
| 1 | 7617 ms | 7824 ms | **−2.6%** |
| 2 | 7818 ms | 8116 ms | **−3.7%** |
| 3 | 7645 ms | 10456 ms | −26.9% (the `vec_dot` arm collapsed: 127 draft rounds, 20.0 GB/s) |
| 4 | 7829 ms | 7887 ms | −0.7% |

**The kernel wins 4 of 4 pairs**; median −3.2%, −2.6% with the outlier round dropped. The gate/up
phase is **−7.6%** (median), and the sustained row-read rate is 28.6 GB/s against 25.6.

Small, and the reason is the machine: both arms sit at the DDR4-2666 wall (~28 GB/s), so this is not
an ALU win. What it buys is that a verify window walks each expert row **once** instead of once per
token in the window.

## Caveats

- **No AVX-512 kernel for IQ4_XS**: `iq512_supported` is unchanged, so an AVX-512 CPU keeps using
  ggml-cpu for this format. Adding the AVX-512 variant is a separate change.
- **IQ1_M still stays on ggml-cpu.**
- The `down` rows of an IQ4_XS pack are IQ4_NL (type 20) and already have their AVX-2 multi-token
  kernel; this patch only covers gate/up.
- The end-to-end number is one AVX2 machine, RAM-bound. On a machine with faster memory the gate/up
  share should matter more; on one with AVX-512 this format still does not reach a multi-token
  kernel at all.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。