Pull requests / #863

#863 cpu kernels: gathered decode on the cores where it is faster, AVX-VNNI rows, IQ3_S one-token kernel

closed · @Hardin22 · 0 comentarios · En GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quants

Descripción

Three changes to the AVX2 expert kernels. Each one gives the same bits as the scalar path, except (c), which follows the #152 rule.

1. **Gathered grid decode, chosen per core** (`STRATA_IQ256_GATHER`: unset = auto, 0 = off, 1 = every core).
   - Fmt32<118>/<121>/<122> become the only gathered decodes. They replace the in-decode branch from dd24a55.
   - They also gather the IQ3_XXS sign masks and the IQ2_S grid.
   - **Auto** gathers on Intel P-cores of Alder Lake / Sapphire Rapids and newer:
     - GenuineIntel with AVX-VNNI is the generation marker. That excludes Haswell/Broadwell and the Downfall-microcode generations.
     - The E-core-only parts are excluded by model.
     - On a hybrid CPU, only a thread on a core of CPUID 1Ah type 40h.
     - AMD keeps the scalar decode.
2. **AVX-VNNI**: `vpdpwssd` / `vpdpbusd` forms of the i-quant, IQ4_NL and Q2_0 rows. Integer sums, so the same bits.
   - Gated by `cpu_avxvnni_ok()`, which lives in `expert_layout.cpp` (not compiled for AVX2) and respects `STRATA_FORCE_ISA` and `STRATA_ISA_FLOOR`. `STRATA_NO_AVXVNNI=1` turns it off.
   - On GCC/Clang only the VNNI copies get `target("avxvnni")`, and only if CMake finds the intrinsics. No file gets `-mavxvnni`.
3. **IQ3_S takes the multi-token kernel for a lone token** wherever its grid is gathered. #152 sent one token to ggml's dot because the kernel was no faster there; with the gather it is much faster. This changes the rounding of a lone token's IQ3_S gate/up rows to the group's. On these CPUs that makes IQ3_S independent of the window size. `STRATA_IQ_MT_MIN` keeps its meaning.

**Test**: `iq_avx2_parity`, a CPU-only CTest. Random blocks for IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ3_S and IQ4_XS gate/up rows and IQ4_NL / Q2_0 down rows. Every variant at 1..8 tokens must match the scalar variant bit for bit; the scalar rows are checked against ggml's vec_dot (Q2_0 against a double reference).
- 0 failures on an i9-14900KF, also with `STRATA_IQ256_GATHER=0`, `STRATA_IQ_MT_MIN=1`, `STRATA_NO_AVXVNNI=1` and `STRATA_FORCE_ISA=avx2`, and pinned to a P-core and to an E-core.
- With `STRATA_FORCE_ISA=avx` it skips cleanly.
- A deliberately broken gather was caught.

## Measured (`iq_avx2_parity --bench --dispatch`, i9-14900KF, one thread, more expert blobs than L3, median of 5)

Engine path, ms per expert (gate/up + down) at 1-4 tokens. Upstream main is shown at its default (`STRATA_IQ256_GATHER` unset = 0) and with `=1`:

| format (gate/up / down) | core | tokens | main | main, gather=1 | this PR (auto) | vs main |
|---|---|---|---|---|---|---|
| IQ3_S / IQ4_NL | P | 1 | 0.431 | 0.405 | **0.257** | −40% |
| IQ3_S / IQ4_NL | P | 2 | 0.493 | 0.380 | **0.312** | −37% |
| IQ3_S / IQ4_NL | P | 4 | 0.642 | 0.535 | **0.455** | −29% |
| IQ3_XXS / IQ4_NL | P | 1 | 0.271 | 0.272 | 0.276 | +2% (noise) |
| IQ3_XXS / IQ4_NL | P | 2 | 0.362 | 0.325 | **0.291** | −20% |
| IQ2_S / Q2_0 | P | 2 | 0.377 | 0.380 | **0.298** | −21% |
| IQ2_XXS / Q2_0 | P | 4 | 0.419 | 0.405 | **0.369** | −12% |
| IQ3_S / IQ4_NL | E | 2 | 1.060 | 1.472 | **0.971** | −8% |
| IQ3_XXS / IQ4_NL | E | 4 | 1.196 | 1.741 | **1.046** | −13% |
| IQ2_S / Q2_0 | E | 4 | 1.252 | 1.264 | **1.138** | −9% |

The full grid has 5 formats × 1-4 tokens × P/E core. Nowhere is it slower than main beyond noise. On E-cores, `gather=1` is 40-75% slower, which is why auto leaves them alone.

**End to end**: Swift 1.5 IQ3_XXS on an RTX 5080 alone (most of the window is CPU experts) showed no difference beyond noise. That is expected: its CPU experts mostly see one token, and IQ3_XXS at one token is unchanged. I don't have an IQ3_S model here, so the end-to-end effect on IQ3_S, where one token goes from 0.43 to 0.26 ms per expert, is **not measured**. A report from an IQ3_S rig would settle it.

## Not tested

- The GCC/Clang path: no GCC or Clang on this PC. The `target("avxvnni")` copies and the CMake check have never been compiled. If the check fails, the build simply leaves the VNNI copies out.
- AMD CPUs: auto keeps them on today's path by construction.

En el sitio

Enlaces a install, modelos, releases.