Pull requests / #930
#930 cpu: IQ3_S / IQ2_S AVX2 decode without the per-index shift/mask/or (1.07-1.6x per row)
closed · @sanastasiou · 0 comentários · No GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quants
Descrição
The AVX2 multi-token kernels in `src/kernels/cpu/iq_avx2.cpp` are compute-bound on the codebook decode, not on the per-token arithmetic: a hot (in-cache) and a cold (256 MiB) working set measure the same ns/row. This PR makes the IQ3_S and IQ2_S decode cheaper. The arithmetic is untouched (`vpsignb`, `vpmaddubsw`, `vpmaddwd`, exact int32 per QK_K block), so the results are bit-identical.
**What changes (one file, IQ3_S and IQ2_S only)**
- **Grid indices.** The `qh` byte indexes a compile-time table that spreads its high bits one per index byte. One `punpcklbw` puts them next to the 8 low index bytes, so each lookup is a 16-bit load plus the grid load instead of a shift, a mask and an or per index.
- **Sign vectors.** The half's four sign bytes are broadcast straight from the block (`vpbroadcastd` from memory, a load uop) and expanded with ONE `pshufb`, instead of `sgn_vec`'s two `pshufb` + `vinserti128`. On Intel cores those are three uops on the single shuffle port.
- **Scales.** The `(2s+1)` vectors come from compile-time tables: one load instead of shift/or/broadcast.
- **No static constructor.** The tables are `constexpr` (like `EvenSigns`), so the AVX2 TU still has no static constructor (#391). `static_assert`s pin the table layout.
- **Gather opt-in kept, branch hoisted.** IQ3_S's `STRATA_IQ256_GATHER` path becomes its own instantiation, chosen once per call in `iq256_rows` / `iq256_gu_rows`. With the check left inside `decode()`, the branch sat in the innermost loop and made the new IQ3_S decode 25-35 % slower at nt=2-3 (default path, Zen 3).
- **IQ3_XXS unchanged.** The same hoist measured -5..+5 % there, which is not worth it.
**Measured** (ns per row of all nt tokens, hot, one pinned thread, best of 3 rounds with base and PR alternated; base = `main` @ 6f32ec0; the kernels compiled with `-march=haswell`):
| CPU | type | nt=1 | nt=2 | nt=3 | nt=4 |
|---|---|---|---|---|---|
| Ryzen 9 5950X (Zen 3) | IQ3_S | 256.1 → 213.7 (**1.20x**) | 304.4 → 267.5 (1.14x) | 360.7 → 313.7 (1.15x) | 399.8 → 374.1 (1.07x) |
| | IQ2_S | 208.0 → 144.1 (**1.44x**) | 225.5 → 192.9 (1.17x) | 255.6 → 222.7 (1.15x) | 279.6 → 271.2 (1.03x) |
| Ryzen 5 3600 (Zen 2) | IQ3_S | 337.7 → 270.4 (**1.25x**) | 386.4 → 320.5 (1.21x) | 451.4 → 378.5 (1.19x) | 524.8 → 437.3 (1.20x) |
| | IQ2_S | 244.5 → 172.7 (**1.42x**) | 278.3 → 222.8 (1.25x) | 320.0 → 265.8 (1.20x) | 360.7 → 327.0 (1.10x) |
| Core i7-7700 (Kaby Lake, AVX2, no AVX-512) | IQ3_S | 405.6 → 354.6 (1.14x) | 616.6 → 382.8 (**1.61x**) | 507.2 → 426.5 (1.19x) | 573.3 → 492.4 (1.16x) |
| | IQ2_S | 310.5 → 208.4 (**1.49x**) | 359.8 → 254.6 (1.41x) | 404.8 → 313.4 (1.29x) | 440.1 → 365.9 (1.20x) |
On the i7-7700, `main`'s IQ3_S at nt=2 is slower than at nt=3. This reproduced on two cores and two runs; with this PR the curve is monotonic again.
I have not measured end to end in the engine, and I have not measured on a Haswell Xeon. The Kaby Lake core has the same single-shuffle-port layout, so I expect Haswell to land in the same range, but that is not measured. Numbers from an E5 v3 owner (#913) would close that gap.
**Tests**
- **Parity harness.** A standalone harness (our private test bench; happy to contribute it as a CPU-only `tests/` target if useful) compiles this exact `iq_avx2.cpp` and checks `iq256_rows` against ggml's `_generic` references and a float64 oracle.
- Every type, nt 1..9, row ranges with odd counts and matrix borders: 844 tests shuffled, pass on Zen 3.
- The `strata_iq256` subset: 180 tests, pass on Zen 2 and Kaby Lake.
- The same 844 tests also pass with `STRATA_IQ256_GATHER=1`.
- **Mutation-checked.** Breaking the IQ3_S high bits or the IQ2_S sign offset fails the parity tests. Swapping the IQ2_S scale nibbles fails at compile time on the `static_assert`.
- **Upstream build.** CPU-only `cmake -DSTRATA_WERROR=ON -DSTRATA_BUILD_TESTS=ON` builds clean.
- ctest: 15/16 pass. `expert_multi_test` fails on this Zen 3 box because it requires AVX-512 ("this CPU cannot run the expert kernel"), not because of this change.
- `native_expert_parity` needs CUDA plus the llama.cpp oracle and does not cover the IQ formats; I did not run it.
- **No AVX-512.** Nothing in the shipped object uses EVEX or `zmm`; this is checked by disassembly.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
No site
Links install, modelos, releases.