Pull requests / #622

#622 iq: gather the AVX2 codebook from the grid table (STRATA_IQ256_GATHER)

closed · @yannickloth · 0 comentarios · En GitHub

BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quants

Descripción

## What

Add an opt-in AVX2 gather for the two IQ3 gate/up codebooks, the mirror of the AVX-512 kernels' existing `STRATA_IQ_GATHER`.

`STRATA_IQ256_GATHER=1` replaces the scalar `_mm256_set_epi32` grid assembly in `Fmt32<18>` (IQ3_XXS, `iq3xxs_grid`) and `Fmt32<21>` (IQ3_S, `iq3s_grid`) with one `_mm256_i32gather_epi32`. Default **off**; the engine is unchanged unless the variable is set.

## Why

The AVX2 multi-token kernel (`iq_avx2.cpp`) is the path for the non-AVX-512 CPUs this engine targets (Intel Core 12th-14th gen, Core Ultra, Zen 2/3). P4 measured i-quant gate/up as the largest CPU term, and this file documents it as codebook-lookup bound. A scratch build that substitutes a constant for the grid vector shows the grid assembly is ~half the single-core kernel:

| format | full | grid bypassed | speedup |
| --- | ---: | ---: | ---: |
| IQ3_XXS (18) | 248 us, 5.1 GB/s | 130 us, 9.7 GB/s | 1.91x |
| IQ3_S (21) | 357 us, 4.0 GB/s | 140 us, 10.0 GB/s | 2.55x |

IQ3_XXS + IQ3_S are 19 of 48 layers and ~58% of the gate/up lookup work on the Swift 1.5 IQ3_XXS pack, so this is where a runtime decode change can matter.

## Numbers (i7-12850HX, Alder Lake, no AVX-512; RTX A3000; engine 0.1.38)

Isolated `native_expert_parity --synthetic`, one thread, pinned to a P-core, weights L2-resident, gate+up one token:

| format | scalar | gather | speedup |
| --- | ---: | ---: | ---: |
| IQ3_XXS | 245 us, 5.13 GB/s | 185 us, 6.77 GB/s | 1.32x |
| IQ3_S | 334 us, 4.21 GB/s | 250 us, 5.64 GB/s | 1.34x |

Serve A/B, same binary, `bench/e2e.sh`, fixed prompt, interleaved, `decode_tok_s` from `/metrics` (off / on): 36.4/36.8, 36.9/37.0, 36.5/36.9, 37.1/37.1, 35.9/36.9 — median 36.5/36.9, mean **36.56/36.94 (+1.0%), on ≥ off 5/5**.

The deterministic `--stats` pool is flat (gate/up off 9.756 / on 9.773 ms/round): at 15 workers the pool is not decode-bound enough for the gather to move the whole token, so this is a per-core win that compounds with other work, not a structural one.

## Parity

- A direct same-binary output dump, scalar vs gather, is **bit-identical** for all five gate/up formats (`cmp` clean; 7680 bytes each).
- `native_expert_parity` stays green with the switch on and off (same `rel ≈ 3.1e-08` against ggml's `vec_dot`), width invariance 0 rows.
- The AVX-512 path is untouched. IQ2_XXS / IQ2_XS / IQ2_S use 64-bit grids, where an AVX2 `i64gather` measured a wash (IQ2_XXS 184→185 us, IQ2_S 246→284 us), and are unchanged.

## Caveats

- AVX2 gather throughput is microarchitecture-dependent, which is why this is opt-in exactly like the AVX-512 `STRATA_IQ_GATHER` (that one measured 3-5% slower on Zen 4). It should not be defaulted on.
- Measurement caveat: pin the harness thread or use the engine A/B; unpinned isolated runs are bimodal on this host.

Full write-up with the raw runs: #616 (and the plan's P7). Model/config: Swift 1.5 IQ3_XXS pack, `--adapt-every 1 --pool-workers 15`, fixed prompt `hit_rate 0.5509`.

En el sitio

Enlaces a install, modelos, releases.