Pull requests / #1697

#1697 cpu: AVX1 128-bit expert kernels (iq_avx1) - the i-quant rows on CPUs with no AVX2

open · @nashcap · 0 comentarios · En GitHub

BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quants

Descripción

## What

Completes the AVX1 floor for native packs. Since #394 an AVX-only CPU (Sandy Bridge through Ivy Bridge Xeons) starts and runs a native pack, but every i-quant dot in ggml-cpu is `#if defined(__AVX2__)` and falls to the scalar `_generic` path, and the fork's Q2_0 has no vectorized dot either - so the CPU-computed experts of a native pack run scalar there. This adds the missing sibling of `kq_avx1.cpp` for the expert rows:

- `src/kernels/cpu/iq_avx1.cpp` (+ header): gate/up for **IQ3_XXS, IQ3_S, IQ2_S** and down for **IQ4_NL, Q2_0**, multi-token (nt=1..8), in the same separate-TU pattern as `kq_avx1.cpp`: compiled `-msse4.2 -mavx -mno-avx2 -mno-fma -mno-f16c` (a `#error` guards against FMA/F16C leaking into the object), gated at run time behind `cpu_avx1_ok()` **and only when AVX2 is absent** - no AVX2/AVX-512 kernel is touched, so no mainstream path changes rounding or speed.
- 128-bit integer lanes (`pshufb`/`sign_epi8`/`maddubs` are SSSE3), 256-bit float accumulation, software half-to-float (no F16C), no gathers (the non-gather `Fmt32` path; `vpgatherdd` is AVX2).
- `q8k_quant_avx1`: the byte-identical Q8_K activation quantizer in SSE4.2/AVX1 (the AVX-2 copy's 256-bit packs become SSE2 packs, which need no lane-crossing fixup at 128-bit). 2.0x over ggml's scalar reference.
- Dispatch in `native_expert.cpp`: gate/up rows behind the same `mt_min` policy as the AVX-2 kernels (`STRATA_IQ_MT_MIN`), down rows always; kill switch `STRATA_NO_IQ128=1`.
- **IQ4_XS gate/up deliberately stays on ggml's generic dot**: its generic auto-vectorizes well and beats the 128-bit grid kernel (0.7-1.0x measured). The kernel is built and parity-tested; only the dispatch declines it. That layer's down rows still take the new kernels.

This is the shape #227 was asked to take before it could land: the native (IQ) paths covered, the choice made by the existing run-time probes rather than a force flag, and an end-to-end run on a real AVX-only machine (below).

## Tests

- `iq_avx1_parity` (new, registered in CTest): all six formats, nt=1..8 - every row bit-identical across widths (#152) and within the repo's 1e-5 aggregate bar against ggml's generic dots. It calls the kernels directly, so it exercises them on AVX2/AVX-512 CI as-is.
- `q8k_quant_parity` extended: the AVX1 copy is byte-identical to `quantize_row_q8_K_ref` over all 16 degenerate/NaN/inf cases x 4 sizes; registered with a `STRATA_FORCE_ISA=avx` variant.
- `iq_multi_parity`, `router_top10_parity`, `router_dot_parity`, `dequant_s2_parity`, `dequant_q5_1_test`, `expert_layout_test` pass on this machine.
- Engine-level, end to end: `strata serve` on the hardware below runs with the kernels active (the startup line reports `the expert kernels run on AVX1 128-bit (iq_avx1, the older-CPU build)`), the start's GPU-vs-CPU expert check (`slot 0 verified`) passes, and it has served real traffic since.

## Measured on the hardware this exists for

Intel **Xeon E5-2470 v2** (Ivy Bridge, AVX1 only, 10c/20t, DDR3), RTX 5060 Ti 16 GB, Qwen3.8-Flash-Next IQ3_S native pack, engine 0.1.40.3.

Per expert, weights streamed from DRAM, one thread (`iq_avx1_parity --bench`; gate/up + down together):

| pair (gate/up / down) | nt=1 | nt=2 | nt=4 |
|---|---|---|---|
| iq3_xxs / iq4_nl | 1.0x* | 1.3x | 1.5x |
| iq3_s / iq4_nl     | 1.0x* | 1.3x | 1.6x |
| iq2_s / iq4_nl     | 0.9x* | 1.2x | 1.5x |
| iq3_xxs / q2_0     | 2.0x  | 2.4x | 2.9x |
| iq3_s / q2_0       | 1.8x  | 2.4x | 3.0x |
| iq2_s / q2_0       | 1.8x  | 2.4x | 2.9x |

\* at nt=1 the gate/up rows are gated by `mt_min=2` (the AVX-2 policy, #152) - ggml's generic is already competitive for these formats at one token, the win is in amortizing the row decode over the draft window.

Serve, same 5276-token synthetic prompt, greedy, 150 tokens, before -> after: cold prefill **430 -> 581-607 tok/s**, warm-reuse decode **22-25 -> 35-46 tok/s**; real interactive traffic 22-23 -> 25-26.5 tok/s.

## Found while testing (pre-existing, not touched here)

`pool_tasks_test` dies with SIGILL on non-AVX-512 machines even on a clean tree (v0.1.41, this hardware): it predates the ISA floor and appears to reach an AVX-512 path unconditionally, so its CTest registration would fail a runner without AVX-512.

En el sitio

Enlaces a install, modelos, releases.