Pull requests / #227
#227 AVX1: the Q2_0 expert kernels for CPUs that have AVX but no AVX2
closed · @demetree · 0 コメント · GitHub で見る
Setup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
本文
## What this adds
A third rung below the AVX2 one in the CPU expert ladder, for CPUs that have AVX but no AVX2 —
Sandy Bridge and Westmere Xeons, and the i7/i5 parts of the same vintage. They are still in service,
and today a pre-AVX-512 CPU either fails at startup or takes an illegal instruction.
AVX-512 (expert.cpp) -> AVX2 (q2_avx2.cpp) -> AVX1 (q2_avx1.cpp, this PR)
Each rung is its own translation unit compiled for exactly that ISA, and the choice is made at run
time by the probes in `expert_layout.cpp`. That containment is the point: a probe that returns the
wrong answer costs speed, not a trap.
This was asked for on X, and built on a dual Xeon E5-2680 (Sandy Bridge-E) that happens to sit under
the CUDA floor — so it is contributed untested against a real model load, and the verification below
is kernel-level.
## What it does *not* do
Worth being explicit, because it is the obvious first question:
- **It does not make Strata run on an AVX1 CPU by itself.** The CUDA arch guard is at compute
capability 7.5, so a pre-Ampere *card* is refused regardless of CPU support. This PR is about the
CPU side only.
- **It does not make the canonical path faster than AVX2 where both are available.** The AVX-512 kernel
reduces eight float lanes and the AVX1 one reduces four, so the sums are grouped differently. The
two are not bit-identical to each other and are not expected to be; each is checked against the
scalar reference separately.
- **It does not add FMA3 or F16C.** Those arrived one generation after AVX (Ivy Bridge, 2011) and
`q2_avx1.cpp` does a software fp16 decode and a mul+add instead.
## Commits
1. **`2ed3b8c`** — the NATIVE-pack kernel (`q2_avx1.cpp`), the `cpu_avx1_ok()` probe, three-way dispatch
in `q2_rows_any()` / `act_quant_any()`, the startup gate, and `q2_avx1_parity`.
2. **`d5eda68`** — the CANONICAL expert path (`s2_expert_avx1.cpp`), the `s2_expert_*_any`
dispatchers, and `s2_avx1_parity`. This path is not optional: `generate.cpp` calls the startup gate
for it unconditionally, and every function `pool.cpp` calls lives in an `/arch:AVX512` translation
unit — so the whole compute path needs a rung below AVX-512, not just the dequant.
3. **`1c2f8bc`** — two defects found in the above while validating, both described below.
## Why the tests here are self-contained
`expert_parity` needs the 38 GB pack and an AVX-512 CPU, so it **skips itself on exactly the hardware
this PR targets**. Without `q2_avx1_parity` / `s2_avx1_parity`, a build of this would ship with the
new kernel untested. Both new tests synthesise their own expert blobs and check against an independent
longhand scalar reference, so they need neither the pack nor AVX-512 — they are the only tests in the
tree that run on a pre-AVX2 CPU.
## Verification
Reference CPU: **Xeon E5-2680, family 6 model 45 stepping 7 (Sandy Bridge-E)** — no FMA3, no F16C, no
AVX2, no AVX-512, no BMI2. CPU-only build, MSVC 19.44, CMake 3.31.6, Ninja.
```
ctest -R avx1 -> 2/2 passed
q2_avx1_parity 0 failures
act_quant_q8_1_avx1 == scalar rule, exactly 0/2560 codes differ, max scale delta 0.00e+00
all 640 gate rows match reference rel L1 3.23e-07
nt=2 batching matches nt=1 per token token0 0.00e+00, token1 0.00e+00
s2_avx1_parity 0 failures
s2_expert_vnni_q_avx1 vs reference rel L1 2.31e-07
rows outside [r0,r1) untouched 0 clobbered
s2_expert_vnni_multi_avx1 (nt=2) == nt=1 token0 0.00e+00, token1 0.00e+00
s2_expert_vnni_q_any (what pool.cpp calls) rel L1 2.31e-07
```
`nt=2` is bit-identical to `nt=1` in both, which matters because the multi-token path is the one that
decodes more than one token at a time.
**ISA containment checked by disassembly** of the built objects (`dumpbin`), not by reading flags:
```
q2_avx1.cpp.obj no zmm, no BMI2, no FMA, no F16C
s2_expert_avx1.cpp.obj no zmm, no BMI2, no FMA, no F16C
```
The probe decisions were established by **executing** the candidate instructions on the reference CPU
(`vfmadd*` and `vcvtph2ps` each raise `0xC000001D`, `pmaddubsw` works) rather than trusting CPUID
alone — which is what caught the AVX2 bug below.
## Two bugs found in this work, both mine
Worth flagging because both were invisible to reading the code and only showed up when the tests were
actually run:
**1. `cpu_avx2_ok()` probed the wrong bit.** It read leaf 1 ECX bit 5 and treated it as AVX2. Leaf 1
ECX bit 5 is plain AVX (Sandy Bridge, 2011); AVX2 is leaf 7 subleaf 0 EBX bit 5 (Haswell, 2013). On a
CPU with AVX but no AVX2 it returned true, so the ladder would have handed it the `/arch:AVX2` kernel on
hardware that cannot execute it. `s2_avx1_parity` caught it by asserting the rung choice and
reporting `avx2=1 avx1=1`. Now reads leaf 7 EBX bit 5 with a max-leaf ≥ 7 guard.
**2. A BMI2 fault in a static initializer.** `s2_avx1_parity` exited `0xC000001D` before `main()`
printed a line. `expert_layout_load()` calls `native_fmt()`, so linking the test against
`strata_kernels_cpu` pulled in `native_expert.cpp` → `iq_avx2.cpp`, which is `/arch:AVX2` and contains
BMI2 (`shlx`). The reference CPU has no BMI2, and the fault was in a **static initializer**, so it ran
before any dispatch mattered. `iq_avx2.cpp` did not exist at the 0.1.20 base (`df6980d` added it),
which is exactly why the test passed before the rebase onto 0.1.28 and failed after.
Fixed by stubbing `native_fmt` to break the link edge, in the same abort-on-call style as the existing
AVX-512 stubs. `native_fmt` is the one exception and returns false with a reason rather than aborting,
because `expert_layout_load()` legitimately calls it during setup; the other five `native_*` entry
points abort, since reaching them means the AVX1 path was bypassed.
## Notes for review
- Rebased onto **0.1.28** (`bbaaabb`) — the current tip. No conflicts, despite that release touching
`CMakeLists.txt` and `generate.cpp`, which this also touches.
- The `/arch:AVX` flags are set **per source file**, matching how `expert.cpp` and `q2_avx2.cpp` are
already handled. The MSVC branch deliberately omits `/mfma` and `/mf16c`; the GCC/Clang branch
deliberately omits `-mfma` and `-f16c`, because the target CPU has neither and enabling them
produces instructions that trap on the very machine this is written for.
- The Zig toolchain in this tree cannot build 0.1.28 at all — pristine upstream fails identically,
since `expert.cpp` and `iq_avx512.cpp` need `evex512`. That is a toolchain limitation, not something
in this branch; MSVC, which is what upstream targets, builds it fine.
関連リンク
インストール・モデル・リリースへの站内リンク。