Pull requests / #394
#394 AVX1: run on a CPU with AVX but no AVX2, with an end-to-end run on Sandy Bridge-E
closed · @demetree · 0 commentaires · Sur GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quants
Description
Follow-up to #227, which was closed with four specific objections. Each is addressed below, and the
end-to-end run is on a **pre-Haswell CPU**.
## The end-to-end run
Hardware: **Xeon E5-2680 (Sandy Bridge-E)** — no AVX2, no FMA3, no F16C, no BMI2, no AVX-512 — and a
**GTX 1080 Ti (sm_61, Pascal)**, below this project's sm_75 floor, so the engine was built with
`STRATA_EXPERIMENTAL_SM60=ON` (#236). Model: native IQ2_XS pack, 33.02 GiB loaded, CUDA 12.6, MSVC 19.44.
```
strata generate: PCIe probe: 0.8 GB/s host->device -> pcie_frac 0.00 (default 0.55)
strata generate: 1466 MiB of weights loaded (302 canonical tensors skipped: served natively)
strata generate: loaded 33.02 GiB at 3.68 GiB/s
strata generate: expert cache 4176 slots, 5.60 GiB of VRAM
strata generate: pre-filled 4176 of 4176 slots from the profile; slot 0 verified
strata generate: R4 hit path ON - resident experts are computed on the GPU
strata generate: session is up (engine 0.1.31)
strata generate: prefill 2 tokens in 1 chunks, 3023.8 ms; experts streamed 552 (552 by DMA, resident 193)
strata verify: captured the 1-token window (upload no error, sync no error)
strata verify: captured the 2-token window (upload no error, sync no error)
prompt : 1 2 3
output : 4 5
decode 2 tokens in 770.3 ms -> 2.60 tok/s
prefill 2 tokens in 3392.3 ms -> 0.59 tok/s (time to first token 4652.0 ms)
ENGINE_EXIT=0
```
`CPU experts 8.98 distinct / 11.58 routed per layer` — the CPU path is genuinely exercised, not
bypassed.
### An actual answer, not just an exit code
Same machine, same build, same pack — a real prompt, 80 tokens generated. The ids were produced and
checked with the tokenizer `tools/iq_pack.py` shipped in the pack (`vocab.json` + `merges.txt`,
byte-level BPE); the round-trip reproduces the prompt character-for-character, so the ids are the
model's own.
```
prompt : Explain in two sentences why the sky appears blue:
output : "\n\n<think>\nThe user wants a concise two-sentence explanation of why the sky appears
blue. I need to cover Rayleigh scattering and why blue specifically. Let me craft a clear,
accurate two-sentence response.\n</think>\n\nSunlight contains all colors, but as it
passes through Earth's atmosphere, shorter-wavelength blue light is scattered much more
strongly by gas molecules than longer-wavelength red light"
decode 80 tokens in 27918.3 ms -> 2.87 tok/s
prefill 10 tokens in 6491.8 ms -> 1.54 tok/s (time to first token 7796.0 ms)
```
The answer is correct — Rayleigh scattering, in two sentences, as asked.
Speculative decoding accepts little here, and that is the hardware rather than the change:
```
speculation 78 rounds of 4, drafts accepted 2 of 80 (0.025), 1.03 tokens per round
verify window ... CPU experts 5.88 distinct / 6.28 routed per layer
pool multi gate/up 19.300 quantize 0.427 down 7.235 ms/round; CPU pool call 28.412 ms/round
```
The CPU pool is the bottleneck, so the draft head cannot run ahead of it. On a machine whose CPU path
runs at full speed the acceptance rate should be far higher; 2.87 tok/s is a floor set by the CPU, not
by this change.
## The four objections
### 1. "native (IQ) packs ... aren't covered (`iq_avx2.cpp` has BMI2 and would still trap on an AVX-only CPU)"
Correct, and the trap was worse than described: it happened **before `main()`**. `iq_avx2.cpp` builds a
static sign table with `<< (8 * k)`, a *variable* shift, which MSVC under `/arch:AVX2` emits as BMI2
`shlx`. `even_signs` is a namespace-scope static, so the constructor runs at process start — `strata.exe`
exited `0xC000001D` with no output at all, not even for `--help`. (Submitted separately as #391, since
it needs none of this.)
Then there was a second trap. `native_gu_rows()` gated the AVX-512 kernel correctly and the AVX2 one
not at all:
```cpp
static const bool avx512 = cpu_avx512_ok() && getenv("STRATA_NO_IQ512") == nullptr;
static const bool avx2 = getenv("STRATA_NO_IQ256") == nullptr; // no cpu_avx2_ok()
```
so a verify window with two or more tokens reached `iq256_gu_rows` and trapped on `vfmadd231ps`. The
same asymmetry guarded `native_down_rows()`'s two multi-token kernels. All three now require
`cpu_avx2_ok()`.
**This does not need an AVX1 kernel, and that is the useful part.** `native_expert.cpp` already computes
native IQ experts *"through ggml-cpu"*, and ggml-cpu ships a `vec_dot` for every one of these types —
`ggml_vec_dot_iq2_xxs_q8_K`, `_iq2_s_q8_K` and the rest — compiled for whatever baseline the build
selected. That path sits *below* both multi-token kernels and loops over tokens itself, so it is
correct for any `nt`. It was unreachable only because skipping the multi-token kernels required an
environment variable. With the gate fixed, a pre-Haswell CPU falls through to it.
So the native path is covered by **three one-line guards**, not by 655 lines of re-derived quant
kernels — which is also why the numbers above are exact rather than approximate.
### 2. "the release links ggml with an AVX2 baseline anyway"
Understood and unchanged — this is a deliberate portability choice, not something to override. Two
distinct regimes, and the second is what an AVX1 build needs:
- `STRATA_PORTABLE=ON` → `GGML_NATIVE=OFF` + `GGML_AVX/AVX2/FMA/F16C/BMI2` all `ON`: one engine for any
AVX2 PC. Correct as it stands.
- `STRATA_PORTABLE=OFF` (default) → `GGML_NATIVE=ON`, host-detected. An AVX1 baseline build then works,
and configure confirms it: `Adding CPU backend variant ggml-cpu: /arch:AVX GGML_AVX`, with
`HAS_AVX2_1 - Failed`, `HAS_FMA_1 - Failed`, `HAS_AVX512_1 - Failed`.
So the run above is an AVX1 ggml build. What this branch does *not* do is add an explicit
`GGML_NATIVE=OFF` + AVX1-only option for someone who wants a portable engine *and* an old CPU. That
would be a separate, small CMake change and I would rather ask than assume.
### 3. "`STRATA_FORCE_AVX2=1` would exit on AVX-512 machines"
Real bug, and mine. Our `cpu_avx1_ok()`/`cpu_avx2_ok()` tested the **mere presence** of the variable while
`cpu_avx512_ok()` tests `f[0] == '1'`. All three rungs therefore returned false at once on an AVX-512
machine and the startup gate exited. Both now use `f[0] == '1'`, with the reason in the source.
### 4. "no end-to-end run"
See above.
## What this branch contains
1. **The AVX1 rung** for the canonical Q2_0 expert path — `q2_avx1.cpp` and `s2_expert_avx1.cpp`, three-way
dispatch, `cpu_avx1_ok()` probing AVX + SSSE3 + SSE4.1 and deliberately *not* FMA3/F16C (verified by
**executing** `vfmadd*` and `vcvtph2ps` on the target CPU, where each raises `#UD`).
2. **The startup floor** moved from `cpu_avx2_ok()` to `cpu_avx1_ok()`, message updated to name the new
floor.
3. **The three AVX2 multi-token gates** above.
4. **`cpu_avx2_ok()` corrected** to leaf 7 EBX bit 5 — leaf 1 ECX bit 5 is plain AVX (Sandy Bridge), not
AVX2 (Haswell). Our version of this was wrong and would have routed a pre-Haswell CPU into the AVX2
kernel; the test caught it as `avx2=1 avx1=1` on a CPU with no AVX2.
## Tests
Two self-contained parity tests, `q2_avx1_parity` and `s2_avx1_parity`, both of which need neither the
38 GB pack nor an AVX-512 CPU — `expert_parity` **skips itself on exactly the hardware this targets**,
so without them this kernel would ship untested.
```
ctest -R avx1 -> 2/2 passed
nt=2 batching is bit-identical to nt=1 per token (both tests)
rows outside [r0,r1) untouched: 0 clobbered
ISA containment by disassembly: no zmm, no BMI2, no FMA, no F16C in q2_avx1.cpp.obj / s2_expert_avx1.cpp.obj
```
## Caveats
- Speed is not the point of the AVX1 path. It is roughly half the AVX2 kernel's throughput (16 int8 per
instruction instead of 32) and nothing in accuracy. Its purpose is to stop an illegal instruction.
- The AVX1 rung is for the **canonical** Q2_0 pack. For a **native** IQ pack, correctness on a
pre-Haswell CPU comes from ggml-cpu's own kernels via the fixed gate, not from anything in this branch.
- The sm_61 GPU is unrelated to the CPU work and was only ever a means to run the engine here. If
`STRATA_EXPERIMENTAL_SM60` is useful to anyone else on a card below the community-tested set, the
configuration in the log above is a working data point for it.
Sur le site
Liens install, modèles, releases.