Pull requests / #257

#257 q2_avx2: two-block 256-bit unpack for the AVX2 Q2_0 expert rows (bit-exact)

closed · @hireymage · 0 评论 · 在 GitHub 查看

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quants

描述

Title:
q2_avx2: two-block 256-bit unpack for the AVX2 Q2_0 expert rows (bit-exact)

Body:

## What this does

On CPUs without AVX-512 (Intel Core 12th–14th gen, AMD Zen 2/3 — the
machines `q2_avx2.cpp` is built for), the Q2_0 expert rows run a
128-bit-only shuffle ladder that unpacks every 64-code weight block from
its 16 code bytes. This PR halves the ladder work for the most common
routed shape: rows that serve 1–2 activations. A 256-bit pass covers two
consecutive blocks at once, so the ladder runs once for a pair of weight
blocks instead of twice.

- `unpack64x2`: lane 0 holds block b's 16 code bytes, lane 1 holds block
  b+1's (Q2_0 blocks are 18 bytes apart, so lane 1 loads at `codes + 18`);
  every shift/and is per-128-bit lane, i.e. numerically the SSE ladder run
  twice.
- The per-lane halves are re-homed so each block keeps the original byte
  order → the same `madd`/`maddubs` lane mapping → **bit-identical results**.
  Verified by `memcmp` against the original code in a standalone harness
  (all of gate/up and down, tokens 1/2/4/8).
- Active only for 1–2 tokens per row (`rows<1>`/`rows<2>`). At 4+ tokens
  the extra live unpack vectors spill registers and it measurably loses,
  so the original ladder stays in place there. AVX-512 machines are
  unaffected (they use a different path dispatch).

## Measured on my machine

i7-8700K (AVX2, 6C/12T), RTX 2070 (sm_75), Qwen3.8-Flash-Next native Q2_0
pack (spec verified from the GGUF header, ~125.7B-parameter MoE), engine
0.1.27 build, temperature 0, 5 prompts × 3 warm passes per variant, and a
fresh baseline rerun in the same time block to rule out machine drift.

- Standalone harness (bench5, bit-exact-verified): **+10 % gate/up rows,
  +5–7 % down rows at 1 token/expert**; +2–4 % at 2 tokens; −5–11 % at
  4 tokens (why the guard exists).
- End to end: verify-window total **median ~115.7 → ~107.1 ms (−6–7 ms,
  ~−6 %)**, coming cleanly out of the CPU expert-pool stage (stage
  profiler `waitCPU`: GDN ~18.9 → ~13.1 ms, QSA ~4.4 → ~3.2 ms per
  window). GPU kernel stages unchanged, and generation token counts are
  identical across all runs (same prompts, temperature 0).
- Decode: +1–3 tok/s per prompt.

## Please test on other devices

I can only measure one machine, and both the gain and the end-to-end
visibility depend on the platform:

- **Other AVX2-only CPUs** (Intel Core 12th/13th/14th gen, Core Ultra
  without AVX-512 enabled, AMD Zen 2/3): the per-row gain is expected to
  be universal, but its end-to-end size depends on how much of decode is
  the CPU expert pool.
- **Models with fewer experts / more active tokens** (rows serving 3+
  tokens are more common there) — the `nt <= 2` guard should keep them on
  the original ladder; a check that nothing changes on that side is
  welcome.
- **AVX-512 machines** should behave exactly as before (different code
  path) — a confirmation is also useful.

Greedy outputs should be identical before/after; if you see a difference,
that's a bug, please report it.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。