Pull requests / #257
#257 q2_avx2: two-block 256-bit unpack for the AVX2 Q2_0 expert rows (bit-exact)
closed · @hireymage · 0 comentarios · En GitHub
BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quants
Descripción
Title: q2_avx2: two-block 256-bit unpack for the AVX2 Q2_0 expert rows (bit-exact) Body: ## What this does On CPUs without AVX-512 (Intel Core 12th–14th gen, AMD Zen 2/3 — the machines `q2_avx2.cpp` is built for), the Q2_0 expert rows run a 128-bit-only shuffle ladder that unpacks every 64-code weight block from its 16 code bytes. This PR halves the ladder work for the most common routed shape: rows that serve 1–2 activations. A 256-bit pass covers two consecutive blocks at once, so the ladder runs once for a pair of weight blocks instead of twice. - `unpack64x2`: lane 0 holds block b's 16 code bytes, lane 1 holds block b+1's (Q2_0 blocks are 18 bytes apart, so lane 1 loads at `codes + 18`); every shift/and is per-128-bit lane, i.e. numerically the SSE ladder run twice. - The per-lane halves are re-homed so each block keeps the original byte order → the same `madd`/`maddubs` lane mapping → **bit-identical results**. Verified by `memcmp` against the original code in a standalone harness (all of gate/up and down, tokens 1/2/4/8). - Active only for 1–2 tokens per row (`rows<1>`/`rows<2>`). At 4+ tokens the extra live unpack vectors spill registers and it measurably loses, so the original ladder stays in place there. AVX-512 machines are unaffected (they use a different path dispatch). ## Measured on my machine i7-8700K (AVX2, 6C/12T), RTX 2070 (sm_75), Qwen3.8-Flash-Next native Q2_0 pack (spec verified from the GGUF header, ~125.7B-parameter MoE), engine 0.1.27 build, temperature 0, 5 prompts × 3 warm passes per variant, and a fresh baseline rerun in the same time block to rule out machine drift. - Standalone harness (bench5, bit-exact-verified): **+10 % gate/up rows, +5–7 % down rows at 1 token/expert**; +2–4 % at 2 tokens; −5–11 % at 4 tokens (why the guard exists). - End to end: verify-window total **median ~115.7 → ~107.1 ms (−6–7 ms, ~−6 %)**, coming cleanly out of the CPU expert-pool stage (stage profiler `waitCPU`: GDN ~18.9 → ~13.1 ms, QSA ~4.4 → ~3.2 ms per window). GPU kernel stages unchanged, and generation token counts are identical across all runs (same prompts, temperature 0). - Decode: +1–3 tok/s per prompt. ## Please test on other devices I can only measure one machine, and both the gain and the end-to-end visibility depend on the platform: - **Other AVX2-only CPUs** (Intel Core 12th/13th/14th gen, Core Ultra without AVX-512 enabled, AMD Zen 2/3): the per-row gain is expected to be universal, but its end-to-end size depends on how much of decode is the CPU expert pool. - **Models with fewer experts / more active tokens** (rows serving 3+ tokens are more common there) — the `nt <= 2` guard should keep them on the original ladder; a check that nothing changes on that side is welcome. - **AVX-512 machines** should behave exactly as before (different code path) — a confirmation is also useful. Greedy outputs should be identical before/after; if you see a difference, that's a bug, please report it.
En el sitio
Enlaces a install, modelos, releases.