贡献 / #1701
#1701 cpu: q2_avx1 - the Q2_0 rows and the Q8_1 quantizer for CPUs with AVX but no AVX2 (fixes #1699)
open · @nashcap · 0 评论 · 去 GitHub 看
说明
## Fixes #1699 `q2_rows_any` and `act_quant_any` (expert_layout.cpp) had two rungs - AVX-512 and AVX2 - and the `else` branch checked nothing: an AVX-only CPU (Sandy Bridge through Ivy Bridge Xeons, the #623/#1375 class) was handed AVX2 code and died with SIGILL. On a clean v0.1.41 tree an E5-2470 v2 kills both `pool_tasks_test` and `q2_bitplane_parity` at the first frame (`act_quant_q8_1_avx2+123`), and any future caller of the dispatchers inherits the trap. ## The fix The third rung, `src/kernels/cpu/q2_avx1.cpp`, in the containment pattern the repo already uses for `kq_avx1.cpp`/`iq_avx1.cpp`: its own translation unit compiled `-msse4.2 -mavx -mno-avx2 -mno-fma -mno-f16c` (a `#error` guard keeps AVX2/FMA/F16C out of the object even on a `-march=native` host), called only behind `cpu_avx1_ok() && !cpu_avx2_ok()`. The AVX-2 and AVX-512 paths are untouched, so nothing changes on any CPU that has AVX2. - 128-bit integer lanes (`maddubs`/`madd` are SSSE3/SSE2), 256-bit float accumulation, software fp16 scale conversion (no F16C). - The legacy row kernel keeps the AVX-2 kernel's lane pairing and reduction tree exactly; the only difference is its FMA becoming a mul+add (last bits, unavoidable without FMA). - The bit-plane kernel and its image builder port as-is (the image builder is already pure SSE2/SSSE3). - `q2_bitplane_parity` now tests the AVX-2 pair where the CPU has AVX2 and the AVX1 pair below it, and skips cleanly (pass) without AVX. ## Measured on the hardware from the issue E5-2470 v2 (Ivy Bridge, AVX only), v0.1.41 + this commit: - `pool_tasks_test`: 168 bitwise comparisons passed (SIGILL before) - `q2_bitplane_parity`: PASS (SIGILL before) - `iq_multi_parity`, `router_dot_parity`, `dequant_s2_parity`, `expert_layout_test`, `q8k_quant_parity`: unchanged, pass Independent of #1697 (the i-quant AVX1 kernels); either can merge first.
本站相关内容
相关页面的快捷入口。