Pull requests / #1415
#1415 cpu: add bit-exact AVX2 singleton expert rows
open · @InB4DevOps · 0 comentarios · En GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentation
Descripción
## Summary - Add an opt-in AVX2 singleton gate/up path for IQ2_XXS, IQ2_XS, IQ3_XXS and IQ3_S, keeping ggml's floating-point accumulation and reduction order. The default stays unchanged. - Check the new path bitwise against ggml in existing synthetic and real-expert parity tests. ## Measurements - On an RTX 3060 12 GB / i7-12700KF / 62 GB PC, P-core, E-core and non-VNNI CPU parity had 0 differences across 5,120 outputs per format. Real GGUF checks had 0 differences on six tested layers (1,920 outputs per layer). - With engine 0.1.40.2 and the IQ3_XXS pack, four unprofiled A/B pairs of a 52-token chat prompt and 512 generated tokens measured 45.28 -> 46.42 tok/s (median paired change +2.45%). Tokens and meaningful state matched in every pair. See bench/results/2026-10-07-cpu-singleton/README.md for methods and microbenchmarks. - A separate refactoring-prompt run changed draft acceptance in one pair; no speed claim is made for it. ## Validation - Built and ran iq_avx2_parity on P-core 2 and E-core 16, including a portable AVX2 build and a non-VNNI run. - Ran native_expert_parity on four real IQ3_S and two real IQ3_XXS layers, and an end-to-end A/B against the same binary. This remains opt-in; performance on other CPUs and packs is not claimed.
En el sitio
Enlaces a install, modelos, releases.