Pull requests / #1415

#1415 cpu: add bit-exact AVX2 singleton expert rows

open · @InB4DevOps · 0 コメント · GitHub で見る

BenchmarksNVIDIA / CUDAModels & quantsDocumentation

本文

## Summary
- Add an opt-in AVX2 singleton gate/up path for IQ2_XXS, IQ2_XS, IQ3_XXS and IQ3_S, keeping ggml's floating-point accumulation and reduction order. The default stays unchanged.
- Check the new path bitwise against ggml in existing synthetic and real-expert parity tests.

## Measurements
- On an RTX 3060 12 GB / i7-12700KF / 62 GB PC, P-core, E-core and non-VNNI CPU parity had 0 differences across 5,120 outputs per format. Real GGUF checks had 0 differences on six tested layers (1,920 outputs per layer).
- With engine 0.1.40.2 and the IQ3_XXS pack, four unprofiled A/B pairs of a 52-token chat prompt and 512 generated tokens measured 45.28 -> 46.42 tok/s (median paired change +2.45%). Tokens and meaningful state matched in every pair. See bench/results/2026-10-07-cpu-singleton/README.md for methods and microbenchmarks.
- A separate refactoring-prompt run changed draft acceptance in one pair; no speed claim is made for it.

## Validation
- Built and ran iq_avx2_parity on P-core 2 and E-core 16, including a portable AVX2 build and a non-VNNI run.
- Ran native_expert_parity on four real IQ3_S and two real IQ3_XXS layers, and an end-to-end A/B against the same binary.

This remains opt-in; performance on other CPUs and packs is not claimed.

関連リンク

インストール・モデル・リリースへの站内リンク。