Pull requests / #1544

#1544 cpu: IQ2_S in the bit-exact singleton rows, and the singleton path only below the #152 rule (stacked on #1415)

open · @stuchapin909 · 0 comentarios · En GitHub

BenchmarksNVIDIA / CUDAModels & quantsWindowsLinux

Descripción

## Summary
Stacked on #1415 (InB4DevOps), one commit on top of 1ba010e. Two changes, both behind #1415's opt-in `STRATA_CPU_IQ_ALL_EXACT1=1`; with the switch unset nothing changes.

1. **IQ2_S (type 22) joins the singleton list.** ggml's AVX2 `ggml_vec_dot_iq2_s_q8_K` sums the same int32 lanes as `row_dot_ggml_one` (one scale per 16 values, even and odd 32-value groups in two accumulators, one FMA per block) and ends with `0.125f * hsum_float_8`, so the existing routine gives its bits once 22/122 take the 0.125 return. IQ2_S is the gate/up format of 34 of 48 layers in Swift IQ2_XS and 20 of 48 in the GSQ-RCO IQ3_S pack, the largest share in both.
2. **The singleton path only where ggml's dot runs today** (`nt < native_gu_mt_min`). In #1415 the gate comes before the #152 rule, so with `STRATA_IQ_MT_MIN=1` a lone token took ggml's rounding while a group kept the multi-token kernel's, and `native_expert_parity --synthetic` failed its width-invariance check with the switch on (reported in #1415). With this, `STRATA_IQ_MT_MIN=1` behaves as without the switch.

## What changed
- `src/kernels/cpu/iq_avx2_rows.inl`: `row_dot_ggml_one` returns `0.125f * ggml_hsum8` for 22 / 122 too.
- `src/kernels/cpu/iq_avx2.cpp`: `iq256_gu_rows_exact_one` dispatches 22 (122 with the gathered decode), plain and AVX-VNNI.
- `src/kernels/cpu/native_expert.cpp`: type 22 in the gate, and `nt < native_gu_mt_min(f.gu_type)`.
- `src/kernels/cpu/iq_avx2_parity.cpp`, `src/kernels/native_expert_parity.cpp`: IQ2_S in the bit-for-bit checks.

## Measured
i9-9920X (Skylake-X, 12 cores; `cpu_avx512_ok()` false, so the AVX-2 kernels, scalar decode, no AVX-VNNI), 96 GB DDR4, Windows 11, one RTX 3090. Built with `STRATA_PORTABLE=ON` (ggml-cpu `/arch:AVX2`, as the release).

Parity:
- `iq_avx2_parity`: exact singleton vs ggml 0 of 5,120 for IQ2_XXS, IQ2_XS, **IQ2_S**, IQ3_XXS and IQ3_S; IQ2_S also 0 of 5,120 with `STRATA_IQ256_GATHER=1`. 0 failures.
- `native_expert_parity --synthetic` for `iq3_s/iq4_nl`, `iq2_s/q2_0`, `iq3_xxs/iq4_nl`, `iq2_xs/iq4_nl`, `q4_K/q5_1`, each with `STRATA_CPU_IQ_ALL_EXACT1` 0 and 1: exit 0, width invariance 0 rows differ, exact singleton 0 of 1,920 (with #1415 alone the `=1` runs of the three i-quant pairs exited 1, gate/up 1,028 rows).

Decode: Swift IQ2_XS, expert cache sized like a 12 GB card (`--vram-reserve-mib 12300`), `--spec 4 --spec-min-p 0.5 --kv k8v4 --kv-resident 32768 --max-context 262144 --pcie-frac 0.25`, one prompt ("Explain how to compute 17*19 + 23*29 by hand, then give the integer."), greedy, 256 tokens, `strata generate --tokens-file`, a fresh process per run. `#1415` = the PR head; `+ this` = this commit; both with the switch set:

| | tok/s | pool gate/up ms per window |
| --- | --- | --- |
| switch unset | 69.14 | 23.65 |
| #1415 | 71.17, 68.84 | 23.53, 23.70 |
| #1415 + this | 73.03, 71.32, 72.76 | 21.68, 21.64, 21.03 |

The pool's gate/up time falls ~2 ms a window (-8.5%) against #1415; tok/s +3-4%. The same tokens in all six runs. The last row's third run is with the `nt < mt_min` change, the first two without it.

## Extra Notes
- Not measured on AVX-VNNI, the gathered decode in the engine (only in the parity test), Zen 2/3, or Linux.
- The `!cpu_avx512_ok()` exclusion from #1415 is kept as it is.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

En el sitio

Enlaces a install, modelos, releases.