Pull requests / #207
#207 E-2 on the AVX-2 path too: prefetch the i-quant expert rows (same switch, CPUs without AVX-512)
closed · @pipeob0 · 0 comentarios · En GitHub
BenchmarksNVIDIA / CUDAModels & quantsWindows
Descripción
## What `df6980d` put the software prefetch into `iq_avx512.cpp`, so it only reaches CPUs that have AVX-512. The AVX-2 file - which is what a Zen 3, a Haswell or a Skylake-client box actually runs the i-quant expert rows on, see #103 - had none of it. This adds it there, with **the same switch, the same default and the same hint**: `STRATA_IQ_PREFETCH` is the distance in bytes, `0` = off, default `2048`. No new knob. Three sites, the three loops that stream rows in that file: - `row_dot` (the i-quant gate/up rows, `Fmt32<TY>::bytes` per block) - `row_dot_iq2xs` (the IQ2_XS gate/up rows) - `iq4nl_rows` (the IQ4_NL down rows) 21 lines added, nothing removed, no change to what is computed. ## Measured Ryzen 7 5700X3D (AVX-2 only, 8 cores), 64 GB DDR4-2666, RTX 5060 Ti, Windows. IQ3_S, 131072 context, `--pcie-frac 0.30`, `--spec 4`, 300 tokens, the same args for every run. `ms/round` is the metric, since `tok/s` moves with draft acceptance. Runs interleaved, so the thermal drift hits every variant the same way; the comparison below is **within the same round**, against the same patched binary with `STRATA_IQ_PREFETCH=0`. | | gate/up ms/round | rows GB/s | whole round ms/round | |---|---|---|---| | round 2, prefetch off | 13.571 | 25.1 | 45.54 | | round 2, prefetch 2048 | **13.024 (-4.0%)** | **26.1 (+1.0)** | **44.94 (-1.3%)** | | round 1 (pool cold), off | 13.726 | 24.4 | 46.91 | | round 1 (pool cold), 2048 | 13.487 (-1.7%) | 25.1 (+0.7) | 45.77 (-2.4%) | The unpatched binary and the patched one with the prefetch off land within 0.1% of each other in both rounds, so the patch itself costs nothing. `native_expert_parity` on the IQ3_S shard: **0 failures** with the prefetch on and with it off. The effect is bigger with a cold pool (-2.4% on the round) than warm (-1.3%), which is what you would expect from a prefetch. ## Things I checked and did not send - **The non-temporal hint is worse here**, at the same distance: -0.7% vs -1.3% on the round, +0.5 vs +1.0 GB/s. So this keeps `_MM_HINT_T0` like the AVX-512 file. - **4096 is a little worse than 1024 and 2048** (-0.7% vs -1.3%), so the default stays 2048 as in `iq_avx512.cpp`. - **The gather decode is not ported.** You measured it 3-5% slower on Zen 4; Zen 3 gathers are no better, so I left it out rather than add a second env nobody would set. - Prefetching past the end of a row is fine (the hint never faults, and the rows of one expert are contiguous in the blob); parity over the last row of each tensor confirms it. ## Caveat This is **2 rounds of 6 variants**, not the sweep you usually ask for: the sign is the same in both rounds and on all three metrics, and the control is clean, but n is 2. Happy to run more rounds on this machine if you want the number firmed up before merging - it is ~25 min per 12 runs.
En el sitio
Enlaces a install, modelos, releases.