Pull requests / #207

#207 E-2 on the AVX-2 path too: prefetch the i-quant expert rows (same switch, CPUs without AVX-512)

closed · @pipeob0 · 0 comentários · No GitHub

BenchmarksNVIDIA / CUDAModels & quantsWindows

Descrição

## What

`df6980d` put the software prefetch into `iq_avx512.cpp`, so it only reaches CPUs that have AVX-512. The AVX-2 file - which is what a Zen 3, a Haswell or a Skylake-client box actually runs the i-quant expert rows on, see #103 - had none of it. This adds it there, with **the same switch, the same default and the same hint**: `STRATA_IQ_PREFETCH` is the distance in bytes, `0` = off, default `2048`. No new knob.

Three sites, the three loops that stream rows in that file:

- `row_dot` (the i-quant gate/up rows, `Fmt32<TY>::bytes` per block)
- `row_dot_iq2xs` (the IQ2_XS gate/up rows)
- `iq4nl_rows` (the IQ4_NL down rows)

21 lines added, nothing removed, no change to what is computed.

## Measured

Ryzen 7 5700X3D (AVX-2 only, 8 cores), 64 GB DDR4-2666, RTX 5060 Ti, Windows. IQ3_S, 131072 context, `--pcie-frac 0.30`, `--spec 4`, 300 tokens, the same args for every run. `ms/round` is the metric, since `tok/s` moves with draft acceptance. Runs interleaved, so the thermal drift hits every variant the same way; the comparison below is **within the same round**, against the same patched binary with `STRATA_IQ_PREFETCH=0`.

| | gate/up ms/round | rows GB/s | whole round ms/round |
|---|---|---|---|
| round 2, prefetch off | 13.571 | 25.1 | 45.54 |
| round 2, prefetch 2048 | **13.024 (-4.0%)** | **26.1 (+1.0)** | **44.94 (-1.3%)** |
| round 1 (pool cold), off | 13.726 | 24.4 | 46.91 |
| round 1 (pool cold), 2048 | 13.487 (-1.7%) | 25.1 (+0.7) | 45.77 (-2.4%) |

The unpatched binary and the patched one with the prefetch off land within 0.1% of each other in both rounds, so the patch itself costs nothing. `native_expert_parity` on the IQ3_S shard: **0 failures** with the prefetch on and with it off.

The effect is bigger with a cold pool (-2.4% on the round) than warm (-1.3%), which is what you would expect from a prefetch.

## Things I checked and did not send

- **The non-temporal hint is worse here**, at the same distance: -0.7% vs -1.3% on the round, +0.5 vs +1.0 GB/s. So this keeps `_MM_HINT_T0` like the AVX-512 file.
- **4096 is a little worse than 1024 and 2048** (-0.7% vs -1.3%), so the default stays 2048 as in `iq_avx512.cpp`.
- **The gather decode is not ported.** You measured it 3-5% slower on Zen 4; Zen 3 gathers are no better, so I left it out rather than add a second env nobody would set.
- Prefetching past the end of a row is fine (the hint never faults, and the rows of one expert are contiguous in the blob); parity over the last row of each tensor confirms it.

## Caveat

This is **2 rounds of 6 variants**, not the sweep you usually ask for: the sign is the same in both rounds and on all three metrics, and the control is clean, but n is 2. Happy to run more rounds on this machine if you want the number firmed up before merging - it is ~25 min per 12 runs.

No site

Links install, modelos, releases.