Pull requests / #1374

#1374 experts: IQ1_S on GPU/AVX-2, IQ1_S+IQ1_M on AVX-512, and a measured PCIe share

open · @Yxmura · 0 comments · View on GitHub

BenchmarksServer & APIAMD / HIPModels & quantsDocumentationWindows

Description

Round 2: IQ1_S made first-class across the pipeline, plus a PM-driven PCIe share.

## What

- **IQ1_S (type 19) on the GPU and on AVX-2.** Unsloth UD-IQ1_S packs did not run at all before: `dq_iq1_s`, `vec_dot_iq1_s_q8_1`, `Fmt<19>`, `Split<19>`, the grouped kernels, the reader geometry, and the AVX-2 `row_dot_iq1s`/`row_dot_iq1m`.
- **IQ1_S + IQ1_M on 512-bit lanes.** `iq512_supported` listed neither, and `native_expert.cpp` only takes the AVX-2 kernel off an AVX-512 CPU, so AVX-512 CPUs ran both through ggml-cpu single-token — slower than an AVX-2 CPU. No EVEX `vpsignb`, so the grid sign is `vpabsb` + a masked subtract.
- **IQ1_S on the MMQ prompt path.** IQ1_S was absent from `_strata_mmq_types`, `mmq::supported()` and the dispatch switch, so UD-IQ1_S prefill fell back to per-expert dequant→FP16. Added; the `mmq-instance-iq1_s.cu` already exists in the pinned ggml.
- **The PCIe share measures itself (serve).** `--pcie-frac` was the link probe alone (capped 0.55); the best share also depends on the CPU/GPU and runs *both* ways. A `--serve` sweeps the probed share, 0 and 1.0 over its first verify windows, keeps the lowest (CPU pool + GPU wait) per routed expert entry, and holds it. `--pcie-frac <value>` pins; `--pcie-frac auto` for the CLI.

## Measured

Radeon PRO W7800 (gfx1100, ROCm 7.13), Ryzen 7 3800X, 30 GiB RAM; UD-IQ1_S (67.5 GiB); 600-token prompt, MTP.

| | before | after |
|---|---|---|
| UD-IQ1_S loads/answers | no | yes |
| decode (this box) | 26.5 tok/s (probed share) | **40.2 → 75.3 tok/s** (measured, held) |
| prefill | 169.7 tok/s | **326.8 tok/s** (+93%: +16% the other formats, **+66% from IQ1_S MMQ alone**) |
| AVX-512 CPU, IQ1_S/IQ1_M gate/up | ggml-cpu single-token | multi-token AVX-512 |

The share runs both ways, which is why it is measured: here the GPU taking every miss is fastest, while in the [RDNA2 x8 bench](../blob/main/bench/results/2026-10-04-rdna2-helper-pcie-share/README.md) `--pcie-frac 0` beat the probed 0.39.

## Validation

- HIP gfx1100 + full `strata`; `iq_parity` (IQ1_S/IQ1_M dequant rel 0.00e+00), `iq_multi_parity`, `native_expert_parity` (IQ1_S 3.1e-08, IQ1_M 4.1e-08 vs ggml-cpu), `native_grouped_parity` — 0 failures.
- AVX-512 path under Intel SDE `-spr` (this Zen 2 cannot run it): rel ~1e-07 vs ggml-cpu, width-invariant, 0 failures.
- `hip_prefill_mmq_parity` extended with a synthetic IQ1_S case: rel_l2 0.0038 (limit 0.04), max_abs/ref_rms 0.0145, 0 zero rows.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.