Pull requests / #802

#802 Run IQ1_S (Unsloth UD-IQ1_S) on the GPU and AVX-2, and read the expert file tier faster

closed · @Yxmura · 0 commentaires · Sur GitHub

BenchmarksAMD / HIPModels & quantsWindowsLinux

Description

## What

- **IQ1_S on the GPU:** `dq_iq1_s`, `vec_dot_iq1_s_q8_1`, `Fmt<19>`, `Split<19>`, dispatchers, reader geometry.
- **IQ1_S + IQ1_M on AVX-2:** multi-token CPU kernels (these two formats previously fell back to ggml-cpu).
- **Tests:** `iq_parity` / `iq_multi_parity` / `native_grouped_parity` / fixture cover the new type.
- **`gfx1100-hipblaslt-100300.txt`** (calibrated on a W7800) + name the Radeon PRO W7900 / W7800.
- **Faster file tier:** parallel read-ahead for the mapped tier on every platform, and a Linux buffered-pread path behind the same `set_unbuffered` gate as Windows (mapping fallback). Windows untouched.

## Measured

Radeon PRO W7800 32 GB, 30 GiB RAM, ROCm 7.13; Unsloth UD-IQ1_S; engine 0.1.39. Without this PR the model does not load
(IQ1_S is unsupported), so the CPU rows hold the GPU support present and force the CPU experts to ggml-cpu
(`STRATA_NO_IQ256=1`) for a like-for-like "before".

| | before | after |
|---|---|---|
| UD-IQ1_S loads and answers | no (IQ1_S unsupported) | yes |
| prefill, 600-token prompt | 103 tok/s (CPU experts on ggml) | 117 tok/s (AVX-2) |
| decode, MTP | 14.5 tok/s (CPU experts on ggml) | 20.4 tok/s (AVX-2) |
| CPU gate+up, 1 expert, 1 thread — IQ1_S | 5085 us (ggml) | 664 us (7.7x) |
| CPU gate+up, 1 expert, 1 thread — IQ1_M | 6910 us (ggml) | 499 us (13.8x) |
| cold start, expert shard on FUSE-NTFS | ~193 s (mapped, serial) | ~110 s (preads) |

HIP gfx1100 build + full `strata` target; all parity suites pass; vs gguf-py 0.00e+00, vs ggml 3-4e-08.

Sur le site

Liens install, modèles, releases.