Pull requests / #802
#802 Run IQ1_S (Unsloth UD-IQ1_S) on the GPU and AVX-2, and read the expert file tier faster
closed · @Yxmura · 0 コメント · GitHub で見る
BenchmarksAMD / HIPModels & quantsWindowsLinux
本文
## What - **IQ1_S on the GPU:** `dq_iq1_s`, `vec_dot_iq1_s_q8_1`, `Fmt<19>`, `Split<19>`, dispatchers, reader geometry. - **IQ1_S + IQ1_M on AVX-2:** multi-token CPU kernels (these two formats previously fell back to ggml-cpu). - **Tests:** `iq_parity` / `iq_multi_parity` / `native_grouped_parity` / fixture cover the new type. - **`gfx1100-hipblaslt-100300.txt`** (calibrated on a W7800) + name the Radeon PRO W7900 / W7800. - **Faster file tier:** parallel read-ahead for the mapped tier on every platform, and a Linux buffered-pread path behind the same `set_unbuffered` gate as Windows (mapping fallback). Windows untouched. ## Measured Radeon PRO W7800 32 GB, 30 GiB RAM, ROCm 7.13; Unsloth UD-IQ1_S; engine 0.1.39. Without this PR the model does not load (IQ1_S is unsupported), so the CPU rows hold the GPU support present and force the CPU experts to ggml-cpu (`STRATA_NO_IQ256=1`) for a like-for-like "before". | | before | after | |---|---|---| | UD-IQ1_S loads and answers | no (IQ1_S unsupported) | yes | | prefill, 600-token prompt | 103 tok/s (CPU experts on ggml) | 117 tok/s (AVX-2) | | decode, MTP | 14.5 tok/s (CPU experts on ggml) | 20.4 tok/s (AVX-2) | | CPU gate+up, 1 expert, 1 thread — IQ1_S | 5085 us (ggml) | 664 us (7.7x) | | CPU gate+up, 1 expert, 1 thread — IQ1_M | 6910 us (ggml) | 499 us (13.8x) | | cold start, expert shard on FUSE-NTFS | ~193 s (mapped, serial) | ~110 s (preads) | HIP gfx1100 build + full `strata` target; all parity suites pass; vs gguf-py 0.00e+00, vs ggml 3-4e-08.
関連リンク
インストール・モデル・リリースへの站内リンク。