Issues / #1165
#1165 Ship an MMQ-enabled engine build (or opt-in loading) for k-quant prompts: a full-RAM machine is GPU-bound on the FP16 dequant path
closed · @bys13 · 0 评论 · 在 GitHub 查看
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
描述
**Summary.** On a machine where the UD-IQ4_XS experts are fully resident (no SSD anywhere in the
decode or prompt path), prompt reading is limited by the FP16-dequant expert products the released
engine uses for Q4_K / Q5_1. docs/UNSLOTH_Q4.md (lines 173-177) says the MMQ prompt kernels are off
in the released engine because they cost start-up time and 1-3 expert slots of VRAM — but the
measured "no difference" behind that decision comes from a 64 GB machine where the prompt waits for
the SSD either way. On a full-RAM machine there is no SSD to wait for: the FP16 path *is* the
bottleneck, and the 4x GPU-time saving would land in full. I'd like a way to pay for it voluntarily.
## Environment
- **Engine:** 0.1.40 (`engine/BUILD.json`), Python server 0.1.40.1, Windows 11, driver 610.88
- **GPU:** RTX 4090 48 GB (modded), CUDA 13.0 build
- **CPU / RAM:** i9-12900KF (AVX2, no AVX-512), 128 GB DDR4
- **Model:** `unsloth/Qwen3.8-Flash-Next-GGUF`, **UD-IQ4_XS**, 3 shards, vision on, 256K context
## Configuration (abridged)
```jsonc
"args": [
"--pack", "D:\\Strata-data\\packs\\unsloth-ud-iq4_xs",
"--native", "D:\\Strata-data\\models\\unsloth-UD-IQ4_XS\\...-00001-of-00003.gguf",
"--expert-profile", "data\\expert-profile.bin", "--expert-cache", "auto",
"--prefill", "auto", "--mtp", "D:\\Strata-data\\mtp\\rt", "--spec", "4", "--spec-min-p", "0.5",
"--max-context", "262144", "--kv", "int8", "--kv-resident", "32768",
"--resident-budget-gib", "55", "--ple-io", "mmap", "--vram-reserve-mib", "2048"
]
```
The engine log confirms full residency, no file reads on the hot path:
```
expert cache 17126 slots, 38.61 GiB of VRAM
resident RAM mode: 16.82 GiB of experts in RAM (page-locked), 17126 in the GPU cache;
adaptive swaps exchange them with the GPU cache (no file reads)
```
## Measured (this machine)
| Run | Prompt | Result |
| --- | --- | --- |
| Fresh prefill, `--ple-io direct` | 27,889 tokens, 0 reused | **862.6 tok/s** |
| Fresh prefill, `--ple-io mmap` (after warm-up) | same 27,889 tokens, 0 reused | **1,168.8 tok/s** |
| Decode (reasoning off, MTP accepted) | — | ~85 tok/s |
| **IQ3_S (GSQ-RCO) on the same machine, same context** | 211,567 tokens, 0 reused | **3,814 tok/s** |
The IQ3_S comparison is the tell: same architecture, same PLE table, same machine, same
`--kv`/`--spec`/`--mtp` settings — 3.3x faster prompt reading, because the i-quant formats have the
native MMVQ prompt path while the k-quant experts of UD-IQ4_XS fall back to dequantizing to FP16.
Disk during prefill drops from ~380 MB/s (file-cache warm-up) to near zero once warm, so nothing
here is waiting for the drive.
## What I'd like
Any of these, in order of preference:
1. A release engine build with `-DSTRATA_MMQ_KQUANTS=ON` offered alongside the current one (or a
separate `engine-mmq` download), so full-RAM k-quant users can opt in; or
2. Loading the k-quant MMQ kernels lazily (on the first k-quant prompt) instead of at start-up, so
non-k-quant models never pay the start-up or VRAM cost; or
3. A documented note in UNSLOTH_Q4.md telling 128 GB-class users that building with the flag is the
expected way to run this pack, with the measured delta.
The 1-3 expert-slot VRAM cost and the kernel-load start-up cost are fine to eat when your prompt
path is the long pole — they should just not be charged to users whose prompts are SSD-bound.
Happy to benchmark any test build on this machine (fresh 28K-token prefill, warm and cold file
cache, decode, and the IQ3_S pair for reference).站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。