Issues / #1165

#1165 Ship an MMQ-enabled engine build (or opt-in loading) for k-quant prompts: a full-RAM machine is GPU-bound on the FP16 dequant path

closed · @bys13 · 0 comentários · No GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

Descrição

**Summary.** On a machine where the UD-IQ4_XS experts are fully resident (no SSD anywhere in the
   decode or prompt path), prompt reading is limited by the FP16-dequant expert products the released
   engine uses for Q4_K / Q5_1. docs/UNSLOTH_Q4.md (lines 173-177) says the MMQ prompt kernels are off
   in the released engine because they cost start-up time and 1-3 expert slots of VRAM — but the
   measured "no difference" behind that decision comes from a 64 GB machine where the prompt waits for
   the SSD either way. On a full-RAM machine there is no SSD to wait for: the FP16 path *is* the
   bottleneck, and the 4x GPU-time saving would land in full. I'd like a way to pay for it voluntarily.

   ## Environment

   - **Engine:** 0.1.40 (`engine/BUILD.json`), Python server 0.1.40.1, Windows 11, driver 610.88
   - **GPU:** RTX 4090 48 GB (modded), CUDA 13.0 build
   - **CPU / RAM:** i9-12900KF (AVX2, no AVX-512), 128 GB DDR4
   - **Model:** `unsloth/Qwen3.8-Flash-Next-GGUF`, **UD-IQ4_XS**, 3 shards, vision on, 256K context

   ## Configuration (abridged)

   ```jsonc
   "args": [
     "--pack", "D:\\Strata-data\\packs\\unsloth-ud-iq4_xs",
     "--native", "D:\\Strata-data\\models\\unsloth-UD-IQ4_XS\\...-00001-of-00003.gguf",
     "--expert-profile", "data\\expert-profile.bin", "--expert-cache", "auto",
     "--prefill", "auto", "--mtp", "D:\\Strata-data\\mtp\\rt", "--spec", "4", "--spec-min-p", "0.5",
     "--max-context", "262144", "--kv", "int8", "--kv-resident", "32768",
     "--resident-budget-gib", "55", "--ple-io", "mmap", "--vram-reserve-mib", "2048"
   ]
   ```

   The engine log confirms full residency, no file reads on the hot path:

   ```
   expert cache 17126 slots, 38.61 GiB of VRAM
   resident RAM mode: 16.82 GiB of experts in RAM (page-locked), 17126 in the GPU cache;
   adaptive swaps exchange them with the GPU cache (no file reads)
   ```

   ## Measured (this machine)

   | Run | Prompt | Result |
   | --- | --- | --- |
   | Fresh prefill, `--ple-io direct` | 27,889 tokens, 0 reused | **862.6 tok/s** |
   | Fresh prefill, `--ple-io mmap` (after warm-up) | same 27,889 tokens, 0 reused | **1,168.8 tok/s** |
   | Decode (reasoning off, MTP accepted) | — | ~85 tok/s |
   | **IQ3_S (GSQ-RCO) on the same machine, same context** | 211,567 tokens, 0 reused | **3,814 tok/s** |

   The IQ3_S comparison is the tell: same architecture, same PLE table, same machine, same
   `--kv`/`--spec`/`--mtp` settings — 3.3x faster prompt reading, because the i-quant formats have the
   native MMVQ prompt path while the k-quant experts of UD-IQ4_XS fall back to dequantizing to FP16.
   Disk during prefill drops from ~380 MB/s (file-cache warm-up) to near zero once warm, so nothing
   here is waiting for the drive.

   ## What I'd like

   Any of these, in order of preference:

   1. A release engine build with `-DSTRATA_MMQ_KQUANTS=ON` offered alongside the current one (or a
      separate `engine-mmq` download), so full-RAM k-quant users can opt in; or
   2. Loading the k-quant MMQ kernels lazily (on the first k-quant prompt) instead of at start-up, so
      non-k-quant models never pay the start-up or VRAM cost; or
   3. A documented note in UNSLOTH_Q4.md telling 128 GB-class users that building with the flag is the
      expected way to run this pack, with the measured delta.

   The 1-3 expert-slot VRAM cost and the kernel-load start-up cost are fine to eat when your prompt
   path is the long pole — they should just not be charged to users whose prompts are SSD-bound.

   Happy to benchmark any test build on this machine (fresh 28K-token prefill, warm and cold file
   cache, decode, and the IQ3_S pair for reference).

No site

Links install, modelos, releases.