Pull requests / #416
#416 bench: IQ4_XS on a 64 GB PC - pinned arena vs experts.bin mmap vs the RAM budget
closed · @pipeob0 · 0 评论 · 在 GitHub 查看
BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentationWindows
描述
## What A benchmark entry for the case the docs do not cover: **a model whose routed experts are bigger than the machine's RAM**, on an AVX-2-only CPU. `orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF` (IQ4_XS) has **60.94 GiB of routed experts on a 64 GB PC**, and the engine offers three ways to serve them. I measured all three on the same file, plus the knobs that decide whether the model is usable at all. Ryzen 7 5700X3D (Zen 3, no AVX-512), 64 GB DDR4-2666, RTX 5060 Ti 16 GB at PCIe 4.0 x8 (the engine's own probe: 14.1 GB/s host→device), Windows 11, engine 0.1.31 and 0.1.32 built from source. ## Headline | path | tok/s | |---|---:| | `--mmap-experts` over the pack's `experts.bin` + `--kv-resident 20480` + `STRATA_IQ_PREFETCH=16384` | **37.5 – 39.0** | | pinned arena (the default, no `--mmap-experts`) | 33.85 warm / 27.90 cold | | GGUF read in place + `--resident-budget-gib 40` | 26.0 – 27.5 | | `experts.bin` + `--resident-budget-gib 40` | 24.7 – 25.4 | | GGUF read in place, no budget | **9.7** | | reference: the same model at IQ3_S, arena | 51.5 | **4× between the best and the worst way to run the same file.** Four mechanisms behind the numbers, each quoted from the engine's own log in the README: - **The arena does not pin, and at this size the ranking flips.** The comment in `generate.cpp` measures the arena at 1.79× better than warm mmap, on a 34 GB `experts.bin` on 63 GB, and it already says the arena is not pinned there (`cudaHostRegister` on 31.64 GiB fails). The same failure here at 60.94 GiB (`SetProcessWorkingSetSizeEx(14814 MiB) failed (error 1450)`, 37 slices, large pages refused too) has a different consequence: the arena pays ~30 s of load on every start and now loses to warm mmap. `cudaHostRegister` *succeeds* for the IQ3_S arena (`PORTABLE ok` on 50,295,996,416 B), so the arena is right one quantization step down and wrong one step up. - **The RAM budget costs ~30% against plain mmap**, and both costs are visible: `adaptive tier 2688 experts swapped into the VRAM tier (every 4 rounds, 7.580 ms/round)`, and the CPU pool reads the page-locked complement at **15.5 GB/s** against **27.4 GB/s** from the OS page cache. - **A page-locked budget evicts the page cache the same model reads from.** Same arm, same arguments, two rounds with two 40 GiB budgeted runs in between: 38.06 tok/s / 27.4 GB/s → 18.59 / **10.6**. - **The prefetch distance is not a lever.** Turning `STRATA_IQ_PREFETCH` off, or setting it to 8192, moves nothing: off / 2048 / 8192 land within 0-3% across three models and both expert paths. The lever is `--kv-resident` (20480 instead of the 32768 default, ~15% on this model). Also reproduced: #403 from the other side (budget 50 refused after a run at the same budget succeeded — 40 is the ceiling on 64 GB here), and #369 (`--expert-cache-per-layer` does not start on a native IQ pack in mmap). `STRATA_LOOKAHEAD=0` changes nothing at this size. Context 220000 costs 2.2%. ## Method Offline `strata generate --tokens-file <frozen prompt> --max-new 300 --stats`, arms interleaved with **the arm order rotated one position per round**, warmup passes not counted. Compared as **ms per draft round at equal draft-round counts**, and as **total time for the run** when the arms differ in draft acceptance: the GPU hit path rounds differently from the CPU (the engine says so at start), so the same arm with the same arguments produced 112–115 rounds in five runs and 127–133 in two, which alone moves tok/s by ~10%. Three measurement traps worth recording, all hit during this work and all fixed in the final numbers: - the first arm of a round pays the page-cache warming on the mapped path, so later arms look better; - one warmup pass does not warm a 61 GB file (the first measured round still read at 9.7 GB/s); - **an arm that differs in two flags at once measures neither.** The first version of this entry reported the prefetch default as ~9% low; the 2048 B arm in that comparison also ran with the default `--kv-resident 32768`. Holding kv-resident constant, the prefetch effect disappears. **Which version, and what has since moved.** Every number here is from 0.1.31 and 0.1.32 built from source. 0.1.33 changed the `--resident-budget-gib` path (#403: a clamped budget now leaves a margin and the safety check re-reads free RAM) and made `--expert-cache-per-layer` start on native packs (#369). The budget-mode columns (24.7-27.5 tok/s) therefore describe the code as it was before that fix; the mapped-mode numbers, which are the point of the entry, do not involve the budget path at all. ## Corrections after posting Two claims in the first version of this entry were wrong, and the history of the branch says so: - I wrote that the source comment does not know the arena fails to pin. It does: it reports `cudaHostRegister` on 31.64 GiB failing, and its 1.79x/3.7x numbers are measured at 34 GB of experts on 63 GB. What this entry adds is that at 61 GiB the ranking flips. - I wrote that the prefetch default is ~9% low on the mapped path. It is not: that arm also differed in `--kv-resident`. Section 4 has the re-run on three models. ## Caveats One machine, one model, offline mode (no server, no KV streaming across requests), 300 tokens per run, greedy + `--spec 4`. Reproducible to ±1.5% at fixed draft-round counts; the cold columns depend on what the previous run left in RAM.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。