Pull requests / #416

#416 bench: IQ4_XS on a 64 GB PC - pinned arena vs experts.bin mmap vs the RAM budget

closed · @pipeob0 · 0 comments · View on GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentationWindows

Description

## What

A benchmark entry for the case the docs do not cover: **a model whose routed experts are bigger than
the machine's RAM**, on an AVX-2-only CPU. `orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF` (IQ4_XS)
has **60.94 GiB of routed experts on a 64 GB PC**, and the engine offers three ways to serve them.
I measured all three on the same file, plus the knobs that decide whether the model is usable at all.

Ryzen 7 5700X3D (Zen 3, no AVX-512), 64 GB DDR4-2666, RTX 5060 Ti 16 GB at PCIe 4.0 x8 (the engine's
own probe: 14.1 GB/s host→device), Windows 11, engine 0.1.31 and 0.1.32 built from source.

## Headline

| path | tok/s |
|---|---:|
| `--mmap-experts` over the pack's `experts.bin` + `--kv-resident 20480` + `STRATA_IQ_PREFETCH=16384` | **37.5 – 39.0** |
| pinned arena (the default, no `--mmap-experts`) | 33.85 warm / 27.90 cold |
| GGUF read in place + `--resident-budget-gib 40` | 26.0 – 27.5 |
| `experts.bin` + `--resident-budget-gib 40` | 24.7 – 25.4 |
| GGUF read in place, no budget | **9.7** |
| reference: the same model at IQ3_S, arena | 51.5 |

**4× between the best and the worst way to run the same file.** Four mechanisms behind the numbers,
each quoted from the engine's own log in the README:

- **The arena does not pin, and at this size the ranking flips.** The comment in `generate.cpp`
  measures the arena at 1.79× better than warm mmap, on a 34 GB `experts.bin` on 63 GB, and it already
  says the arena is not pinned there (`cudaHostRegister` on 31.64 GiB fails). The same failure here at
  60.94 GiB (`SetProcessWorkingSetSizeEx(14814 MiB) failed (error 1450)`, 37 slices, large pages
  refused too) has a different consequence: the arena pays ~30 s of load on every start and now loses
  to warm mmap. `cudaHostRegister` *succeeds* for the IQ3_S arena (`PORTABLE ok` on 50,295,996,416 B),
  so the arena is right one quantization step down and wrong one step up.
- **The RAM budget costs ~30% against plain mmap**, and both costs are visible: `adaptive tier 2688
  experts swapped into the VRAM tier (every 4 rounds, 7.580 ms/round)`, and the CPU pool reads the
  page-locked complement at **15.5 GB/s** against **27.4 GB/s** from the OS page cache.
- **A page-locked budget evicts the page cache the same model reads from.** Same arm, same arguments,
  two rounds with two 40 GiB budgeted runs in between: 38.06 tok/s / 27.4 GB/s → 18.59 / **10.6**.
- **The prefetch distance is not a lever.** Turning `STRATA_IQ_PREFETCH` off, or setting it to 8192,
  moves nothing: off / 2048 / 8192 land within 0-3% across three models and both expert paths. The
  lever is `--kv-resident` (20480 instead of the 32768 default, ~15% on this model).

Also reproduced: #403 from the other side (budget 50 refused after a run at the same budget
succeeded — 40 is the ceiling on 64 GB here), and #369 (`--expert-cache-per-layer` does not start on
a native IQ pack in mmap). `STRATA_LOOKAHEAD=0` changes nothing at this size. Context 220000 costs
2.2%.

## Method

Offline `strata generate --tokens-file <frozen prompt> --max-new 300 --stats`, arms interleaved with
**the arm order rotated one position per round**, warmup passes not counted. Compared as **ms per
draft round at equal draft-round counts**, and as **total time for the run** when the arms differ in
draft acceptance: the GPU hit path rounds differently from the CPU (the engine says so at start), so
the same arm with the same arguments produced 112–115 rounds in five runs and 127–133 in two, which
alone moves tok/s by ~10%.

Three measurement traps worth recording, all hit during this work and all fixed in the final numbers:

- the first arm of a round pays the page-cache warming on the mapped path, so later arms look better;
- one warmup pass does not warm a 61 GB file (the first measured round still read at 9.7 GB/s);
- **an arm that differs in two flags at once measures neither.** The first version of this entry
  reported the prefetch default as ~9% low; the 2048 B arm in that comparison also ran with the
  default `--kv-resident 32768`. Holding kv-resident constant, the prefetch effect disappears.

**Which version, and what has since moved.** Every number here is from 0.1.31 and 0.1.32 built from
source. 0.1.33 changed the `--resident-budget-gib` path (#403: a clamped budget now leaves a margin and
the safety check re-reads free RAM) and made `--expert-cache-per-layer` start on native packs (#369).
The budget-mode columns (24.7-27.5 tok/s) therefore describe the code as it was before that fix; the
mapped-mode numbers, which are the point of the entry, do not involve the budget path at all.

## Corrections after posting

Two claims in the first version of this entry were wrong, and the history of the branch says so:

- I wrote that the source comment does not know the arena fails to pin. It does: it reports
  `cudaHostRegister` on 31.64 GiB failing, and its 1.79x/3.7x numbers are measured at 34 GB of experts
  on 63 GB. What this entry adds is that at 61 GiB the ranking flips.
- I wrote that the prefetch default is ~9% low on the mapped path. It is not: that arm also differed
  in `--kv-resident`. Section 4 has the re-run on three models.

## Caveats

One machine, one model, offline mode (no server, no KV streaming across requests), 300 tokens per
run, greedy + `--spec 4`. Reproducible to ±1.5% at fixed draft-round counts; the cold columns depend
on what the previous run left in RAM.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.