Pull requests / #1237
#1237 perf(expert): the cache fill's stage buffers pinned with cudaHostAlloc
open · @ZhongUncle · 0 comments · View on GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quants
Description
`blob()` serves experts from the stage buffers whenever the resident expert arena is opted out of (`--mmap-experts`) or too small for the expert set. The stage buffers are the pool that assembles each blob from the GGUF's three role slices. They are allocated with a bare `new[]`, so they are pageable. A `cudaMemcpyAsync` from pageable memory first copies the data into a driver-internal bounce buffer, on the calling thread, before the DMA can start. Every consumer pays this per expert: the profile prefill's blocking fill ([`generate.cpp:4114`](https://github.com/Niko1221/Strata/blob/82f46a8c8f475f001ad76d92f58f4a4f8ffb0253/src/program/generate.cpp#L4114)), the lent-slot refills, the adaptive tier's swaps, and the admit path where it runs. This PR allocates the stage buffers with `cudaHostAlloc` instead, falling back to pageable per buffer when the driver refuses (pinned memory is scarce on exactly the small-RAM machines this path serves). Nothing else changes: same pool, same claim/reuse policy, same bytes copied — a buffer whose only job is to be the source of an async H2D copy should be pinned. **TL;DR** — pinning the stage buffers makes every staged H2D copy cheaper. The cost is ≈0.3 s slower startup (profile prefill ≈7%; ≈0.2 s of it measured pinning cost, the rest unexplained) and ~386 MiB of pinned RAM. It only matters when `blob()` serves from the stage buffers; otherwise they sit idle. Decode tok/s is roughly flat. - micro-benchmark, one 1.5 MiB blob: the caller blocks 109–167 µs inside `cudaMemcpyAsync` from pageable vs 2–3 µs from pinned; end-to-end 219–256 → 144–145 µs (1.5–1.8×), the pinned copy at 10.8 of the card's probed 11.1 GB/s - engine, `--mmap-experts` in both arms: the adaptive tier's swap round 4.20–4.30 → 0.90–1.18 ms (**3.6–4.8×**, four interleaved pairs); the lent-slot refill 165–169 → 104–105 ms for 530 slots (**1.6×**) - the profile prefill loses ≈7% (four-run means 3277 → 3055 MB/s): ≈0.21 s of it is the measured pinning of the pool's ~257 buffers, the rest unexplained (Findings) - scope: on a machine whose RAM holds the whole expert set, the pinned complement (a host mirror of the experts the GPU cache doesn't hold) serves `blob()` directly and the stage buffers sit idle — this PR is for the configs where they don't ### What changes - `claim_stage`: `new uint8_t[stage_blob_]` → `cudaHostAlloc(..., cudaHostAllocDefault)`, with a per-buffer pageable fallback - the pool's `unique_ptr` gains a small deleter that frees each buffer the way it was allocated - two one-time stderr notes make the state observable in the log: the first pinned buffer; the first fallback, with the driver's error string - `cudaHostAllocDefault`, not `Mapped`: a stage buffer is only ever the *source* of an H2D copy, never addressed by the device - allocation stays under `stage_mu_` exactly where the `new[]` was; a failed `cudaHostAlloc`'s sticky error is cleared on the spot - pool growth is concentrated in the first distinct-expert sweep — the profile prefill grows the pool to its ~257-buffer steady state, and afterwards buffers are only recycled, never allocated — so no `cudaHostAlloc` serialization reaches decode - the buffers are freed through the same `close()` path that already `cudaFreeHost`s the complement and exchange buffers — no new lifetime questions ### Measured results Measured on: - RTX 3080 20 GB (driver 595.91.07, CUDA 13.2) - Xeon E5-2680 v4 (no AVX512) - 62 GB RAM - Ubuntu 24.04 - Qwen3.8-Flash-Next IQ2_XS pack (48 layers × 512 experts, largest blob 1.51 MB) - `--mmap-experts` in **both** arms (see the scope note), `--spec 4 --spec-min-p 0.5` + the MTP drafter **Micro-benchmark — per-copy cost of one 1.5 MiB blob, 64-buffer rotating ring (no buffer is re-read within 64 copies), 512 copies per arm per variant, quiet machine:** | source | variant | host µs inside `cudaMemcpyAsync` | end-to-end µs (GPU events) | |---|---|---|---| | pageable `new[]` | copy-only | 166.9 | 256.3 | | pinned | copy-only | 2.0 | 145.2 | | pageable `new[]` | fill-then-copy | 108.7 | 218.9 | | pinned | fill-then-copy | 3.2 | 144.4 | - fill-then-copy memcpys a fresh blob into the buffer right before its copy — the engine's pattern, where `fill_stage` assembles the blob and the cache fill copies it right away - the pageable arm's host time is lower there than in copy-only, likely because the just-written blob is still cache-resident when the driver bounce-copies it - pinned end-to-end is the copy engine at 10.8 GB/s against the card's probed 11.1 GB/s; pageable end-to-end is the bounce staging plus the DMA, serialized **Engine A/B — four interleaved baseline/patched pairs, 512 generated tokens after an 873-token prompt (raw lines in the appendix):** | metric (per pair) | baseline | patched | × | |---|---|---|---| | adaptive tier, ms per swap round | 4.20 / 4.26 / 4.30 / 4.27 | 1.18 / 1.17 / 0.90 / 1.17 | **3.6 / 3.6 / 4.8 / 3.6** | | lent-slot refill, ms (530 slots) | 169.2 / 165.1 / 164.6 / 166.0 | 104.6 / 104.7 / 104.6 / 104.1 | **1.62 / 1.58 / 1.57 / 1.59** | | profile prefill, MB/s (10,244 slots) | 3294 / 3262 / 3252 / 3300 | 3038 / 3040 / 3039 / 3103 | 0.92 / 0.93 / 0.93 / 0.94 | | decode, tok/s | 36.00 / 40.03 / 39.95 / 40.35 | 39.75 / 41.01 / 39.11 / 40.83 | roughly flat | ### Findings 1. **The swap round's host time drops 3.6–4.8×.** The adaptive tier's round loop ([`generate.cpp:10610`](https://github.com/Niko1221/Strata/blob/82f46a8c8f475f001ad76d92f58f4a4f8ffb0253/src/program/generate.cpp#L10610)) issues one async H2D per swapped expert from `blob()`'s stage buffer, on its own thread; the logged ms/round times that loop. From pageable memory each issue costs the ~110–170 µs bounce copy measured in the micro-benchmark, from pinned memory ~3 µs. Decode tok/s does not move (finding 4): at this swap cadence (every 4 rounds) the copies are already off the critical path. The gain is headroom on the swap thread, which should matter at a higher cadence or on a slower CPU; the copies are also expected to overlap the MTP drafting that follows (the loop's own comment, [`generate.cpp:10615`](https://github.com/Niko1221/Strata/blob/82f46a8c8f475f001ad76d92f58f4a4f8ffb0253/src/program/generate.cpp#L10615)). 2. **The lent-slot refill gains 1.6×, matching the micro-benchmark's end-to-end ratio.** The refill ([`generate.cpp:10180`](https://github.com/Niko1221/Strata/blob/82f46a8c8f475f001ad76d92f58f4a4f8ffb0253/src/program/generate.cpp#L10180)) queues 530 staged copies and waits once: pageable pays ~117 µs more per slot than pinned (313 → 197 µs), the bounce-staging term. 3. **The profile prefill's ≈7% loss is mostly the pool's one-time pinning cost, not a per-copy regression.** The fill loop is serial — `blob()` (three memcpys out of the mapped GGUF) then a *blocking* `cudaMemcpy`. The sweep of 10,244 distinct experts grows the stage pool to its ~257-buffer steady state before recycling kicks in; a buffer becomes reusable 256 assemblies after its last use. `cudaHostAlloc` of 1.5 MiB measures 831 µs here vs 0.2 µs for `new[]`, so pinning the ~257 buffers accounts for ≈0.21 s — about two-thirds of the measured mean slowdown (4.5 → 4.8 s, +0.3 s); the remaining ~0.1 s is unexplained. Per-copy costs are otherwise strictly lower when pinned (findings 1–2), the fill loop is dominated by the mmap read (~440 µs per slot), and the loss is paid once at startup. 4. **Decode is roughly flat; the trajectory split in the logs is unresolved.** Excluding the cold first run, decode differs by at most ±1 tok/s between arms, within the run-to-run spread. The lookup counts split into groups that mostly follow the arm (the forensics are in the appendix), and the swap totals differ with them, so the patch may have shifted this engine's timing-dependent expert refusals; but the cold baseline run shares the patched group's lookup total, which points to an independent rig state. These runs cannot settle it. Either way, ms/round is normalized per round, so findings 1–2 are unaffected. ### Validation - **What was compared:** baseline (this PR's parent, 82f46a8c) vs patched binaries on four interleaved pairs of identical `--mmap-experts` runs; the micro-benchmark isolates the mechanism - **Pinning confirmed in the log:** every patched run prints `FileExpertSource: the stage buffers are pinned (cudaHostAlloc)` exactly once; no baseline run does - **The engine's own fill check:** `slot 0 verified` passes in every run (the prefill reads one slot back and compares bytes) - **Full suite:** `ctest` 87/89 — `ple_parity` and `expert_multi_test` fail identically on the base commit on this host (the first needs the Q2_0 pack's GGUF, not present here; the second needs AVX512, which this Xeon lacks) ### Limits - **Scope:** with the default config on this rig the 33 GiB complement is fully pinned and serves every `blob()` — the stage buffers are idle, and nothing measurable changes. The measured `--mmap-experts` config stands in for any machine whose arena is smaller than its expert set. - **Pinned footprint:** the pool's steady state here is ~257 buffers × 1.5 MiB ≈ 386 MiB of non-pageable RAM — real money on the small-RAM machines this path serves. A driver refusal downgrades only that one buffer to pageable, so under pin pressure the pool runs mixed: slower copies for those buffers, nothing breaks. - The prefill's one-time pinning cost scales with the pool's growth (~257 buffers here), not with the model. With `--prefill` off, that growth — and its ~0.2 s — moves out of startup and into the first stage-buffer users: the lent-slot refill alone queues 530 slots, more than the pool's steady state, so the refill would pay it inside its own ~0.1 s window. That configuration was not measured. - **The decode-time admit path ([`expert_source.cpp:2721`](https://github.com/Niko1221/Strata/blob/82f46a8c8f475f001ad76d92f58f4a4f8ffb0253/src/core/expert_source.cpp#L2721)) is not measurable in this pack's `generate` configs:** `--spec` requires the token graph, and the graph decides hits on the device with residency frozen per token — "nothing is admitted during a token" ([`expert_source.cpp:2391`](https://github.com/Niko1221/Strata/blob/82f46a8c8f475f001ad76d92f58f4a4f8ffb0253/src/core/expert_source.cpp#L2391); all four runs logged `0 admitted`). That path consumes the same buffers through the same `fill_slot`, so it stands to gain what findings 1–2 measure; it is cited by code inspection, not measured here. - One model, one quant, one GPU. ### Reproduce ```bash # engine A/B (host with the model; pretokenize via the pack's chat template) ./engine/strata --pack Strata-data/packs/iq2_xs --native <shard1.gguf> --ple-gguf <shard2.gguf> \ --expert-profile data/expert-profile.bin --expert-cache auto --prefill auto \ --spec 4 --spec-min-p 0.5 --mtp Strata-data/mtp/rt \ --max-context 24576 --kv int8 --kv-resident 32768 --mmap-experts \ --tokens-file PROMPT.tokens --max-new 512 --stats # the "pre-filled" and "lent slots refilled" startup lines and the "adaptive tier" stat line # carry the figures; interleave baseline and patched runs ``` The micro-benchmark is **not** part of this PR; its parameters and build line are in the appendix, and its ~100-line source can be posted as a comment here if anyone wants to rerun it. <details> <summary>Raw `--stats` lines and the micro-benchmark</summary> Engine A/B, all eight runs — 873-token task prompt, 512 generated tokens, interleaved baseline/patched pairs; both arms `--mmap-experts` (supports findings 1–4): ``` baseline r1: pre-filled 10244/10244 in 4.5 s (3294 MB/s) — 530 lent slots in 169.2 ms — adaptive tier 4185 swaps, 4.198 ms/round — decode 36.00 tok/s (14223.8 ms) — hits 293709/326400, 0 admitted patched r1: pre-filled 10244/10244 in 4.9 s (3038 MB/s) — 530 lent slots in 104.6 ms — adaptive tier 4078 swaps, 1.177 ms/round — decode 39.75 tok/s (12879.4 ms) — hits 293322/326400, 0 admitted baseline r2: pre-filled 10244/10244 in 4.5 s (3262 MB/s) — 530 lent slots in 16
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.