Pull requests / #1190
#1190 resident RAM mode with a layer split: the copy keeps each stage's lend region
open · @evaanp · 0 コメント · GitHub で見る
Server & APINVIDIA / CUDAModels & quants
本文
## The problem #848 gave `--resident-experts` with a layer split a RAM copy of only the experts **no stage's cache holds**. That left out the experts the prompt path borrows: `generate.cpp` hands `pin_cache_complement` the later stages' cache pairs as `additional_gpu_pairs`, and with that list non-empty the function switches the single lend region off (`expert_source.cpp:1752`, `lend = ... && additional_gpu_pairs.empty()` — the comment at `generate.cpp:5820`: "with them the copy keeps no lend region: pin_cache_complement turns the loan off"). So under a split, the experts in the tail slots of **every** card's cache — CUDA0's part included — are in neither VRAM nor the RAM copy during a prompt, and the prompt reads them from the model files through page faults. Measured 2026-10-06 (v0.1.40 build, config: 17/48 layer split, RTX 2000 Ada stage 0 = layers 0-16 + RTX 5060 Ti stage 1 = 17-47, IQ3_S native pack without experts.bin so `--mmap-experts` maps the GGUF shards, `--prefill auto` = 8192-token chunks; each card's prompt path borrows 4.31 GiB of its expert-cache slots): | | arena (default) | `--resident-experts` (before) | |---|---|---| | RAM pinned for experts | ~47 GiB (all) | 25.13 GiB (complement only) | | available RAM | 3.5 GB | 24.8 GB | | decode | 62.8 t/s | 62.3 t/s | | prefill 73K cold | 2339 t/s | 1892 t/s (-19%) | | short prompts (70 tokens, cold) | ~90-110 t/s | ~50-65 t/s | | NVMe read during one 73K prefill | ~1.1 GB | **~3.7 GB** | `STRATA_PREFILL_TIMING=1`, one 8192-token chunk (GPU wall 3132 ms arena vs 3865 ms resident): dequant 232 vs 678 ms, host grouping 33 vs 187 ms, host staging 365 vs 1504 ms. The expensive experts were the file-read ones: with the borrowed slots' bytes already in RAM, grouping/staging/dequant all stream the same bytes the arena streams, and the walk collapses to `memcpy` + `cudaMemcpy` instead of a per-blob page-fault trail through the mmap'd GGUF. (`--no-prefill-borrow` restored the short prompts but cost the 73K prefill — 1428 t/s — and decode: not the fix.) ## The change - `FileExpertSource::stage_lend_regions(...)` (header): each stage's (layer, expert) pairs, listed from the cache's highest slot down — the order the borrowing takes them. - `detail::append_stage_lend_regions` (`expert_source.cpp:186`): adds the regions to the compact plan after the complement; a pair already in the copy is passed over, a stage's walk stops at the first pair that does not fit, and the next stage's walk still runs. One budget caps the whole copy — the same available-minus-headroom reading as the complement (#403), or an explicit `--resident-budget-gib`. Experts past it keep the mapped-file fallback they have today. - `pin_cache_complement` folds the regions in after its own lend walk (`expert_source.cpp:1833`) and says how many of the split's lent slots stayed in RAM (`the split stages lend N slots to the prompt path: M keep their experts in RAM too`). - `generate.cpp:5827`: sizes each part's region at the largest chunk the serve scan can ever pick — auto's ceiling `auto_ceiling`, or an explicit `--prefill`, which the scan only walks down — with the scan's own tail-slot arithmetic (`cache_slots_for`, the 128-slot floor, the auto `kAutoLendPct` cap), and honours #340's `own_ok` rule: a part whose card can fund the bound chunk's buffers outright gets no region and no RAM is spent on it. RAM arithmetic on the rig above: 25.13 GiB (complement) + ~8.6 GiB (two regions of 4.31 GiB) ≈ 33.7 GiB — ~13.5 GiB less pinned than the arena, with none of the 2.6 GiB extra NVMe reads the one prompt cost. Answers are unchanged: the copy holds the same file bytes the fallback would read. Untouched: the arena default, single-GPU behaviour (no regions set; `pin_cache_complement`'s own lend walk stays the only one), disk/keyless parking, and the expert-order/lookup semantics — resident pairs resolve through the existing `cache_complement_blob_or_fallback`. ## Tests - `split_lend_complement_test` (new, CPU-runnable): a two-stage split built the way `generate` builds the call (CUDA0's cache as the primary, the later stage's pairs as `additional_gpu_pairs`, empty complement). The device calls are wrapped at the link — `conversation_transfer_test`'s `--wrap` pattern — and `pin = false` keeps the copy pageable. It checks the copy holds exactly the two stages' lend-region pairs (`has_resident`, `resident_lent_slots`), that the borrowed experts are served from the RAM copy (`blob(l,e) == resident_blob(l,e)`, matching bytes, and `file_reads() == 0` — not one read reaches the file source), that a `--resident-budget-gib` cap splits the two regions at the first pair that does not fit while the pair past it keeps its file fallback (`file_reads() == 1`), and the old #848 shape with no regions set (copy keeps nothing). Red on HEAD: `'class strata::core::FileExpertSource' has no member named 'stage_lend_regions'`. - `file_expert_source_test.cpp: test_split_lend_regions` (extended): the plan-level walk — both regions compacted after an existing complement entry (passed over, not counted twice), whole-copy cap, a refused walk leaves the plan untouched, out-of-range pair rejected before anything changes. Red on HEAD: `'append_stage_lend_regions' was not declared in this scope`. ## Measured after the change Same rig and config as above, `--pcie-frac 0.20` for all three columns (2026-10-06; interleaved A B C C runs, 5 samples per run, medians): | | arena (default) | `--resident-experts` before | `--resident-experts` with this change | |---|---|---|---| | RAM copy | ~47 GiB (all experts) | 25.13 GiB | 32.32 GiB (25.1 complement + 7.20 GiB lend regions, 3849 slots) | | available RAM (58 GiB host) | 4.2 GB | 24.8 GB | 17.5 GB | | prefill, 73K cold | 2341 t/s | 1892 t/s | 2287 / 2285 t/s (two runs) | | short prompts (66-85 tokens, cold) | 96-105 t/s | 50-65 t/s | 87-104 t/s | | NVMe read over one 73K prompt | 2.1 GB | 3.7 GB | 2.1 GB | The 2.1 GB that remains in every column is the PLE n-gram table (`--ple-io direct`, unbuffered reads of a 26.8 GiB shard, off the critical path), not experts. `conversation_cache_parity.py` reuse / disk / mixed pass byte-exact with this change on the split. Decode is unchanged by it (the regions only matter while a prompt borrows; with resident vs arena decode differs by the cache-miss path, not by this change). 🤖 Generated with [Claude Code](https://claude.com/claude-code)
関連リンク
インストール・モデル・リリースへの站内リンク。