Pull requests / #1190

#1190 resident RAM mode with a layer split: the copy keeps each stage's lend region

open · @evaanp · 0 comentários · No GitHub

Server & APINVIDIA / CUDAModels & quants

Descrição


## The problem

#848 gave `--resident-experts` with a layer split a RAM copy of only the experts **no stage's cache holds**.
That left out the experts the prompt path borrows: `generate.cpp` hands `pin_cache_complement` the later
stages' cache pairs as `additional_gpu_pairs`, and with that list non-empty the function switches the
single lend region off (`expert_source.cpp:1752`, `lend = ... && additional_gpu_pairs.empty()` — the comment
at `generate.cpp:5820`: "with them the copy keeps no lend region: pin_cache_complement turns the loan off").
So under a split, the experts in the tail slots of **every** card's cache — CUDA0's part included — are in
neither VRAM nor the RAM copy during a prompt, and the prompt reads them from the model files through page
faults.

Measured 2026-10-06 (v0.1.40 build, config: 17/48 layer split, RTX 2000 Ada
stage 0 = layers 0-16 + RTX 5060 Ti stage 1 = 17-47, IQ3_S native pack without experts.bin so
`--mmap-experts` maps the GGUF shards, `--prefill auto` = 8192-token chunks; each card's prompt path borrows
4.31 GiB of its expert-cache slots):

| | arena (default) | `--resident-experts` (before) |
|---|---|---|
| RAM pinned for experts | ~47 GiB (all) | 25.13 GiB (complement only) |
| available RAM | 3.5 GB | 24.8 GB |
| decode | 62.8 t/s | 62.3 t/s |
| prefill 73K cold | 2339 t/s | 1892 t/s (-19%) |
| short prompts (70 tokens, cold) | ~90-110 t/s | ~50-65 t/s |
| NVMe read during one 73K prefill | ~1.1 GB | **~3.7 GB** |

`STRATA_PREFILL_TIMING=1`, one 8192-token chunk (GPU wall 3132 ms arena vs 3865 ms resident): dequant 232 vs
678 ms, host grouping 33 vs 187 ms, host staging 365 vs 1504 ms. The expensive experts were the file-read
ones: with the borrowed slots' bytes already in RAM, grouping/staging/dequant all stream the same
bytes the arena streams, and the walk collapses to `memcpy` + `cudaMemcpy` instead of a
per-blob page-fault trail through the mmap'd GGUF. (`--no-prefill-borrow` restored the short
prompts but cost the 73K prefill — 1428 t/s — and decode: not the fix.)

## The change

- `FileExpertSource::stage_lend_regions(...)` (header): each stage's (layer, expert) pairs, listed from
  the cache's highest slot down — the order the borrowing takes them.
- `detail::append_stage_lend_regions` (`expert_source.cpp:186`): adds the regions to the compact
  plan after the complement; a pair already in the copy is passed over, a stage's walk stops at the first
  pair that does not fit, and the next stage's walk still runs. One budget caps the whole copy — the same
  available-minus-headroom reading as the complement (#403), or an explicit `--resident-budget-gib`.
  Experts past it keep the mapped-file fallback they have today.
- `pin_cache_complement` folds the regions in after its own lend walk (`expert_source.cpp:1833`) and says
  how many of the split's lent slots stayed in RAM (`the split stages lend N slots to the prompt path: M
  keep their experts in RAM too`).
- `generate.cpp:5827`: sizes each part's region at the largest chunk the serve scan can ever pick — auto's
  ceiling `auto_ceiling`, or an explicit `--prefill`, which the scan only walks down — with the scan's own
  tail-slot arithmetic (`cache_slots_for`, the 128-slot floor, the auto `kAutoLendPct` cap), and honours
  #340's `own_ok` rule: a part whose card can fund the bound chunk's buffers outright gets no region and no
  RAM is spent on it.

RAM arithmetic on the rig above: 25.13 GiB (complement) + ~8.6 GiB (two regions of 4.31 GiB) ≈ 33.7 GiB —
~13.5 GiB less pinned than the arena, with none of the 2.6 GiB extra NVMe reads the one prompt cost.
Answers are unchanged: the copy holds the same file bytes the fallback would read.

Untouched: the arena default, single-GPU behaviour (no regions set; `pin_cache_complement`'s own lend walk
stays the only one), disk/keyless parking, and the expert-order/lookup semantics — resident pairs
resolve through the existing `cache_complement_blob_or_fallback`.

## Tests

- `split_lend_complement_test` (new, CPU-runnable): a two-stage split built the way `generate` builds
  the call (CUDA0's cache as the primary, the later stage's pairs as `additional_gpu_pairs`, empty
  complement). The device calls are wrapped at the link — `conversation_transfer_test`'s `--wrap`
  pattern — and `pin = false` keeps the copy pageable. It checks the copy holds exactly the two stages'
  lend-region pairs (`has_resident`, `resident_lent_slots`), that the borrowed experts
  are served from the RAM copy (`blob(l,e) == resident_blob(l,e)`, matching bytes, and
  `file_reads() == 0` — not one read reaches the file source), that a `--resident-budget-gib` cap splits the
  two regions at the first pair that does not fit while the pair past it keeps its file fallback
  (`file_reads() == 1`), and the old #848 shape with no regions set (copy keeps nothing).
  Red on HEAD: `'class strata::core::FileExpertSource' has no member named 'stage_lend_regions'`.
- `file_expert_source_test.cpp: test_split_lend_regions` (extended): the plan-level walk — both
  regions compacted after an existing complement entry (passed over, not counted twice), whole-copy cap,
  a refused walk leaves the plan untouched, out-of-range pair rejected before anything changes.
  Red on HEAD: `'append_stage_lend_regions' was not declared in this scope`.

## Measured after the change

Same rig and config as above, `--pcie-frac 0.20` for all three columns (2026-10-06; interleaved A B C C runs,
5 samples per run, medians):

| | arena (default) | `--resident-experts` before | `--resident-experts` with this change |
|---|---|---|---|
| RAM copy | ~47 GiB (all experts) | 25.13 GiB | 32.32 GiB (25.1 complement + 7.20 GiB lend regions, 3849 slots) |
| available RAM (58 GiB host) | 4.2 GB | 24.8 GB | 17.5 GB |
| prefill, 73K cold | 2341 t/s | 1892 t/s | 2287 / 2285 t/s (two runs) |
| short prompts (66-85 tokens, cold) | 96-105 t/s | 50-65 t/s | 87-104 t/s |
| NVMe read over one 73K prompt | 2.1 GB | 3.7 GB | 2.1 GB |

The 2.1 GB that remains in every column is the PLE n-gram table (`--ple-io direct`, unbuffered reads of a
26.8 GiB shard, off the critical path), not experts. `conversation_cache_parity.py` reuse / disk / mixed
pass byte-exact with this change on the split. Decode is unchanged by it (the regions only matter while a
prompt borrows; with resident vs arena decode differs by the cache-miss path, not by this change).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

No site

Links install, modelos, releases.