Pull requests / #848
#848 Resident RAM mode with a layer split, --vram-reserve-later-mib (#642)
closed · @Hardin22 · 0 コメント · GitHub で見る
BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsWindows
本文
Refs #642. Two of the pieces you listed there as still open: the resident RAM mode on a layer split, and `--vram-reserve-later-mib`. Three commits, each one reviewable by itself: 1. **`--vram-reserve-later-mib N`**: the reserve of every card after the first (default: `--vram-reserve-mib`'s value), used by the split search and by each later card's cache sizing. The card that drives the monitors needs more headroom than one that drives none. 2. **Resident RAM mode with a layer split** (engine). `--resident-experts` no longer turns into `--mmap-experts` on a split: - the RAM copy leaves out the experts every stage's cache holds, not only CUDA0's; - when the rest does not fit whole (#467), it keeps the hottest by the whole profile. Ranked by CUDA0's share alone, it kept only layers 0..K-1 (4.0 GiB in RAM instead of 21.2 below); - an adaptive swap copies the evicted expert back from the cache of the card that owns its layer, on that card's stream. `host_res` holds slot numbers of whichever cache owns the layer, so reading a later stage's slot out of CUDA0's cache put another expert's bytes in RAM, and the answers turned to garbage after the first swaps. This was latent upstream because the guard kept the mode off a split. - Remote expert caches are still refused with it. 3. **setup**: with an engine that has it, the low-RAM mode on several GPUs takes the resident variant (same rule as on one card) instead of recommending one card. At a start, `split_mmap` leaves a resident config alone. Gated on `MIN_ENGINE >= RESIDENT_SPLIT_ENGINE`, which I set to `(0, 1, 40)` as a guess at the release that would carry it: until then setup behaves as today and the golden configs don't move. Please change the number if it's another release. ## Measured Swift 1.5 IQ3_XXS at 160K (q4_0 KV, `--prefill 4096`, `--spec 4` with the stock draft layer, `--pool-workers 15`). RTX 4060 Ti (layers 0-19) + RTX 5080 (20-47, display), i9-14900KF, 32 GB of RAM, Windows 11. Four greedy prompts (code, Italian prose, C, English prose), 400 tokens max, decode tok/s averaged over the four: | | round 1 | round 2 | round 3 | round 4 | |---|---|---|---|---| | upstream 6f32ec0, 5080 alone, `--resident-experts` | 24.8 | 27.5 | 27.0 | 28.8 | | upstream 6f32ec0, split 20, `--mmap-experts` (what `--resident-experts` becomes) | 31.9 | 53.8 | 64.7 | 64.0 | | this PR, split 20, `--resident-experts` (21.8-22.1 GiB of experts locked in RAM) | 70.6 | 69.1 | 70.3 | 69.0 | The mmap split catches up once the OS file cache holds the experts, and only because nothing else was running on the PC. The resident copy is there from the first request, and it stays locked when other programs want the RAM. - **Correctness**: - With `--pcie-frac 0 --adapt-every 0 STRATA_IQ_MT_MIN=1`, the split's greedy output is identical to upstream's `--mmap-experts` on the same split (4/4 prompts, 400 tokens each). - With adaptive swaps on, 33 varied requests at temperature 0.6 (thinking on/off, five languages, code, each checked for degenerate output) exchanged 81,783 experts between RAM and the two caches. Every answer was clean. - **One card**: unchanged; same numbers as upstream within noise. - **`--vram-reserve-mib 300 --vram-reserve-later-mib 1800`** on the same PC: the 4060 Ti's cache goes from 4551 to 4867 slots (vs `--vram-reserve-mib 1000` for both), and the 5080 keeps 1.3 GB free with everything loaded. - **setup**: `tools/test_setup_*.py` pass (260 tests). Two new tests in `test_setup_risk.py` run the new paths with `MIN_ENGINE` patched. ## Notes - #833 (`STRATA_RESIDENT_REMOTE=1`) edits the same guard in generate.cpp, so whichever goes second has a small textual conflict there. - Not in this PR: the pipelined windows from the fork. They sit on top of this and need more porting onto the batch-slot code. That will be a separate PR.
関連リンク
インストール・モデル・リリースへの站内リンク。