Pull requests / #988

#988 Resident RAM mode with a layer split: only the experts no stage's cache holds (UD-Q4_K_XL on 2x 3090 with 64 GB)

closed · @MirkoMorello · 0 commentaires · Sur GitHub

BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Description

## What

With a layer split, Strata keeps **every** expert in RAM - also the ones the cards' caches already hold. For UD-Q4_K_XL that is 72 GiB, so setup asks for ~135 GB of RAM before it offers two GPUs (#498, #737, #806). The alternative, `--mmap-experts` with a split, needs less RAM but leaves every miss to the CPU: the cards have no mapped alias of the file pages.

This PR lets the resident RAM mode run with a split, opt-in, with no change to setup or to any default:

- `--resident-cpu-experts` is accepted with `--layer-split` (still refused with the helper caches, `--expert-cache-remote`).
- `pin_cache_complement` gets the later stages' cached experts as its `additional_gpu_pairs` (the parameter already existed for a second GPU tier), so the page-locked copy holds only what **no** card holds.
- `resident_stage_swaps` copied an evicted expert back from CUDA0's cache for every layer. With a split, a swap on a later stage's layer would have read the wrong slot. It now takes the cache, device and stream of the stage that owns the layer.

`--resident-experts` alone with a split still runs as the plain mmap mode (#364, #384); adding `--resident-cpu-experts` makes it explicit. Removing that downgrade, and letting setup offer the split from the RAM this mode needs (experts minus what the caches hold, plus the headroom), are left to you.

## Measured

2x RTX 3090 (PCIe 4.0 x16, no P2P), Ryzen 9 5900X (AVX2, no AVX-512), 64 GB DDR4 (62.7 GiB usable), Ubuntu 24.04, driver 580, engine built from `6f32ec0` for sm_86 in Docker (CUDA 13.0). UD-Q4_K_XL at revision `38bb39e`, 200K context, fp16 KV, MTP (q2_0 draft), `--prefill auto`, split auto (K=24; the caches hold 10,013 experts, 29 GiB).

Three requests with a ~16.6K-token shared prefix, 768 tokens out each (Python, Italian prose, JSON; temperature 0.7, seed 42), measured from the client over the OpenAI streaming API:

| | RAM copy | Decode tok/s | Prompt tok/s (16.6K) | Three requests |
| --- | ---: | ---: | ---: | ---: |
| split, `--mmap-experts` | page cache | 30.9 | 461 | 122.0 s |
| split, `--mmap-experts --pool-workers 22` | page cache | 29.4 | 472 | 124.9 s |
| split, resident (this PR) | 42.5 GiB page-locked | 63.9 | 735 | 63.8 s |

- `STRATA_DECODE_TIMING=1`: CPU expert jobs per window went from 41-64 ms (mmap) to ~0.02 ms; per window 37-45 ms in all. 2-3% of the routed experts are read over PCIe, about 90% hit the caches.
- A 175K-token prompt (source files with three needles, then a second turn): read in 89 s, decode 51.8 tok/s at 175K, 3/3 needles found, and the second turn reused the prefix (1.6 s to its first token). `MemAvailable` never went under 11.3 GiB, with `STRATA_RESIDENT_HEADROOM_GIB=2`.
- For scale, #498 measured 64-78 tok/s for the same split with all the experts in RAM (165 GiB).

## Tests

- `file_expert_source_test`: PASS, with a new case for the split, where the later stage's experts come in as the additional tier.
- `strata` builds with `-DSTRATA_BUILD_TESTS=ON` (sm_86); no new warnings.
- Through the server: a streaming tool call and its follow-up turn, and a coding agent (Pi) reading a file with its tools; both correct. The engine log of the serving instance has no `adaptive swap` / `refill` failures (a short session, not a soak test).

## Not done

- No perplexity or top-k comparison against llama.cpp *in this mode*. The numerics are the resident mode's (#403), so I would expect the agreement documented in UNSLOTH_Q4.md, but I have not measured it here.
- Not run on AMD or on more than two GPUs. Nothing in the change is CUDA-only, but the HIP stream and device switch paths are untested.
- Docs: a section in MULTI_GPU.md and a pointer from UNSLOTH_Q4.md.

Sur le site

Liens install, modèles, releases.