Pull requests / #1665
#1665 resident RAM: spend --resident-budget-gib before an auto helper cache
closed · @W1nge · 0 comments · View on GitHub
BenchmarksMulti-GPUNVIDIA / CUDA
Description
## The problem A prompt reads back every routed expert's blob from wherever it lives. With `--expert-cache-device1..3` that is one copy per expert over *that card's* PCIe link, while the same bytes served by the host tier travel the primary's link. On a 2080 Ti (Gen 3 x16) + P100 (Gen 3 **x4**) rig the auto cache gave CUDA1 8782 experts (14.29 GiB) and the prompt path read 14.63 GiB back per prompt at the x4 link's saturated 3.4 GB/s: **4.3 s of an 8 s prompt, its largest single term**. The primary card was idle for most of it (`nvidia-smi` sampled at 300 ms: 5-57 % utilization, 24-122 W of 310 W). `--resident-budget-gib` already exists to put experts in host RAM, but the helper caches are filled *first* from the same ranked list, so the RAM copy only ever gets the complement the helpers leave behind - raising the budget alone did nothing (a 12 GiB and a 17 GiB budget both produced the same 10.84 GiB copy). ## The change The ranked list is now spent **primary cache, RAM, helpers**: the resident budget's walk marks its pairs in `claimed`, which is exactly what a helper's fill already skips (a layer split's pairs are recorded there the same way). The reservation uses the same bound the resident copy does (`detail::clamp_resident_budget`), so it cannot reserve more than the RAM will hold. An explicit `--expert-cache-device1 N` is left alone, and so is the whole tier set around one. ## Measured 5K prompt fixture (cold prompt cache), 4 baseline runs vs 3 with the change; the harness's run-to-run spread on the same config is +-8 %: | | before | after | | --- | --- | --- | | 2K prompt | 6367, 6775, 7385, 8185 ms | **5554, 5610, 5687 ms** | | 5K prompt | 7595, 7989, 8139, 8243 ms | **7234, 7523, 7671 ms** | | helper bytes per prompt | 14631.5 MiB | **11157.0 MiB** | | CUDA1 cache | 8782 experts | 6671 experts | | host RAM copy | 10.84 GiB | 14.23 GiB | | 256-token decode | 63.4, 66.0, 69.3, 71.8 tok/s | 61.9, 64.2, 67.0, 68.3 tok/s | The decode side is the tradeoff: the helper's cache is also what lets it compute a decode expert locally, so a smaller cache means more decode experts fetched over PCIe. It lands ~3 % lower here, inside the fixture's own spread, against a 7-20 % prompt gain - and the change only applies when an explicit budget asks for the RAM residency, so a rig that values the helper cache can leave `--resident-budget-gib` off. Correctness: the bytes are the same bytes, only their tier changes; the daily end-to-end fixture (5 cases, cold) passes, and the existing `--layer-split` path's `claimed` use is unchanged.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.