贡献 / #1665

#1665 resident RAM: spend --resident-budget-gib before an auto helper cache

closed · @W1nge · 0 评论 · 去 GitHub 看

BenchmarksMulti-GPUNVIDIA / CUDA

说明

## The problem

A prompt reads back every routed expert's blob from wherever it lives. With
`--expert-cache-device1..3` that is one copy per expert over *that card's* PCIe
link, while the same bytes served by the host tier travel the primary's link.

On a 2080 Ti (Gen 3 x16) + P100 (Gen 3 **x4**) rig the auto cache gave CUDA1
8782 experts (14.29 GiB) and the prompt path read 14.63 GiB back per prompt at
the x4 link's saturated 3.4 GB/s: **4.3 s of an 8 s prompt, its largest single
term**. The primary card was idle for most of it (`nvidia-smi` sampled at
300 ms: 5-57 % utilization, 24-122 W of 310 W).

`--resident-budget-gib` already exists to put experts in host RAM, but the
helper caches are filled *first* from the same ranked list, so the RAM copy only
ever gets the complement the helpers leave behind - raising the budget alone did
nothing (a 12 GiB and a 17 GiB budget both produced the same 10.84 GiB copy).

## The change

The ranked list is now spent **primary cache, RAM, helpers**: the resident
budget's walk marks its pairs in `claimed`, which is exactly what a helper's
fill already skips (a layer split's pairs are recorded there the same way). The
reservation uses the same bound the resident copy does
(`detail::clamp_resident_budget`), so it cannot reserve more than the RAM will
hold. An explicit `--expert-cache-device1 N` is left alone, and so is the whole
tier set around one.

## Measured

5K prompt fixture (cold prompt cache), 4 baseline runs vs 3 with the change;
the harness's run-to-run spread on the same config is +-8 %:

| | before | after |
| --- | --- | --- |
| 2K prompt | 6367, 6775, 7385, 8185 ms | **5554, 5610, 5687 ms** |
| 5K prompt | 7595, 7989, 8139, 8243 ms | **7234, 7523, 7671 ms** |
| helper bytes per prompt | 14631.5 MiB | **11157.0 MiB** |
| CUDA1 cache | 8782 experts | 6671 experts |
| host RAM copy | 10.84 GiB | 14.23 GiB |
| 256-token decode | 63.4, 66.0, 69.3, 71.8 tok/s | 61.9, 64.2, 67.0, 68.3 tok/s |

The decode side is the tradeoff: the helper's cache is also what lets it
compute a decode expert locally, so a smaller cache means more decode experts
fetched over PCIe. It lands ~3 % lower here, inside the fixture's own spread,
against a 7-20 % prompt gain - and the change only applies when an explicit
budget asks for the RAM residency, so a rig that values the helper cache can
leave `--resident-budget-gib` off.

Correctness: the bytes are the same bytes, only their tier changes; the daily
end-to-end fixture (5 cases, cold) passes, and the existing `--layer-split`
path's `claimed` use is unchanged.

本站相关内容

相关页面的快捷入口。