Pull requests / #237
#237 Use a second GPU's VRAM as an opt-in expert store for the prompt path
closed · @lukmanfauzie · 0 Kommentare · Auf GitHub
BenchmarksServer & APIMulti-GPUModels & quants
Beschreibung
Supersedes the `--vram-experts` part of #151, per @Niko1221's request to split the
three changes. Rebased as a single commit onto `main` = v0.1.29 (`d6708a4`) and
**opt-in**.
On a machine whose RAM cannot cache the 31.6 GiB expert set, every prompt re-reads the experts from the SSD and that read is the whole cost of the prompt (measured: 27.4 GiB re-read, 214.7 s for a 22,266-token prompt). A second card's VRAM holds the same bytes and hands them to the CPU ~19x faster than the engine's fault path can (0.334 ms against 6.31 ms per expert, measured on this box), while computing nothing.
`--vram-experts` (opt-in, needs `--mmap-experts`) uploads whole layers of routed experts into the second card's VRAM and serves the prompt path from there.
- `VramExpertStore` (`src/core/expert_source.{hpp,cpp}`) owns complete layers on one device and uploads them straight out of the mapped file, one blob at a time, so it needs NO host allocation at all - the host is the scarce resource on the machine this exists for. The layer set is chosen by routing frequency (the profile) until the card is full, greedily by each layer's OWN extent: a native pack's blobs differ layer to layer (1.5-2.3 MB across the Coder's 48), so one average would waste room or over-commit.
- `ExpertSource` gains `vram_held` / `vram_fetch` (defaulting to false), so every other source and caller is unchanged and a store is an ACCELERATOR, not a dependency: a fetch that misses falls back to the file.
- `prefill.cpp` fetches into a small ring of pinned handoff buffers and DMAs from there into the ring slot the H2D already reads, so `--vram-experts` costs no new device memory and ~5 MB of host RAM. Both prompt paths are wired: the staged path AND the streamed path (`--prefill 8192` is >= the stream gate, so a low-RAM run takes the streamed one for every expert - wiring only the staged path left the card filled, reported as holding whole layers, and never read).
- The store answers the PROMPT path only: decode's verify window is below its batch gate, so its layers deliberately STAY in device 0's cache - excluding them there sent 22 of 48 layers to the CPU at decode (42 -> 30 tok/s on Q2_0, 30 -> 15 on the Coder). Adaptation is disabled while the store is on, because it moves experts the store already owns.
- `--vram-expert-layers N` caps the layer count and `--vram-expert-device D` picks the card (default 1); both default to what the machine can hold.
Nothing here runs unless `--vram-experts` is given. The pinned handoff buffers are only allocated when the source actually has a store, and every store call site is guarded by that allocation, so a machine without the flag - or with one GPU - takes the previous code path unchanged. With one GPU, or without `--mmap-experts`, the flag warns and the run continues without a store.
`serve/server.py` no longer adds `--layer-split auto` to a config that already names `--vram-experts`. That flag wants the second card's whole VRAM as a copy store, while the split puts compute layers there and the two compete for the same memory; `--layer-split` is also refused by the engine unless it was started with `--serve` and can put a distinct GPU behind every K, so the generated line left the server waiting for a READY that never came. A config that wants the split is unchanged, and the new `GpuChoice` test pins all three cases.Mehr auf der Site
Links zu Install, Modellen, Releases.