Pull requests / #110
#110 Multi-GPU: a second GPU as an expert tier for decode and prompts (--peer-device), + multi-conversation cache
closed · @q8atnight · 0 commentaires · Sur GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsLinux
Description
## What Two-GPU support (closes #36), tested on 2x RTX 3090 with NVLink (P2P also works over PCIe, slower). The second card is an **expert tier**, not a pipeline stage - the model's dense part stays on the primary, so the output is bit-identical to a single card holding the same experts (see Exactness): - `--peer-device N` (+ `--peer-reserve-mib`, `--peer-slots`, `--peer-adapt-swaps`): a second adaptive expert cache on the other GPU. Decode: the pool launches the peer's share right after publishing the primary's plan; its rows come back through the same mapped rows. Prompts: per MoE layer the peer computes the rows of the experts it holds (activations in over P2P, results back group by group), streams a share of the primary's non-resident experts over **its own** PCIe link (`STRATA_PF_PEER_STREAM`, default 0.35), and does the selection + attention of the last half of each chunk's queries in QSA layers. Group-sized (compact) peer buffers keep it at ~0.26 GB on the helper card. - `--expert-cache-fill-all`: auto sizing fills the whole card past the ranked profile pairs - `--ple-io ram`: the n-gram table mmap'd + mlock'ed (no O_DIRECT SSD reads on the prompt/token path) - `--conv-cache N` (+ `-gib`, `-min`): up to N conversations another request displaced keep their positional cells (KV host copy, pooled indexer keys, drafter K/V) + checkpoints in RAM; switching back reads only the new part (a 111K chat back in ~1 s instead of ~64 s). Complements #62 (which keeps the shared prefix of ONE branch). - mid-prompt checkpoints copied out asynchronously (pinned staging on the prompt stream + a thread) instead of a device sync + pageable copies (~370 ms of GPU idle each) - the drafter's prompt pass runs on a thread beside the next chunk - the config's `vision` block may carry `env` (e.g. `CUDA_VISIBLE_DEVICES=0` puts the image encoder on the helper card) - includes #108 (prefill kernels) and #109 (decode batching); if they are merged first this diff shrinks accordingly - diagnostics: `STRATA_PREFILL_TIMING` now also prints the peer GPU's own timeline ## Measured (flashbench 1.0, IQ3_S, 262K context, tokens/s, median) | | prefill 8K / 32K / 128K | decode greedy 1K / 32K / 128K | decode sampled | |---|---|---|---| | stock 0.1.13, one 3090 | 1376 / 1361 / 1223 | 45 / 46 / 41 | ~40 | | this branch (on 0.1.20), two 3090s | 1876 / 2095 / 1968 | 112 / 102 / 96 | 80-95 | With Q2_0 / IQ2_XS all 24,576 experts fit on the two cards (Q2_0: ~2,170 tok/s prompt at 32K). ## Exactness Gate: a single-GPU config holding N experts vs a two-GPU config holding the same N split across both cards, static residency (`--adapt-swaps 0 --pcie-frac 0`), `STRATA_MMQ_NO_STREAMK=1` (a debug switch that turns off MMQ stream-k, whose tile split depends on the group's total tiles): `flashbench.py gate` greedy outputs IDENTICAL on all prompts, after every step. With stream-k on, results differ only by MMQ's summation order. ## Notes - The primary should be the card without the desktop (`CUDA_VISIBLE_DEVICES=1,0`); P2P is enabled both ways at open. - Memory on the helper: keep >= ~250 MiB "left free" (the engine prints it); `--peer-reserve-mib` sets it. - Developed with an AI coding assistant; all numbers measured on 2x RTX 3090 (NVLink) / Ryzen 9 3950X / 121 GB RAM.
Sur le site
Liens install, modèles, releases.