Pull requests / #229
#229 multi-GPU: second GPU as an opt-in expert-cache tier with P2P rows (--peer-device), part 1 of #204
closed · @q8atnight · 0 コメント · GitHub で見る
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quants
本文
Part 1 of the peer-tier plan agreed in #204: a second GPU as an **opt-in** expert-cache tier beside `--gpus`. Part 2 (the QSA query split) is not included — see the note at the end. ## What it does (only with `--peer-device N`) - **Second expert-cache tier** on GPU N: an adaptive cache of its own, filled from the profile after the primary's cache; decode rows come back through the pool's mapped rows (the verify graph is untouched). - **P2P rows in the prompt path**: the peer computes its experts' rows of each prompt chunk (activations in / rows out over P2P), with compact group buffers, the rows returned group by group on a second stream, and a streaming ring that lets the peer stream a share of the primary's streamed experts over its own PCIe link. - Peer rows are scattered straight into the pool's mapped rows (those buffers are now `cudaHostAllocPortable`, the only unconditional change — no behaviour change without a peer). New flags: `--peer-device N` (default -1 = off), `--peer-reserve-mib` (600), `--peer-slots` (0 = auto), `--peer-adapt-swaps` (-1), `--peer-prefill-rows` (-1 = chunk×K/2, 0 = off). Env switches, all only read when a peer is active: `STRATA_PF_PEER_OUT_PIPE`, `STRATA_PF_PEER_COMPACT`, `STRATA_PF_PEER_STREAM` (0.35), `STRATA_PF_PEER_RING` (48), `STRATA_PEER_DIRECT`, `STRATA_PEER_HOT` (0.45), `STRATA_PEER_HOT_AT`. Rebased on 0.1.28; upstream's `--expert-cache-device1..3` tier, the D-5 issuer thread and D-1 attention are left intact. ## Exactness (#152 method) Env trio `STRATA_NO_IQ512=1 STRATA_NO_IQ256=1 STRATA_NO_IQ4NL=1` + no-stream-k MMQ (local debug patch), greedy, the three gate prompts (chat / code / math), IQ3_S: - **Default path (no `--peer-device`) vs plain 0.1.28: byte-identical (PASS).** - **Peer on vs single GPU with the same GPU expert set** (single `--expert-cache` 7846 slots; peer run 3896 primary + 3950 peer, static, `STRATA_PEER_HOT=0`): **byte-identical (PASS).** ## Numbers — 2× RTX 3090 (NVLink), IQ3_S, 0.1.28 base, flashbench `--quick` (greedy) | setup | prefill 32K (tok/s) | decode 1K | decode 32K | | --- | ---: | ---: | ---: | | single GPU (auto cache) | 2,094 | 91.5 | 88.8 | | `--gpu1-experts` (`--expert-cache-device1 10000`) | 1,425 | 88.8 | 86.4 | | layer split (`--layer-split auto`) | 936 | 104.2 | 94.9 | | **this PR (`--peer-device 1`, auto)** | **2,566** | **106.5** | **106.7** | (Peer run: 8,692 primary + 11,168 peer = 19,860 of 24,576 experts on the GPUs. One run per cell.) ## About part 2 (QSA query split) We ported it too, but on 0.1.28 it does not meet the bar: prefill attention now uses the D-1 tensor-core kernel, which our split cannot divide mid-chunk, so the split falls back to the older kernel — not byte-identical, and slower than this PR alone (2,125 vs 2,566 at 32K). So it's left out. We might come back with a redesign that keeps D-1 on both cards; if so it will be a separate PR against this one.
関連リンク
インストール・モデル・リリースへの站内リンク。