Pull requests / #229

#229 multi-GPU: second GPU as an opt-in expert-cache tier with P2P rows (--peer-device), part 1 of #204

closed · @q8atnight · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quants

描述

Part 1 of the peer-tier plan agreed in #204: a second GPU as an **opt-in** expert-cache tier beside `--gpus`. Part 2 (the QSA query split) is not included — see the note at the end.

## What it does (only with `--peer-device N`)
- **Second expert-cache tier** on GPU N: an adaptive cache of its own, filled from the profile after the primary's cache; decode rows come back through the pool's mapped rows (the verify graph is untouched).
- **P2P rows in the prompt path**: the peer computes its experts' rows of each prompt chunk (activations in / rows out over P2P), with compact group buffers, the rows returned group by group on a second stream, and a streaming ring that lets the peer stream a share of the primary's streamed experts over its own PCIe link.
- Peer rows are scattered straight into the pool's mapped rows (those buffers are now `cudaHostAllocPortable`, the only unconditional change — no behaviour change without a peer).

New flags: `--peer-device N` (default -1 = off), `--peer-reserve-mib` (600), `--peer-slots` (0 = auto), `--peer-adapt-swaps` (-1), `--peer-prefill-rows` (-1 = chunk×K/2, 0 = off).
Env switches, all only read when a peer is active: `STRATA_PF_PEER_OUT_PIPE`, `STRATA_PF_PEER_COMPACT`, `STRATA_PF_PEER_STREAM` (0.35), `STRATA_PF_PEER_RING` (48), `STRATA_PEER_DIRECT`, `STRATA_PEER_HOT` (0.45), `STRATA_PEER_HOT_AT`.

Rebased on 0.1.28; upstream's `--expert-cache-device1..3` tier, the D-5 issuer thread and D-1 attention are left intact.

## Exactness (#152 method)
Env trio `STRATA_NO_IQ512=1 STRATA_NO_IQ256=1 STRATA_NO_IQ4NL=1` + no-stream-k MMQ (local debug patch), greedy, the three gate prompts (chat / code / math), IQ3_S:
- **Default path (no `--peer-device`) vs plain 0.1.28: byte-identical (PASS).**
- **Peer on vs single GPU with the same GPU expert set** (single `--expert-cache` 7846 slots; peer run 3896 primary + 3950 peer, static, `STRATA_PEER_HOT=0`): **byte-identical (PASS).**

## Numbers — 2× RTX 3090 (NVLink), IQ3_S, 0.1.28 base, flashbench `--quick` (greedy)
| setup | prefill 32K (tok/s) | decode 1K | decode 32K |
| --- | ---: | ---: | ---: |
| single GPU (auto cache) | 2,094 | 91.5 | 88.8 |
| `--gpu1-experts` (`--expert-cache-device1 10000`) | 1,425 | 88.8 | 86.4 |
| layer split (`--layer-split auto`) | 936 | 104.2 | 94.9 |
| **this PR (`--peer-device 1`, auto)** | **2,566** | **106.5** | **106.7** |

(Peer run: 8,692 primary + 11,168 peer = 19,860 of 24,576 experts on the GPUs. One run per cell.)

## About part 2 (QSA query split)
We ported it too, but on 0.1.28 it does not meet the bar: prefill attention now uses the D-1 tensor-core kernel, which our split cannot divide mid-chunk, so the split falls back to the older kernel — not byte-identical, and slower than this PR alone (2,125 vs 2,566 at 32K). So it's left out. We might come back with a redesign that keeps D-1 on both cards; if so it will be a separate PR against this one.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。