Pull requests / #531
#531 multi-GPU: second GPU as an opt-in expert-cache tier (--peer-device), resubmit of #229 on v0.1.36
closed · @q8atnight · 0 comments · View on GitHub
Multi-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
Description
Resubmit of #229 part 1, rebased on v0.1.36, per your closing comment there. `--peer-device N` puts an adaptive expert cache on a second visible GPU beside CUDA0's, holding the ranked pairs the primary cache does not hold. Activations and result rows cross NVLink or another P2P path; the peer computes its experts' rows for decode windows and prompt chunks. Without the flag everything is exactly as released. ## Checklist from #229 - **Rebase past #216's session carve and 1024-token streaming:** the base is v0.1.36 (both landed in 0.1.30); the peer tier is rebased on top and keeps them intact. - **`cudaHostAllocPortable` only with a peer:** `core::set_peer_portable(o.peer_device >= 1)` runs once after arg parsing; verify.cpp, mtp.cpp and the prefill group table add the Portable flag only then. No layer-split or single-GPU run pays for it. - **Which `--layer-split` / `--expert-cache-device1` combinations are refused:** `--peer-device` now refuses both `--layer-split` (a different second-GPU mode, use one or the other) and `--expert-cache-device1..3` (both would fill the same card with an expert cache), with a one-line message each; plus the existing refusals (no `--expert-profile` / expert cache, device not visible). Documented in docs/SECOND_GPU.md and in `--help` (new lines for the `--peer-*` flags). - **Auto sizing on 22 GB cards needs #216's lending:** not in this PR. `--peer-reserve-mib` and `--peer-slots` are the manual knobs; the docs say so and name #216 as the follow-up. ## Numbers (2x RTX 3090 + NVLink, IQ3_S, v0.1.36) | v0.1.36 | prefill 8K / 32K / 128K (t/s) | decode 8K / 32K / 128K (t/s) | 3,000-token story | |---|---|---|---| | stock, one 3090 | 2093 / 2201 / 2019 | 96.7 / 91.3 / 86.0 | 88 t/s | | + this PR (`--peer-device 1 --peer-reserve-mib 1850`) | 2401 / 2747 / 2514 (+15 / +25 / +25 %) | 115.2 / 104.5 / 99.6 (+19 / +14 / +16 %) | 97 t/s | Our private build with the follow-ups (GDN head split etc.) reaches 2657 / 3118 / 3010. ## Exactness Gate: `STRATA_IQ_MT_MIN=1` + our local debug `STRATA_MMQ_NO_STREAMK=1`, `--adapt-swaps 0`, same expert set (single `--expert-cache 6000` vs dual `--expert-cache 3000` + `--peer-slots 3950`). - v0.1.36 single reproduces our 0.1.31 single reference (flashbench gate PASS). - This build without `--peer-device` == v0.1.36 single (flashbench gate + 19.5K-token long gate PASS): default unchanged. - This build with `--peer-device 1` == v0.1.36 single (both gates PASS): byte-identical. ## Independent data - Adamyno on #229: 2x 2080 Ti (sm_75, PCIe P2P, no NVLink) ran the old part 1; with the `--peer-slots 256 --peer-reserve-mib 1536` workaround prefill is at parity and decode is -11-14 %, and the x8 link kills the gain — auto sizing needs #216's lending first. - #390: the layer-split camp is converging on the same helper expert tier (9,269 -> 14,172 slots, 4-way). - #448 (fixed in 0.1.35): the advice there — big card alone, small card as extra expert cache — is what `--peer-device` automates. ## Limits / follow-ups - With a peer the prompt path keeps MMQ; the fused int8 path (#136) is not combined with the peer's rows yet (fused ring sizing switches off when a peer is configured). - Tier sizing is manual (`--peer-reserve-mib`, `--peer-slots`) until #216's buffer lending lands. - Follow-up queued right after this PR: GDN head split over both cards, +6-8 % prefill, byte-identical, announced on #204; the branch is already on this fork (`preview/gdn-split`).
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.