Pull requests / #531

#531 multi-GPU: second GPU as an opt-in expert-cache tier (--peer-device), resubmit of #229 on v0.1.36

closed · @q8atnight · 0 comments · View on GitHub

Multi-GPUNVIDIA / CUDAModels & quantsDocumentationWindows

Description

Resubmit of #229 part 1, rebased on v0.1.36, per your closing comment there. `--peer-device N` puts an adaptive
expert cache on a second visible GPU beside CUDA0's, holding the ranked pairs the primary cache does not hold.
Activations and result rows cross NVLink or another P2P path; the peer computes its experts' rows for decode
windows and prompt chunks. Without the flag everything is exactly as released.

## Checklist from #229

- **Rebase past #216's session carve and 1024-token streaming:** the base is v0.1.36 (both landed in 0.1.30);
  the peer tier is rebased on top and keeps them intact.
- **`cudaHostAllocPortable` only with a peer:** `core::set_peer_portable(o.peer_device >= 1)` runs once after arg
  parsing; verify.cpp, mtp.cpp and the prefill group table add the Portable flag only then. No layer-split or
  single-GPU run pays for it.
- **Which `--layer-split` / `--expert-cache-device1` combinations are refused:** `--peer-device` now refuses both
  `--layer-split` (a different second-GPU mode, use one or the other) and `--expert-cache-device1..3` (both would fill
  the same card with an expert cache), with a one-line message each; plus the existing refusals (no `--expert-profile` /
  expert cache, device not visible). Documented in docs/SECOND_GPU.md and in `--help` (new lines for the `--peer-*` flags).
- **Auto sizing on 22 GB cards needs #216's lending:** not in this PR. `--peer-reserve-mib` and `--peer-slots`
  are the manual knobs; the docs say so and name #216 as the follow-up.

## Numbers (2x RTX 3090 + NVLink, IQ3_S, v0.1.36)

| v0.1.36 | prefill 8K / 32K / 128K (t/s) | decode 8K / 32K / 128K (t/s) | 3,000-token story |
|---|---|---|---|
| stock, one 3090 | 2093 / 2201 / 2019 | 96.7 / 91.3 / 86.0 | 88 t/s |
| + this PR (`--peer-device 1 --peer-reserve-mib 1850`) | 2401 / 2747 / 2514 (+15 / +25 / +25 %) | 115.2 / 104.5 / 99.6 (+19 / +14 / +16 %) | 97 t/s |

Our private build with the follow-ups (GDN head split etc.) reaches 2657 / 3118 / 3010.

## Exactness

Gate: `STRATA_IQ_MT_MIN=1` + our local debug `STRATA_MMQ_NO_STREAMK=1`, `--adapt-swaps 0`, same expert set
(single `--expert-cache 6000` vs dual `--expert-cache 3000` + `--peer-slots 3950`).

- v0.1.36 single reproduces our 0.1.31 single reference (flashbench gate PASS).
- This build without `--peer-device` == v0.1.36 single (flashbench gate + 19.5K-token long gate PASS): default unchanged.
- This build with `--peer-device 1` == v0.1.36 single (both gates PASS): byte-identical.

## Independent data

- Adamyno on #229: 2x 2080 Ti (sm_75, PCIe P2P, no NVLink) ran the old part 1; with the
  `--peer-slots 256 --peer-reserve-mib 1536` workaround prefill is at parity and decode is -11-14 %, and the
  x8 link kills the gain — auto sizing needs #216's lending first.
- #390: the layer-split camp is converging on the same helper expert tier (9,269 -> 14,172 slots, 4-way).
- #448 (fixed in 0.1.35): the advice there — big card alone, small card as extra expert cache — is what
  `--peer-device` automates.

## Limits / follow-ups

- With a peer the prompt path keeps MMQ; the fused int8 path (#136) is not combined with the peer's rows yet
  (fused ring sizing switches off when a peer is configured).
- Tier sizing is manual (`--peer-reserve-mib`, `--peer-slots`) until #216's buffer lending lands.
- Follow-up queued right after this PR: GDN head split over both cards, +6-8 % prefill, byte-identical,
  announced on #204; the branch is already on this fork (`preview/gdn-split`).

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.