Pull requests / #110

#110 Multi-GPU: a second GPU as an expert tier for decode and prompts (--peer-device), + multi-conversation cache

closed · @q8atnight · 0 Kommentare · Auf GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsLinux

Beschreibung

## What

Two-GPU support (closes #36), tested on 2x RTX 3090 with NVLink (P2P also works over PCIe, slower). The second card
is an **expert tier**, not a pipeline stage - the model's dense part stays on the primary, so the output is
bit-identical to a single card holding the same experts (see Exactness):

- `--peer-device N` (+ `--peer-reserve-mib`, `--peer-slots`, `--peer-adapt-swaps`): a second adaptive expert cache
  on the other GPU. Decode: the pool launches the peer's share right after publishing the primary's plan; its rows come
  back through the same mapped rows. Prompts: per MoE layer the peer computes the rows of the experts it holds
  (activations in over P2P, results back group by group), streams a share of the primary's non-resident experts over
  **its own** PCIe link (`STRATA_PF_PEER_STREAM`, default 0.35), and does the selection + attention of the last half
  of each chunk's queries in QSA layers. Group-sized (compact) peer buffers keep it at ~0.26 GB on the helper card.
- `--expert-cache-fill-all`: auto sizing fills the whole card past the ranked profile pairs
- `--ple-io ram`: the n-gram table mmap'd + mlock'ed (no O_DIRECT SSD reads on the prompt/token path)
- `--conv-cache N` (+ `-gib`, `-min`): up to N conversations another request displaced keep their positional cells
  (KV host copy, pooled indexer keys, drafter K/V) + checkpoints in RAM; switching back reads only the new part
  (a 111K chat back in ~1 s instead of ~64 s). Complements #62 (which keeps the shared prefix of ONE branch).
- mid-prompt checkpoints copied out asynchronously (pinned staging on the prompt stream + a thread) instead of a
  device sync + pageable copies (~370 ms of GPU idle each)
- the drafter's prompt pass runs on a thread beside the next chunk
- the config's `vision` block may carry `env` (e.g. `CUDA_VISIBLE_DEVICES=0` puts the image encoder on the helper card)
- includes #108 (prefill kernels) and #109 (decode batching); if they are merged first this diff shrinks accordingly
- diagnostics: `STRATA_PREFILL_TIMING` now also prints the peer GPU's own timeline

## Measured (flashbench 1.0, IQ3_S, 262K context, tokens/s, median)

| | prefill 8K / 32K / 128K | decode greedy 1K / 32K / 128K | decode sampled |
|---|---|---|---|
| stock 0.1.13, one 3090 | 1376 / 1361 / 1223 | 45 / 46 / 41 | ~40 |
| this branch (on 0.1.20), two 3090s | 1876 / 2095 / 1968 | 112 / 102 / 96 | 80-95 |

With Q2_0 / IQ2_XS all 24,576 experts fit on the two cards (Q2_0: ~2,170 tok/s prompt at 32K).

## Exactness

Gate: a single-GPU config holding N experts vs a two-GPU config holding the same N split across both cards, static
residency (`--adapt-swaps 0 --pcie-frac 0`), `STRATA_MMQ_NO_STREAMK=1` (a debug switch that turns off MMQ stream-k,
whose tile split depends on the group's total tiles): `flashbench.py gate` greedy outputs IDENTICAL on all prompts,
after every step. With stream-k on, results differ only by MMQ's summation order.

## Notes

- The primary should be the card without the desktop (`CUDA_VISIBLE_DEVICES=1,0`); P2P is enabled both ways at open.
- Memory on the helper: keep >= ~250 MiB "left free" (the engine prints it); `--peer-reserve-mib` sets it.
- Developed with an AI coding assistant; all numbers measured on 2x RTX 3090 (NVLink) / Ryzen 9 3950X / 121 GB RAM.

Mehr auf der Site

Links zu Install, Modellen, Releases.