Pull requests / #1461

#1461 Peer tier with mutual help: concurrency 2 at 155 t/s aggregate on dual RTX 3090 (elastic peer pair, opt-in)

open · @q8atnight · 0 comentarios · En GitHub

BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Descripción

Follow-up to #723 and #531. As you asked there: opt-in, off by default, gated like the peer pair, numbers first.

**What it does.** With `--peer-device` one engine uses both cards for one conversation and a second request waits.
The elastic pair lets the second card turn into its own engine when a second request arrives, and back into the
first engine's peer tier after it has been idle for a while. While both run, each engine computes the other's rows
for the experts its card holds (`--peer-link`, "mutual help"), so neither falls back to the CPU for the experts the
other card has. Still one sequence per engine process: two processes, one card each, one shared expert arena.

| | lead engine (card 1) | helper engine (card 2) |
|---|---|---|
| one request | `--peer-device`: card 2 is its tier (as without the pair) | asleep: cache given back (#533 `--vram-elastic`), optionally frozen with `cuda-checkpoint` |
| a 2nd request arrives | `PEER_DETACH` between two verify windows, `POOL_HOLD n` | thawed, its cache grown back from the shared arena, `LINK_SERVE 1`, takes the request |
| both run | computes its own experts + the helper's rows for its card's experts | the same, the other way round |
| idle for `hold_s` | `PEER_ATTACH` after its request: card 2 is its tier again | cache given back, frozen |

## The two numbers you asked for (2x RTX 3090 + NVLink, Ryzen 9 3950X, 121 GB DDR4, IQ3_S, 262K, kv int8)

**1. Single request speed unchanged.** Stock v0.1.40.3 `--peer-device 1` vs this PR's elastic pair serving one
request, 3 alternating restarts x 2 runs (n = 6, medians, sampled, cached 54K coding prompt for decode):

| | decode short code / short prose / 54K context | prefill 50K |
|---|---|---|
| v0.1.40.3 `--peer-device 1` | 110.8 / 104.2 / 103.2 t/s | 2,846 t/s |
| elastic pair, one request | 119.4 / 102.8 / 103.7 t/s | 2,844 t/s |

(the lead IS a `--peer-device` engine while it is alone; the only difference is the cache split between the cards:
every second profile rank, close to the default `STRATA_PEER_HOT` 0.45.)

**2. The cost to the first card when the second request starts** (streamed, 3 runs, B starts 8 s into A):

| first request (A) | t/s |
|---|---|
| alone | 102-112 |
| while the helper wakes (4.7 s: detach 0.6 s, thaw 3.1 s, cache 1.1 s) | 72-77 (card 2's experts come from the CPU meanwhile) |
| both running | 82-87 (B: 85-89) |
| after B has ended | 94-104 |
| longest gap in A's stream | 0.07 s - it never stops |

B's first token comes 4.7 s plus its prompt after it was sent.

**Aggregate, two 1,000-token requests at once** (cached prompts, 5 runs): **151-160 t/s** (median 155) vs **99 t/s**
for the same two one after the other on `--peer-device` (+55 %). Two one-card engines without the mutual help: 136-139.

## Exactness

Without `--peer-link` and without the `"elastic"` block nothing of this runs. Gates (greedy, word for word, the
deterministic gate configs: `--expert-cache 3000 --adapt-swaps 0 --pcie-frac 0 --prompt-cache 0 --no-prefill-borrow`,
`STRATA_IQ_MT_MIN=1`): stock v0.1.40.3 saved, stock repeated (PASS), this PR checked:
- unsloth UD-Q4_K_XL, `--layer-split 24`: chat/code/math + 20K-token long prompt x2 + one image: **PASS**
- IQ3_S, one GPU: chat/code/math + 20K long x2: **PASS**

With the link on, a linked expert is computed with the same kernels (q8_1 + `native_expert_grouped`) from the same
activations as on the requester's own card, so its rows are the same; as with any cache, which experts sit in VRAM vs
on the CPU still decides the last bits.

## Where it does NOT help (honest row)

unsloth UD-Q4_K_XL on the same cards: a `--layer-split` already spends both cards' memory bandwidth on one
conversation (95 t/s); the pair gives 87-88 t/s for two together (75-79 for one request on the peer pair). Q4's dense
weights + head (~2.9 GB) are read per verify window per engine, so two engines read them twice. The pair pays off
where `--peer-device` leaves the second card mostly idle (as a peer tier alone it measured 11-18 % busy here).

## How it is built

- `--peer-link FILE --peer-link-role 0|1` (engine): a small shared-memory file. Each engine publishes its card's
  residency; a service thread computes the other engine's rows on its own card (zero-copy through the mapped file,
  high-priority stream). Only the owner ever touches its card: it holds the link while it rewrites or unmaps slots
  (adaptive swaps, the prompt path's loan, VRAM) and publishes after the copies landed; an expert planned a moment
  too late is copied in from the owner's RAM copy for that request (counted in the log line). No CUDA IPC, no P2P
  between processes, no MPS.
- Control lines on stdin: `PEER_DETACH`, `PEER_ATTACH`, `POOL_HOLD n`, `LINK_SERVE 0|1` (answered `CTL ...`).
- `serve/elastic.py`: the desk (two `StrataEngine`s behind one `Service`, `batch = 2`), routing to the idle lane with
  the longest prefix; `serve/test_elastic.py` (fake engines, no GPU); `docs/MULTI_GPU.md`: setup + these numbers.

## Limits / follow-ups

- Linux only (the link file, `sched_setaffinity`); two cards (a pair). The shared arena must fit its tmpfs (`/dev/shm`
  is half of RAM by default).
- The helper's conversation is not handed to the lead when it sleeps: its next turn is read again by the lead.
- The helper's chat prefills on one card (no peer rows in the prompt path while linked).
- The lead detaches only between verify windows: when the second request arrives while the lead is still reading a
  long prompt, the helper wakes after that prompt (measured: 11.5 s into a fresh 54K prompt).
- Two requests is the limit; a third waits for a lane.

Measured with `bench/q4pair.py`-style scripts (one-shot, not part of the PR); happy to add them under `tools/` if useful.

Development of this PR (design, engine code, desk, measurements) was heavily supported by Opus 5.5.

En el sitio

Enlaces a install, modelos, releases.