Pull requests / #1461
#1461 Peer tier with mutual help: concurrency 2 at 155 t/s aggregate on dual RTX 3090 (elastic peer pair, opt-in)
open · @q8atnight · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
描述
Follow-up to #723 and #531. As you asked there: opt-in, off by default, gated like the peer pair, numbers first. **What it does.** With `--peer-device` one engine uses both cards for one conversation and a second request waits. The elastic pair lets the second card turn into its own engine when a second request arrives, and back into the first engine's peer tier after it has been idle for a while. While both run, each engine computes the other's rows for the experts its card holds (`--peer-link`, "mutual help"), so neither falls back to the CPU for the experts the other card has. Still one sequence per engine process: two processes, one card each, one shared expert arena. | | lead engine (card 1) | helper engine (card 2) | |---|---|---| | one request | `--peer-device`: card 2 is its tier (as without the pair) | asleep: cache given back (#533 `--vram-elastic`), optionally frozen with `cuda-checkpoint` | | a 2nd request arrives | `PEER_DETACH` between two verify windows, `POOL_HOLD n` | thawed, its cache grown back from the shared arena, `LINK_SERVE 1`, takes the request | | both run | computes its own experts + the helper's rows for its card's experts | the same, the other way round | | idle for `hold_s` | `PEER_ATTACH` after its request: card 2 is its tier again | cache given back, frozen | ## The two numbers you asked for (2x RTX 3090 + NVLink, Ryzen 9 3950X, 121 GB DDR4, IQ3_S, 262K, kv int8) **1. Single request speed unchanged.** Stock v0.1.40.3 `--peer-device 1` vs this PR's elastic pair serving one request, 3 alternating restarts x 2 runs (n = 6, medians, sampled, cached 54K coding prompt for decode): | | decode short code / short prose / 54K context | prefill 50K | |---|---|---| | v0.1.40.3 `--peer-device 1` | 110.8 / 104.2 / 103.2 t/s | 2,846 t/s | | elastic pair, one request | 119.4 / 102.8 / 103.7 t/s | 2,844 t/s | (the lead IS a `--peer-device` engine while it is alone; the only difference is the cache split between the cards: every second profile rank, close to the default `STRATA_PEER_HOT` 0.45.) **2. The cost to the first card when the second request starts** (streamed, 3 runs, B starts 8 s into A): | first request (A) | t/s | |---|---| | alone | 102-112 | | while the helper wakes (4.7 s: detach 0.6 s, thaw 3.1 s, cache 1.1 s) | 72-77 (card 2's experts come from the CPU meanwhile) | | both running | 82-87 (B: 85-89) | | after B has ended | 94-104 | | longest gap in A's stream | 0.07 s - it never stops | B's first token comes 4.7 s plus its prompt after it was sent. **Aggregate, two 1,000-token requests at once** (cached prompts, 5 runs): **151-160 t/s** (median 155) vs **99 t/s** for the same two one after the other on `--peer-device` (+55 %). Two one-card engines without the mutual help: 136-139. ## Exactness Without `--peer-link` and without the `"elastic"` block nothing of this runs. Gates (greedy, word for word, the deterministic gate configs: `--expert-cache 3000 --adapt-swaps 0 --pcie-frac 0 --prompt-cache 0 --no-prefill-borrow`, `STRATA_IQ_MT_MIN=1`): stock v0.1.40.3 saved, stock repeated (PASS), this PR checked: - unsloth UD-Q4_K_XL, `--layer-split 24`: chat/code/math + 20K-token long prompt x2 + one image: **PASS** - IQ3_S, one GPU: chat/code/math + 20K long x2: **PASS** With the link on, a linked expert is computed with the same kernels (q8_1 + `native_expert_grouped`) from the same activations as on the requester's own card, so its rows are the same; as with any cache, which experts sit in VRAM vs on the CPU still decides the last bits. ## Where it does NOT help (honest row) unsloth UD-Q4_K_XL on the same cards: a `--layer-split` already spends both cards' memory bandwidth on one conversation (95 t/s); the pair gives 87-88 t/s for two together (75-79 for one request on the peer pair). Q4's dense weights + head (~2.9 GB) are read per verify window per engine, so two engines read them twice. The pair pays off where `--peer-device` leaves the second card mostly idle (as a peer tier alone it measured 11-18 % busy here). ## How it is built - `--peer-link FILE --peer-link-role 0|1` (engine): a small shared-memory file. Each engine publishes its card's residency; a service thread computes the other engine's rows on its own card (zero-copy through the mapped file, high-priority stream). Only the owner ever touches its card: it holds the link while it rewrites or unmaps slots (adaptive swaps, the prompt path's loan, VRAM) and publishes after the copies landed; an expert planned a moment too late is copied in from the owner's RAM copy for that request (counted in the log line). No CUDA IPC, no P2P between processes, no MPS. - Control lines on stdin: `PEER_DETACH`, `PEER_ATTACH`, `POOL_HOLD n`, `LINK_SERVE 0|1` (answered `CTL ...`). - `serve/elastic.py`: the desk (two `StrataEngine`s behind one `Service`, `batch = 2`), routing to the idle lane with the longest prefix; `serve/test_elastic.py` (fake engines, no GPU); `docs/MULTI_GPU.md`: setup + these numbers. ## Limits / follow-ups - Linux only (the link file, `sched_setaffinity`); two cards (a pair). The shared arena must fit its tmpfs (`/dev/shm` is half of RAM by default). - The helper's conversation is not handed to the lead when it sleeps: its next turn is read again by the lead. - The helper's chat prefills on one card (no peer rows in the prompt path while linked). - The lead detaches only between verify windows: when the second request arrives while the lead is still reading a long prompt, the helper wakes after that prompt (measured: 11.5 s into a fresh 54K prompt). - Two requests is the limit; a third waits for a lane. Measured with `bench/q4pair.py`-style scripts (one-shot, not part of the PR); happy to add them under `tools/` if useful. Development of this PR (design, engine code, desk, measurements) was heavily supported by Opus 5.5.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。