Issues / #723
#723 Design question before we build: two lanes on a peer pair that help each other ("mutual help") — welcome upstream, or keep it in a fork?
open · @q8atnight · 2 コメント · GitHub で見る
BenchmarksNVIDIA / CUDAModels & quants
本文
Hi Niko — a design question before we write code, not a PR.
**Context.** Thanks again for merging the peer tier as `--peer-device` (#531). On our box (2× RTX 3090 + NVLink, IQ3_S, 262K, Ryzen 3950X, 121 GB DDR4) we now run an experiment along the lines of #127: when a second request arrives, the helper card stops being the lead's peer and becomes its own one-card engine (both processes share the `MAP_SHARED` arena); once it has been idle for a while it goes back to being the peer. Each process still serves exactly one sequence, as in your design.
**What we measured on real agent traffic** (two coding-agent chats at 60–98K context with tool calls, not a synthetic bench):
| mode | decode per chat | expert-cache hit |
|---|---|---|
| one chat, lead + peer (`--peer-device`) | 110–160 t/s | 0.98 |
| two chats, two independent one-card lanes | lead ~75–80, other ~55–65 t/s | 0.74–0.89 |
So the two lanes together are only ~15 % faster than one chat on the peer pair, and each chat runs at roughly half speed. That matches what you wrote in #97 ("two at once would roughly halve each"). The loss comes from the cache split: each lane only has its own card's cache, so 11–26 % of the expert work falls back to the CPU. CPU was ~70 % idle at the same time; the misses look bandwidth-bound.
**The idea ("mutual help").** Keep two processes with one sequence each, but leave the peer tier on in both directions: each lane is the other's `--peer-device`. The two caches hold different experts (deduplicated), so each lane computes an expert on whichever card has it, and both see roughly the peer-pair coverage instead of the CPU fallback. No batching, no shared decode loop; the cost is that each card's GPU time is shared between two sequences, plus the peer sync in both directions. With 3+ cards (72+ GB) all experts would fit on the GPUs, but pairwise peering doesn't scale past an NVLink pair. So the long-term version would probably be one shared expert pool sliced across all cards, which comes close to the multi-slot engine you decided against in #465 / #249.
**Questions:**
1. Would a bidirectional peer tier (two lanes, each the other's peer, deduplicated caches) fit your plans for upstream, if it stays opt-in, byte-identical by default and gated like #531? Or would you rather keep the peer tier strictly lead → helper?
2. For more than two cards: do you see Strata staying at "one sequence per process, one process per card", with any cross-card sharing at the expert-tier level? Or would you consider a shared multi-card expert pool at some point? If it's a clear no, we'll keep that part in our fork and won't come back with it as a PR.
We'll measure first on our side (real-workload bench, then a 2-card prototype) and only come back with numbers.
関連リンク
インストール・モデル・リリースへの站内リンク。