Pull requests / #216
#216 Multi-GPU: the session carve and the per-stage prompt loans
closed · @gopinath87607 · 0 comments · View on GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindows
Description
…ranks' prompt rows Three things a layer split was paying for and not using. * THE SESSION CARVE. A stage carved state for ALL 48 layers whatever slice it ran - every QSA KV pool at max_cells, every GDN recurrence - the whole-model-sized-buffer disease the chunked-QSA-prefill PR cured upstream. `session_bytes`/`session_init` take [layer_lo, layer_hi), and a session now owns only that range: `qsa_states`/`gdn_state` keep GLOBAL ordinal indexing and simply hold fewer rows (`qsa_ord0`/`qsa_alloc`, `gdn_ord0`/`gdn_alloc`), per-layer consumers subtract `gdn_ord0`, and the model-level uses - the shared RoPE table, the prefill staging identity, the MTP rope borrow - read `qsa_states[qsa_primary()]`. A range with no QSA layer still carves ONE primary state so those uses stay valid, and the default full range is byte-identical to the old carve. On the 2x3060 + 2x5060 rig, CUDA1's [30, 48) session costs 0.93 GiB: 7.83 GiB free before the carve, 6.90 GiB after (`session [30, 48) and the head`). * PER-STAGE PROMPT LOANS. `no_prefill_borrow = true` for split stages said "each stage's prompt path has its own buffers" - true, but that is a reason to give each stage its own LOAN, not to make every stage withhold a chunk-sized reserve from its cache for the whole session, and forced on it also collapsed `--prefill auto` to 2048. Each prompt path now borrows the tail of ITS OWN stage's cache (`PfPart`), so outside the prompt the whole cache is expert cache: `--prefill auto` is the largest chunk every participant can lend (8192 tokens here), a participant refills its loan into its own slots and marks only its own layers (a slot refilled into another cache, or another stage's rows marked resident, is silent and produces plausible tokens), and a cache that cannot lend even a 256-token chunk is refused with both numbers rather than over-subscribing a card whose cache has already filled its VRAM. The flat 1 GiB withheld from every stage after the first becomes what is actually held back after the weights, session and drafter exist: 96 MiB of verify windows, plus 1000 MiB of drafter and head on the ONE stage that carries them. CUDA0 lends 2162 slots (3.98 GiB), CUDA1 1993 of its 2653 (3.97 GiB). * THE HELPER RANKS' PROMPT ROWS - SHIPPED OFF, AND MEASURED AS A LOSS. `RemoteExperts::has`/`begin_rows`/ `prompt_input`/`prompt_output` and `kBatchRows` let a helper rank (CUDA2/3) compute prompt rows for the experts it already holds, so those blobs are never staged to a stage; `Prefill::set_remote` drives it behind STRATA_PREFILL_REMOTE (STRATA_PREFILL_REMOTE_SHARE gradients it by dropping a hash-scattered share), the staging is one 4096-row batch per rank (CAP was MAXT*k for decode, now one prompt batch: ~42 MiB pinned and ~54 MiB of device scratch per rank), and `generate.cpp` hands the ranks to exactly ONE prompt path - a rank owns one set of buffers and one stream, and a split runs the next stage on its own thread. IT LOSES. 23,420-token prompt, measured on the 0.1.24 tree this was developed against: 38.1 -> 83.4 s of wall, 2.2x slower. Two causes, both measured: the 3.5M rows batch into 858 fully serialized 4096-row calls at ~52 ms each where the transfer is ~7 ms (1.5 GB/s against an 11.0 GB/s probe), and STRATA_REMOTE_ZEROCOPY (default on) makes the helper quantize from, and write into, BAR1-mapped host memory 41.9 MB at a time; turning it off alone recovers 26 s of timeline (76.9 -> 51.1 s) without closing the gap. It is therefore OFF by default and nothing in the shipped path changes. IT ALSO CORRUPTS UNDER MMQ, SO THAT COMBINATION IS NOW REFUSED. MMQ computes MMQ_GROUP=16 experts per launch and flushes the group's one GEMM from the `compute` call of the group's LAST member, so an expert this stage skips because a helper holds it silently takes its whole group's output with it - a third to a half of a layer's experts missing reads as fluent garbage (`!` x 64). With the per-expert path (STRATA_PREFILL_MMQ=0) the two arms are BIT-IDENTICAL (sha256(reasoning) equal, 542.1 vs 549.3 tok/s), which is what proves the row staging itself is exact. `Prefill::init` now prints the reason and drops the ranks instead of corrupting. * REPORTING. `expert_slots`/`expert_cache_mib` sum every tier: a 4-GPU 256K run reported CUDA0's cache alone (3327 experts) while the four cards held 13310, so the Monitor tab was wrong by 4x for every multi-GPU config. `expert_slots_primary`/`expert_cache_primary_mib` keep the old figure. The helper ranks' ms_begin/ms_wait are cumulative since boot and were printed next to per-request deltas; they take deltas now. KV-streaming stats iterate the session's OWNED ordinals. * REFUSALS. `--gpu-only-full` and `--gpu-stages` replay every layer through CUDA0's session, which a split carves to CUDA0's range, so they are refused with a layer split rather than answering about another stage's state; the token-graph hit path is disabled under a split (such a graph cannot span stages); and the sessions are allocated AFTER the split search, because the search prices each placement by what that range's session will cost on that device (`session_bytes` is pure arithmetic). updates Multi-GPU: the session carve and the per-stage prompt loans Two things a layer split was paying for and not using. * THE SESSION CARVE. A stage carved state for ALL 48 layers whatever slice it ran - every QSA KV pool at max_cells, every GDN recurrence - the whole-model-sized-buffer disease the chunked-QSA-prefill PR cured upstream. `session_bytes`/`session_init` take [layer_lo, layer_hi), and a session now owns only that range: `qsa_states`/`gdn_state` keep GLOBAL ordinal indexing and simply hold fewer rows (`qsa_ord0`/`qsa_alloc`, `gdn_ord0`/`gdn_alloc`), per-layer consumers subtract `gdn_ord0`, and the model-level uses - the shared RoPE table, the prefill staging identity, the MTP rope borrow - read `qsa_states[qsa_primary()]`. A range with no QSA layer still carves ONE primary state so those uses stay valid, and the default full range is byte-identical to the old carve. On the 2x3060 + 2x5060 rig, CUDA1's [30, 48) session costs 0.93 GiB: 7.83 GiB free before the carve, 6.90 GiB after (`session [30, 48) and the head`). * PER-STAGE PROMPT LOANS. `no_prefill_borrow = true` for split stages said "each stage's prompt path has its own buffers" - true, but that is a reason to give each stage its own LOAN, not to make every stage withhold a chunk-sized reserve from its cache for the whole session, and forced on it also collapsed `--prefill auto` to 2048. Each prompt path now borrows the tail of ITS OWN stage's cache (`PfPart`), so outside the prompt the whole cache is expert cache: `--prefill auto` is the largest chunk every participant can lend (8192 tokens here), a participant refills its loan into its own slots and marks only its own layers (a slot refilled into another cache, or another stage's rows marked resident, is silent and produces plausible tokens), and the chunk is the largest one EVERY participant can lend rather than one card's answer imposed on the others. A participant that cannot lend the buffers falls back to the prompt path's own, as a lone one always could; with a single participant the plan reduces to `plan_lend` exactly, so the single-GPU loan is unchanged from main - the percentage cap stays an AUTO-chunk rule (an explicit `--prefill` is the operator's number and a loan of it only has to leave the 128-slot floor). The flat 1 GiB withheld from every stage after the first becomes what is actually held back after the weights, session and drafter exist: 96 MiB of verify windows, plus 1000 MiB of drafter and head on the ONE stage that carries them. CUDA0 lends 2162 slots (3.98 GiB), CUDA1 1993 of its 2653 (3.97 GiB). * REPORTING. `expert_slots`/`expert_cache_mib` sum every tier: a 4-GPU 256K run reported CUDA0's cache alone (3327 experts) while the four cards held 13310, so the Monitor tab was wrong by 4x for every multi-GPU config. `expert_slots_primary`/`expert_cache_primary_mib` keep the old figure. The helper ranks' ms_begin/ms_wait are cumulative since boot and were printed next to per-request deltas; they take deltas now. KV-streaming stats iterate the session's OWNED ordinals. * REFUSALS. `--gpu-only-full` and `--gpu-stages` replay every layer through CUDA0's session, which a split carves to CUDA0's range, so they are refused with a layer split rather than answering about another stage's state; the token-graph hit path is disabled under a split (such a graph cannot span stages); and the sessions are allocated AFTER the split search, because the search prices each placement by what that range's session will cost on that device (`session_bytes` is pure arithmetic). 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.