Pull requests / #1055
#1055 two-PC: a second PC runs a block of layers as a stage of the layer split (--stage-server / --remote-stage, decode only)
closed · @mirifiuto135-debug · 0 Kommentare · Auf GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAWindowsLinux
Beschreibung
## What A second PC can run a block of layers as one more stage of the existing layer split, over plain TCP. Decode only (v1). Two sides, one flag each, in `src/program/generate.cpp` + a new `src/core/stage_link.cpp` / `include/strata/core/stage_link.hpp`: - **Stage engine**: `--serve --stage-server HOST:PORT --stage-layers A,B` loads and runs layers `[A,B)` only (its own pack copy, dense weights, experts, session), no head, no drafter, no HTTP of its own. It dials the driver and answers RUN / COMMIT / ZERO for each verify window. With several cards an explicit `--layer-split M` (`A < M < B`) splits the stage's layers across them. - **Driver**: `--serve --remote-stage PORT --remote-layers A,B --layer-split B,…` keeps layers `[0,A)` and `[B,n)`, the head, sampling and the MTP drafter; its stage chain gets one more hop in the middle. The driver's port only accepts the stage; a stranger connecting first is dropped and the driver keeps waiting. - The hop is a `StageBridge` hook between two `Verifier` stages (`include/strata/core/verify.hpp`): the existing per-window hand-off buffer (host-pinned, ~51 KB per token) is what crosses the cable, in both directions. A loopback bridge (`STRATA_BRIDGE_LOOPBACK=N`, memcpy only) exists to test the hook on one PC. - Refusals are explicit: the two sides cannot be combined, the split must lie inside the stage's layers, `--layer-split auto` is not accepted on the stage. v1 limits, enforced, not hidden: prompts go through verify windows only (no `Prefill` chaining over the cable yet), checkpoints and the conversation cache are off while a remote stage is attached. ## Why A model whose experts do not fit one PC's RAM. The other PC's cards hold its layers' experts; the driver's RAM and cache cover fewer layers. It does not make decode faster (measured below); it moves RAM pressure and raises the driver's card hit rate. ## Measured Windows 11 both sides, direct 2.5 GbE cable. Driver = 3× RTX 5060 Ti 16 GB, 128 GB RAM; stage = 1 or 2× RTX 5060 Ti 16 GB, 64 GB RAM. Qwen3.8 Flash-Next UD-Q6_K_XL converted pack (experts Q8_0), `STRATA_ARENA_MMAP=1`, MTP on, three short prompts, 150 tokens each, temperature 0: - loopback bridge on one PC (hop 0, hop 1) vs the plain split: tokens identical character for character - cable cost: 256 KB there and back = 2.34 ms median (200 round trips, TCP_NODELAY) - driver alone (3 cards): 20.0 / 22.0 / 26.8 tok/s warm - stage holds 8 layers on one card: 21.6 / 22.0 / 24.7 tok/s — same band - stage holds 12 layers on one card: 14.0 / 15.3 / 20.4 tok/s — slower (the card covers 41 % of those experts; the stage's CPU computes the rest) - stage holds 16 layers on two cards (`--layer-split`): 19.1 / 17.6 / 24.8 tok/s — same band, not faster; the stage spent 54.9 ms per window - driver RAM free while serving: 40 GiB alone → 58–75 GiB with a stage; driver expert-cache hit rate 70–80 % Why not faster: the layers run one after the other whichever PC holds them; the window time is set by the experts the cards miss, and the stage's CPU is not faster than the driver's. Two PCs buy RAM headroom, not speed, with this split. ## Not in this PR Prompt path over the cable (a 2048-token chunk is ~105 MB each way; estimated +16 % on prefill at 2.5 GbE), checkpoints / conversation cache with a remote stage, the stage loading only what it uses (today it still loads the PLE table and the head). AMD, Linux and single-PC paths are untouched; `STRATA_ARENA_MMAP` on Windows is PR #1002, independent of this one. Rebased from 0.1.39 onto v0.1.40 for this PR (three conflict regions in `generate.cpp`: the `--kv-grow` block, `--mtp` becoming optional, `--pipeline-windows`); it builds on v0.1.40 with CUDA 13.3 for sm_89 + sm_120. The measurements above are from the 0.1.39 build; I have not re-run the two-PC test on the v0.1.40 build yet and will say so if asked. ## Disclosure The code was written with Claude Code (Anthropic) under my direction; I built and tested it on the hardware above. Happy to adjust anything, or to split it into the bridge hook and the stage/driver halves if that reads better.
Mehr auf der Site
Links zu Install, Modellen, Releases.