Pull requests / #1748
#1748 Remote stage: run the later layers of a layer split on other PCs, over TCP (opt-in; Linux and Windows; Qwen3.8-Flash-Next)
open · @lask3802 · 0 comments · View on GitHub
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Description
## Summary A layer split already hands a token from one stage to the next once per verify window, and each prompt chunk's rows once per chunk. This branch puts that hand-off on a TCP link, so the later layers can run on cards in **other PCs**: - the main process runs layers [0, K), the head, the draft layer, sampling and the server: `--layer-split K --split-device 0 --remote-stage HOST:PORT`; - a worker process on another PC runs [K, 48) without the head: `strata --serve --stage-worker PORT --stage-begin K`; relay workers (`--stage-end K2 --stage-next HOST:PORT`) make three or more stages; - each process loads only its own layers' experts (a ranged expert arena) and dense weights, keeps an expert cache and an adaptive tier for its own layers, and its own part of the session; - the link builds on Linux (POSIX sockets) and Windows (Winsock); a Windows PC can be a worker natively or under WSL2; - `STRATA_REMOTE_TIMING=1` prints each stage's own time per prompt chunk and per window (a worker's split into GPU wait, CPU pool + plan and host staging); - three standard-library Python tools: a node agent that starts workers on each PC (`tools/stage_node.py`), a splitter that ships a node only its layers' bytes as sparse files (`tools/stage_ship.py`), and a tuner that measures configurations for real and writes the split, the prompt chunk and the configs (`tools/stage_tune.py`). Docs: `docs/REMOTE_STAGE.md` (use), `docs/remote-stage/ENGINEERING.md` (design, process, every number). Raw numbers: `bench/results/2026-10-10-remote-stage/`. About #1256: I saw that splitting a model across machines was closed there as not planned. This is fully opt-in and self-contained (a new source file, hooks that are inert by default, three tools, docs), and the measurements below show where it pays. If it is not wanted upstream, that is fine: I will keep maintaining it in my fork, and the Windows clock finding below may be useful on its own. ## Results PC A: RTX 3080 10 GB (PCIe 4.0 at x8), Ryzen 9 5900XT, 128 GB DDR4-3200, Linux. PC B: RTX 2080 Ti 11 GB, Ryzen 9 5900X, 64 GB DDR4-2400, Linux. PC C: RTX 3070 8 GB, Ryzen 7 5700X3D, 64 GB DDR4-3600, Windows 11. Links 1 GbE, and 10 GbE between A and B (~400 MB/s, slot-limited). `--kv int8 --kv-resident 32768 --max-context 131072 --spec 4`, greedy, one run per row. Prefill / decode tok/s; prompts of 7.5K / 30K / 57K / 91K tokens with 320-token answers. | configuration | short decode | 8K | 32K | 100K | |---|---:|---:|---:|---:| | RVN IQ3_S, 3080 alone | 39.9 | 791 / 38.1 | 1,118 / 37.2 | 1,120 / 38.0 | | RVN IQ3_S, A + B (1 GbE), K=26, chunk 5888 | 53.5 | 655 / 48.1 | 1,231 / 49.7 | 1,521 / 45.6 | | RVN IQ3_S, A + B + C (A-B 10 GbE), 24 / 16 / 8, C clocks locked | 51.3 | 658 / 47.8 | 1,110 / 46.5 | 1,640 / 44.7 | | UD-Q4_K_XL, 3080 alone (each number the better of two single-card runs: prefill with a RAM budget, decode without) | 27.3 | 302 / 25.4 | 324 / 24.5 | 326 / 23.7 | | UD-Q4_K_XL, A + B (10 GbE), K=26, chunk 5888 | 38.8 | 659 / 34.4 | 1,147 / 36.9 | 1,164 / 33.6 | | UD-Q4_K_XL, A + B + C, 24 / 16 / 8, C locked | 38.4 | 549 / 35.3 | 945 / 36.6 | 1,334 / 35.2 | | UD-Q4_K_XL, A + B + C, the tuner's pick (24 / 16 / 8 again of 8 splits; its own prompts, C with 1,248 cache slots) | 36.8 | 552 / 33.1 | 975 / 35.2 | 1,383 / 33.6 | - Two PCs: decode +18-44% on IQ3_S; on UD-Q4_K_XL, which one 10 GB card caches poorly (571-620 cache slots, auto chunk 1280-2048), decode +42% and a 100K prompt in 88 s instead of 293. - Long prompts need a large chunk (5888-8192): at 2048 every split read them slower than one card. - A third PC adds a hop to every window: it paid for long prompts (100K) and cost 1-8% of decode on IQ3_S. On UD-Q4_K_XL the tuner tried eight splits of the three stages (from 18 / 16 / 14 to 26 / 14 / 8): all within 45.7-48.0 s for a 32K request, a tie within the run-to-run spread. Over a 30% 8K / 50% 32K / 20% 100K workload (computed from the rows) two PCs take 46.3 s a request and three 47.6-49.1 s; one card ~133 s. **Windows lowers a worker GPU's clocks between decode windows.** A worker idles while the other stages run (25-45 ms a window). Sampled every 100 ms, the RTX 3070 under Windows spent 6% of a decode run in P2 (51% in P5, memory at 810 MHz) and waited 12-22 ms for its 8 layers per window; with `nvidia-smi -lgc 1500,1905 -lmc 7001,7001` it held P2 and waited 3.7-4.1 ms (short decode 33.8 -> 41.8). The 2080 Ti under Linux held P2 throughout (4.6-5.9 ms). WSL2 behaved the same as native Windows. The node agent can lock and reset the clocks around a worker (`"gpu_clocks"`, agent run as administrator) and the tuner's report says when a Windows node has not. This may also affect cards in one Windows PC with an in-process split; I have not measured that. ## What changed | file | change | |---|---| | `include/strata/net/stage_link.hpp`, `src/net/stage_link.cpp` (new) | the protocol (version 3: hello, window, commit, prompt chunk, reset, error), `StageClient`, `StageRelay`, `serve_stage` (a receiving thread, two input and two output buffers, replies on their own thread); POSIX and Winsock behind one small socket layer | | `src/program/generate.cpp` | the flags `--remote-stage`, `--stage-worker`, `--stage-begin`, `--stage-bind`, `--stage-end`, `--stage-next` and their checks; the layer range for the arena, dense weights, profile and session; the worker's serve loop and relay; the main process's head-only stage; the hello (geometry, K, K/V type, pack fingerprint, `STRATA_STAGE_TOKEN`); timing lines; the lend/refill code moved out of the request loop so the worker can call it | | `include/strata/core/verify.hpp`, `src/core/verify.cpp` | `Verifier::set_no_head` and `set_remote` | | `include/strata/prefill/prefill.hpp`, `src/prefill/prefill.cpp` | `Prefill::remote_send` / `remote_recv` (the pipelined chunk hand-off), `set_remote_rows_from`, `set_hand_in` / `set_single_chunk` | | `include/strata/core/expert_source.hpp`, `src/core/expert_source.cpp` | `ArenaExpertSource::set_layer_range` and a layer range in `load_experts_gguf` (buffered and Windows unbuffered readers) | | `CMakeLists.txt` | `src/net/stage_link.cpp` in the `strata` executable; `ws2_32` on Windows | | `tools/stage_node.py`, `tools/stage_ship.py`, `tools/stage_tune.py`, `tools/test_stage_node.py` (new) | the agent, the splitter, the tuner, their tests (`python -m unittest tools.test_stage_node`) | | `docs/REMOTE_STAGE.md`, `docs/remote-stage/ENGINEERING.md` (new), `docs/MULTI_GPU.md` | the docs; a pointer paragraph in MULTI_GPU.md | | `bench/results/2026-10-10-remote-stage/` (new) | every run's JSONL, the timing lines, the clock samples, the benchmark scripts | **What it leaves alone.** Without `--remote-stage` or `--stage-worker` no new code path runs: the new `Verifier`, `Prefill` and arena hooks are inert at their defaults (no range, no remote callbacks, `no_head` false). Checked: one single-GPU config (RVN IQ3_S on the RTX 3080, auto chunk, draft layer on) served by this branch's build twice and by `main`'s build (the merge base) once; greedy answers to five prompts (one of 9,364 tokens, two chunks) were byte-identical across the three (text, token counts, finish reasons). Not compared: an in-process multi-GPU split, other models, Windows against an upstream Windows build. With the flags, these are refused at start: `--batch`, `--pipeline-windows`, `--vision`, `--peer-device`, `--control-vector` (and the speed projection), `--expert-cache-remote`; the prompt cache, the conversation cache, mid-prompt checkpoints and `--kv-grow` are turned off. Not touched: the server, setup, the SYCL port's copies, the HIP-specific code. ## How it was tested - **Builds:** CUDA 13.0 on Linux (GCC, sm_75 + sm_86) and on Windows (MSVC 2022, sm_86); HIP (ROCm 7.0, gfx1100) and the SYCL port (oneAPI 2026.1.1; it compiles the shared `prefill.hpp` this branch changed) in containers. The remote stage itself ran on NVIDIA cards only. - **Two and three PCs** as above, two models, single-card baselines on the same day with the same build. The start of every benchmark answer (100-160 characters kept) was coherent and on topic in every configuration; full answers were not compared against the in-process split. - **Tools end to end:** layers shipped to a Linux host, a WSL2 VM and native Windows, workers started from the shipped sparse files by the agents (one agent serving two models), the Windows agent locking and resetting the clocks (run as administrator during the benchmarks; the final version also run without, where it starts the worker and says "NOT locked"), and the tuner run on two PCs (RVN IQ3_S: it picked the same split as the hand sweep) and three PCs (UD-Q4_K_XL). - `python -m unittest tools.test_stage_node`: 21 tests, on Windows (ReFS) and Linux. - **Probes:** a wrong token is refused; a peer that sends no hello is closed after 10 s. ## Not tested - Running on AMD or Intel cards (HIP and SYCL were built, not run); the SYCL port keeps its own copies of the engine files and has no remote stage. - Windows as the main PC; an in-process split on Windows with locked clocks. - Other quantizations of the family (official IQ3_S, IQ3_XXS, IQ2_XS, UD-IQ4_XS); a pack with `experts.bin` is refused at start. - Correctness against the in-process split with fixed caches; more than one run per configuration (decode moved by up to ~14% between two runs of one configuration). - Reconnecting without restarting the main engine; a ring topology; a bf16 hand-off. ## Extra notes - Plain TCP, not encrypted: token ids and hidden states cross the network in clear. The shared token only keeps other LAN hosts from driving a worker. The docs say to use it on a trusted LAN or inside a tunnel. - The node agent's token lets a request start the engine with whitelisted flags and write files under its data directory: a credential, LAN only. - One feature in one PR; I can split it into the engine hooks, the network stage and the tools if that is easier to review. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.