Pull requests / #482

#482 Networked Pool support

closed · @fragtion · 0 comentarios · En GitHub

BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindowsLinux

Descripción

## Title

Pool: use several PCs together — share requests between them, or split one model's layers across them

## Description

Hi Niko, this adds an optional **Pool** tab that lets two or more PCs on a LAN work together. Nothing changes for
anyone who doesn't use it: an engine started without `--pool-*` flags and a server whose pool role is *Off* behave
exactly as before. I've been running it on my own two machines (RTX 3060 12 GB desktop + RTX 5060 8 GB laptop, both
on the Coder IQ1_M). Please treat it as a proposal: take it, take parts of it, or ask for changes.

### What it does

There are two modes, chosen per PC in the new Pool tab:

| | **Share requests** (routing) | **Split layers** (coordinator + workers) |
| --- | --- | --- |
| Each PC holds | the whole model, as it would alone | only its own layers (experts, dense weights, context) |
| One chat | as fast as one PC | about as fast as one PC (the PCs take turns on every token) |
| Two chats at once | both at full speed, one per PC | one waits |
| What it adds | throughput for agents and several users, and a spare PC | room: experts, context or a model that one PC can't hold |
| Engine changes needed | none (server only) | yes (`--pool-peers` / `--pool-listen`) |

The tab says plainly that neither mode adds up the PCs' tok/s, because that is the first thing people expect.
`docs/POOL.md` explains both modes in the same style as `MULTI_GPU.md`.

**Share requests** (`serve/route.py`): each PC runs Strata as usual. A chat request that reaches any PC runs on the PC
that already holds that conversation, recognised from a hash chain over its messages (clients' moving
`cache_control` markers are ignored), so the prompt cache and parked chats keep working. Otherwise it runs on an idle
PC, preferring the one marked *primary*. A conversation moves off a busy PC only if the other PC would re-read at
most N tokens (default 8,000); a long one waits for its own PC. A PC whose context window is too small for the
request is skipped. If the PCs lose touch, each carries on alone and checks for the others every 3 s. Streams are
relayed as they arrive, and a client that hangs up cancels the request on the other PC (same idea as #430).
Requests between PCs carry an HMAC (pool secret over time, nonce, path and body hash), which replaces the
receiving PC's API key for those requests only.

**Split layers**: pipeline parallelism with the same cut as the multi-GPU layer split, over TCP instead of pinned RAM.
- The coordinator runs layers `[0, K)` plus the head, sampler and MTP draft layer. Each worker runs a later range.
- Each PC loads only its own layers' experts (`ArenaExpertSource::set_layer_range`), dense weights (a `NativeDense`
  skip list) and KV/recurrent state. A PC whose RAM can't hold the whole model can still take part.
- A verify window crosses the network once out and once back (12,804 floats per token), never per layer. Prompt
  chunks go through the workers like a local next stage.
- `--pool-split auto` reuses the multi-GPU cost model, adding a network hop per PC and each PC's RAM and VRAM.
- The protocol (`include/strata/pool/protocol.hpp`):
  - a 24-byte framed header, then `HELLO`, then mutual HMAC-SHA256 auth (the secret never crosses the wire);
  - `CONFIG` / `READY`, with a model fingerprint check;
  - `VERIFY` / `PREFILL` rows in f32 (exact), f16 or bf16;
  - lazily acknowledged `COMMIT`, `RESET` and checkpoint messages.

  A worker keeps its layers loaded between coordinators, and answers a second coordinator with "busy".
- Per-request log line on the coordinator: window time split into this PC, the workers' layers, and the network.

The server side (`serve/pool.py`):
- roles kept in `<config>.pool.json`, switched without restarting the server;
- a supervisor that keeps the worker engine running;
- LAN discovery beacons (UDP 7702) so the tab can list other PCs.

`POOL-FIREWALL.bat` opens the ports on private networks.

### Not supported in Split mode yet (refused with a clear message)

- vision
- control vectors
- the low-RAM modes
- `--shared-expert-arena`
- several GPUs in one pool PC
- conversation parking: per-chat checkpoints (`--prompt-cache`) do work across the pool

Share requests has none of these limits.

### Commits

1. `build: MSVC compiler probe without debug records; STRATA_RELEASE_PDB option`. A small, independent Windows
   build fix: CMake's compiler probe fails with D8050 on PCs whose process monitors block cl.exe debug records. Happy
   to split it into its own PR.
2. `engine: pool - split one model's layers across PCs`: `src/pool`, `include/strata/pool`, and hooks in
   Verifier, Prefill, the expert source, NativeDense and generate.cpp (marked `POOL:`).
3. `serve: Pool tab - share requests ... or split layers ...`: `serve/pool.py`, `serve/route.py`, the tab, and
   the docs.

The diff is large (about 6,800 lines added), but most of it is new files. The changes to existing engine code are
small, opt-in hooks.

### Testing

- New tests that need no GPU:
  - `strata-pool-test`: hashing, wire formats, frames, handshake, split search.
  - `strata-pool-link-test`: the coordinator against scripted workers (wrong secret, other model, busy worker,
    reload, two workers, both wire formats, checkpoints).
  - `serve/test_pool.py`: config, supervisor, discovery, role switching, and two routing servers covering
    stickiness, the primary, context limits, a streamed relay, signatures, and a peer going away.
- The full existing Python suite passes. Engine and tests build with GCC 13 / CUDA 13 on Linux, and with MSVC +
  CUDA 13.0 on Windows (one engine for sm 75;86;89;120).
- Real use, two PCs on gigabit LAN, Coder IQ1_M:
  - **Share requests**: in daily use with a coding agent harness whose subagents run on both PCs at once, each at
    single-PC speed.
  - **Split**: works end to end. With the desktop holding layers 0–34 and the laptop 35–47, the desktop's expert
    hit rate went from 55–60% to 70–78%. In that run the laptop's cache was mis-sized, so tok/s stayed level;
    commit 2 makes a pool PC always size its cache to its free VRAM. I haven't re-measured since.
- **Not tested**: AMD/HIP builds. `src/pool/link.cpp` uses only CUDA runtime calls that `hip_compat` covers, but I have
  no AMD card to check.

### Trying it

Two PCs with the same model:
1. Run `POOL-FIREWALL.bat` once on each.
2. Pool tab on each PC: choose *Share requests*, give both the same secret, and add the other PC (it appears under
   *Found on your network*).
3. For Split, choose *Split: worker* on one PC and *Split: coordinator* on the other.

`POOL-SELFTEST.bat` checks the split path on one PC against the model alone.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_0157fjixQD2nRxvCvtXq66DD

En el sitio

Enlaces a install, modelos, releases.