Pull requests / #526
#526 Allow conversation parking under --layer-split (owner-routed per-stage capture/restore)
closed · @chimpera · 0 Kommentare · Auf GitHub
Server & APIMulti-GPUNVIDIA / CUDA
Beschreibung
## The problem
Today `--layer-split` and `--conversation-cache-mib` cannot be combined — the server refuses to start:
```
strata serve: conversation parking does not yet support --layer-split; disable parking with --conversation-cache-mib 0
```
(`src/program/generate.cpp`, the `return 2` guard.) A multi-GPU split therefore cannot park conversations at all: every conversation switch re-prefills the whole session, exactly the cost parking exists to avoid.
## The change
Teach the snapshot core which GPU stage owns each QSA layer:
- **`conversation_snapshot.hpp`**: `ConversationStageRef {dev, lb, le, ss}` and a trailing `later` parameter (default `{}`) on `bytes`/`capture_bytes`/`save`/`validate`/`restore` — every single-GPU caller and test compiles unchanged.
- **`conversation_state.cpp`**: `owner_of`/`owner_state` map a QSA layer to its stage via `[lb, le)`; per-layer K/V saves/restores go to/from the **owner's** session under an `OnDevice` guard (the primary session's buffers for another stage's layers are allocated but never written). The live checkpoint gains one running-state part per later stage, assembled exactly as the ordinary split checkpoint path does; the draft rides the last stage. Validation requires `stage_parts.size() == later.size()`.
- **`generate.cpp`**: drop the `--layer-split` parking rejection; build the stage refs from the `GpuStage` table once beside the cache and pass them at the estimate/save/validate/restore sites; the `STRATA_SNAPSHOT_VERIFY` read-back runs on the draft's (last stage's) device.
## Validation
Measured on a 2-way split (RTX 3090 ×2, `--layer-split 24`):
- host tests green (1781 + 876 + 1452 checks)
- park 1593 tok / 57 ms and 19,925 tok / 125 ms (1.01 GB)
- restore byte-verified (`STRATA_SNAPSHOT_VERIFY`)
- incremental re-park reuses 23.5 MB of retained K/V through the owner routing
- two conversations parked simultaneously
- exact-token parity of A1/A2/A3 answers between the restore path and a parking-off fresh run
The same change has been running in a daily-driver split deployment with the conversation cache on (parking, restore and incremental re-park exercised by live traffic). Python suite on this branch: 100 tests OK (`python -m unittest serve.test_server`).
🤖 Generated with [Claude Code](https://claude.com/claude-code)Mehr auf der Site
Links zu Install, Modellen, Releases.