Pull requests / #526

#526 Allow conversation parking under --layer-split (owner-routed per-stage capture/restore)

closed · @chimpera · 0 コメント · GitHub で見る

Server & APIMulti-GPUNVIDIA / CUDA

本文

## The problem

Today `--layer-split` and `--conversation-cache-mib` cannot be combined — the server refuses to start:

```
strata serve: conversation parking does not yet support --layer-split; disable parking with --conversation-cache-mib 0
```

(`src/program/generate.cpp`, the `return 2` guard.) A multi-GPU split therefore cannot park conversations at all: every conversation switch re-prefills the whole session, exactly the cost parking exists to avoid.

## The change

Teach the snapshot core which GPU stage owns each QSA layer:

- **`conversation_snapshot.hpp`**: `ConversationStageRef {dev, lb, le, ss}` and a trailing `later` parameter (default `{}`) on `bytes`/`capture_bytes`/`save`/`validate`/`restore` — every single-GPU caller and test compiles unchanged.
- **`conversation_state.cpp`**: `owner_of`/`owner_state` map a QSA layer to its stage via `[lb, le)`; per-layer K/V saves/restores go to/from the **owner's** session under an `OnDevice` guard (the primary session's buffers for another stage's layers are allocated but never written). The live checkpoint gains one running-state part per later stage, assembled exactly as the ordinary split checkpoint path does; the draft rides the last stage. Validation requires `stage_parts.size() == later.size()`.
- **`generate.cpp`**: drop the `--layer-split` parking rejection; build the stage refs from the `GpuStage` table once beside the cache and pass them at the estimate/save/validate/restore sites; the `STRATA_SNAPSHOT_VERIFY` read-back runs on the draft's (last stage's) device.

## Validation

Measured on a 2-way split (RTX 3090 ×2, `--layer-split 24`):

- host tests green (1781 + 876 + 1452 checks)
- park 1593 tok / 57 ms and 19,925 tok / 125 ms (1.01 GB)
- restore byte-verified (`STRATA_SNAPSHOT_VERIFY`)
- incremental re-park reuses 23.5 MB of retained K/V through the owner routing
- two conversations parked simultaneously
- exact-token parity of A1/A2/A3 answers between the restore path and a parking-off fresh run

The same change has been running in a daily-driver split deployment with the conversation cache on (parking, restore and incremental re-park exercised by live traffic). Python suite on this branch: 100 tests OK (`python -m unittest serve.test_server`).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

関連リンク

インストール・モデル・リリースへの站内リンク。