Pull requests / #653
#653 serve: park conversations with a layer split
closed · @anon761 · 0 commentaires · Sur GitHub
Server & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentation
Description
## What
Conversation parking (`--conversation-cache-mib`) refuses `--layer-split` at start:
```
strata serve: conversation parking does not yet support --layer-split; disable parking with --conversation-cache-mib 0
```
On a split every stage holds its own carve of the session, so a parked conversation becomes one image per stage:
- stage 0's image carries the token ids the cache matches on; the later stages' images ride in its new `stage_parts`
- each part is captured, validated and restored on its own GPU, with its own running state, checkpoint states and K/V; the draft layer's K/V goes with the last stage, which holds it
- the checkpoint chain is cut into per-stage views for the capture and put back together (`stage_parts`) on restore
- **every** part is validated before any is restored, so a split image is never half applied
The snapshot core takes the draft as optional (a pointer); the existing reference signatures stay and refuse an image with stage parts, so the single-GPU path is unchanged. Split parks are always full captures (the retained-K/V reuse stays single-GPU). The one-GPU check of the hand-off (`--split-device 0`) shares one session between its stages and still refuses parking. `docs/DETAILS.md` and `--help` say so.
## Tests
`tools/conversation_cache_parity.py` with `--layer-split auto` on 2x RTX 3090, Unsloth UD-Q4_K_XL:
```
reuse, --spec 1: PASS: A/B/A output and checkpoint reuse, byte-exact main-model state
reuse, --spec 4: PASS: A/B/A output and checkpoint reuse
pressure, --spec 1, --cache-mib 300: PASS: pressure fallback, output and byte-exact main-model state
```
`conversation_{cache,memory,snapshot,validation,transfer}_test` pass (`-DSTRATA_BUILD_CONVERSATION_TESTS=ON`).
## Measured
Same box, 262K context, two coding conversations of ~11K tokens taking turns (A, B, A, B, ...), 4 turns each over the OpenAI endpoint. `main` has to run without parking on a split:
| | follow-ups resumed | follow-up prompt time (median) |
| --- | ---: | ---: |
| `main`, split, no parking | 0 of 6 | 5.8 s |
| this branch, split, `--conversation-cache-mib 65536` | 6 of 6 | 0.64 s |
A ~11K-token conversation parks in 360-400 ms (two stages, ~640-770 MB) and is restored in 41-47 ms.
Sur le site
Liens install, modèles, releases.