Pull requests / #653

#653 serve: park conversations with a layer split

closed · @anon761 · 0 commentaires · Sur GitHub

Server & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentation

Description

## What

Conversation parking (`--conversation-cache-mib`) refuses `--layer-split` at start:

```
strata serve: conversation parking does not yet support --layer-split; disable parking with --conversation-cache-mib 0
```

On a split every stage holds its own carve of the session, so a parked conversation becomes one image per stage:

- stage 0's image carries the token ids the cache matches on; the later stages' images ride in its new `stage_parts`
- each part is captured, validated and restored on its own GPU, with its own running state, checkpoint states and K/V; the draft layer's K/V goes with the last stage, which holds it
- the checkpoint chain is cut into per-stage views for the capture and put back together (`stage_parts`) on restore
- **every** part is validated before any is restored, so a split image is never half applied

The snapshot core takes the draft as optional (a pointer); the existing reference signatures stay and refuse an image with stage parts, so the single-GPU path is unchanged. Split parks are always full captures (the retained-K/V reuse stays single-GPU). The one-GPU check of the hand-off (`--split-device 0`) shares one session between its stages and still refuses parking. `docs/DETAILS.md` and `--help` say so.

## Tests

`tools/conversation_cache_parity.py` with `--layer-split auto` on 2x RTX 3090, Unsloth UD-Q4_K_XL:

```
reuse,    --spec 1:                     PASS: A/B/A output and checkpoint reuse, byte-exact main-model state
reuse,    --spec 4:                     PASS: A/B/A output and checkpoint reuse
pressure, --spec 1, --cache-mib 300:   PASS: pressure fallback, output and byte-exact main-model state
```

`conversation_{cache,memory,snapshot,validation,transfer}_test` pass (`-DSTRATA_BUILD_CONVERSATION_TESTS=ON`).

## Measured

Same box, 262K context, two coding conversations of ~11K tokens taking turns (A, B, A, B, ...), 4 turns each over the OpenAI endpoint. `main` has to run without parking on a split:

| | follow-ups resumed | follow-up prompt time (median) |
| --- | ---: | ---: |
| `main`, split, no parking | 0 of 6 | 5.8 s |
| this branch, split, `--conversation-cache-mib 65536` | 6 of 6 | 0.64 s |

A ~11K-token conversation parks in 360-400 ms (two stages, ~640-770 MB) and is restored in 41-47 ms.

Sur le site

Liens install, modèles, releases.