Pull requests / #716

#716 conversation parking with --layer-split

closed · @evaanp · 0 Kommentare · Auf GitHub

Multi-GPUNVIDIA / CUDAModels & quantsLinux

Beschreibung

## What

Conversation parking (`--conversation-cache-mib`, #57) refused to start with `--layer-split`: "conversation parking
does not yet support --layer-split". On two cards every switch between an agent and its sub-agent or side call
then read the whole prompt again. This lets the two work together.

- **One part per stage.** A snapshot holds every stage's running state, used K/V pages and checkpoints, and the
  draft-layer K/V on the stage that carries it. Host K/V stays flat, in stage order with the drafter last, so the
  cache's reuse of retained pages works as before; only running state is per stage, the way
  `ConversationCheckpoint::stage_parts` already stored it.
- **Each part names its carve.** `ConversationCheckpoint` records `layer_lo`/`layer_hi`: two same-sized carves with
  different QSA ordinals cannot be told apart by sizes alone, so a part can never be restored into a stage that runs
  other layers.
- **Validate everything, then write.** `conversation_snapshot_validate` checks every part against its own stage
  before anything is restored, as the single-GPU path already did.
- **Each stage on its own device.** Capture and restore drain each stage inside its own `OnDevice` scope and copy
  there. (`cudaDeviceSynchronize` drains only the current device, so a single sync on stage 0's card would have read
  a later stage's state while its kernels might still be writing it.) Stage 0 is named as device 0 instead of -1.
- **Split checkpoints taken mid-prompt** (`--prompt-cache-every`) now copy the carve too; without it every park of
  a conversation longer than 16K tokens was refused as "a checkpoint from another session layer carve".
- `tools/conversation_cache_parity.py` builds the engine arguments the way the server does, so on several GPUs it
  tests the layer split instead of one card.

## Tests

- `conversation_validation_test` gains a two-stage split fixture (host only): parts, carves, mismatched stage counts
  and carves, and two checks that a later stage is written, and read during capture, only after its own device was
  drained. 1008 checks host-only; `conversation_transfer_test` 1608. That test now also wraps `cudaGetDevice` and
  `cudaSetDevice` (it defined the wrappers, but the link did not use them, so on a machine without a GPU the device
  never changed and the fixture failed on its own harness); with them it passes, and it fails again with the
  per-stage `OnDevice` scopes removed. Built in the CUDA 13.0 image (`-DSTRATA_BUILD_CONVERSATION_TESTS=ON`).
- `conversation_cache_test` (4175 checks) unchanged; `conversation_snapshot_test` call sites take a stage vector.
- On GPUs: `tools/conversation_cache_parity.py --scenario reuse` with the layer split: A/B/A, greedy, tokens and
  main-model state byte-exact against a run without parking (see Measured).

## Measured

IQ3_S on an RTX 5060 Ti 16 GB + RTX 2000 Ada 16 GB (`layer_split 30`), Ryzen 7 9700X, 64 GB RAM. Measured on our
own branch: v0.1.38 merged with this code plus unrelated patches of ours (the parking code is the same as here),
`--conversation-cache-mib 1024`. `--scenario reuse`: PASS, byte-exact. A ~2K-token conversation is a 254 MiB snapshot
across both stages, parked in ~55 ms and restored in ~18 ms (from its live end) to ~20 ms (from a checkpoint). The
same A/B/A also passed on IQ2_XS with the earlier 0.1.34-based version of this code.

Developed with an AI coding assistant; all numbers measured on the machine above (Linux, CUDA 13).

Mehr auf der Site

Links zu Install, Modellen, Releases.