Pull requests / #716
#716 conversation parking with --layer-split
closed · @evaanp · 0 コメント · GitHub で見る
Multi-GPUNVIDIA / CUDAModels & quantsLinux
本文
## What Conversation parking (`--conversation-cache-mib`, #57) refused to start with `--layer-split`: "conversation parking does not yet support --layer-split". On two cards every switch between an agent and its sub-agent or side call then read the whole prompt again. This lets the two work together. - **One part per stage.** A snapshot holds every stage's running state, used K/V pages and checkpoints, and the draft-layer K/V on the stage that carries it. Host K/V stays flat, in stage order with the drafter last, so the cache's reuse of retained pages works as before; only running state is per stage, the way `ConversationCheckpoint::stage_parts` already stored it. - **Each part names its carve.** `ConversationCheckpoint` records `layer_lo`/`layer_hi`: two same-sized carves with different QSA ordinals cannot be told apart by sizes alone, so a part can never be restored into a stage that runs other layers. - **Validate everything, then write.** `conversation_snapshot_validate` checks every part against its own stage before anything is restored, as the single-GPU path already did. - **Each stage on its own device.** Capture and restore drain each stage inside its own `OnDevice` scope and copy there. (`cudaDeviceSynchronize` drains only the current device, so a single sync on stage 0's card would have read a later stage's state while its kernels might still be writing it.) Stage 0 is named as device 0 instead of -1. - **Split checkpoints taken mid-prompt** (`--prompt-cache-every`) now copy the carve too; without it every park of a conversation longer than 16K tokens was refused as "a checkpoint from another session layer carve". - `tools/conversation_cache_parity.py` builds the engine arguments the way the server does, so on several GPUs it tests the layer split instead of one card. ## Tests - `conversation_validation_test` gains a two-stage split fixture (host only): parts, carves, mismatched stage counts and carves, and two checks that a later stage is written, and read during capture, only after its own device was drained. 1008 checks host-only; `conversation_transfer_test` 1608. That test now also wraps `cudaGetDevice` and `cudaSetDevice` (it defined the wrappers, but the link did not use them, so on a machine without a GPU the device never changed and the fixture failed on its own harness); with them it passes, and it fails again with the per-stage `OnDevice` scopes removed. Built in the CUDA 13.0 image (`-DSTRATA_BUILD_CONVERSATION_TESTS=ON`). - `conversation_cache_test` (4175 checks) unchanged; `conversation_snapshot_test` call sites take a stage vector. - On GPUs: `tools/conversation_cache_parity.py --scenario reuse` with the layer split: A/B/A, greedy, tokens and main-model state byte-exact against a run without parking (see Measured). ## Measured IQ3_S on an RTX 5060 Ti 16 GB + RTX 2000 Ada 16 GB (`layer_split 30`), Ryzen 7 9700X, 64 GB RAM. Measured on our own branch: v0.1.38 merged with this code plus unrelated patches of ours (the parking code is the same as here), `--conversation-cache-mib 1024`. `--scenario reuse`: PASS, byte-exact. A ~2K-token conversation is a 254 MiB snapshot across both stages, parked in ~55 ms and restored in ~18 ms (from its live end) to ~20 ms (from a checkpoint). The same A/B/A also passed on IQ2_XS with the earlier 0.1.34-based version of this code. Developed with an AI coding assistant; all numbers measured on the machine above (Linux, CUDA 13).
関連リンク
インストール・モデル・リリースへの站内リンク。