Pull requests / #576

#576 multi-GPU: conversation parking across a layer split

closed · @saikiran-rs · 0 comments · View on GitHub

Multi-GPUNVIDIA / CUDAModels & quants

Description

--conversation-cache-mib refused to start with --layer-split. A parked conversation now carries one SavedConversationStage per later stage: that stage's running state (its carve) and the K/V of the QSA layers it owns, captured and restored on its own device (OnDevice). The draft layer stays the last kv of the first stage's image and is copied on the last stage's GPU. Split checkpoints keep their stage_parts through parking and are validated per stage. The one-GPU API is unchanged (overloads with an empty stage list); retained K/V reuse stays one-GPU, a split always captures in full.

Test: conversation_snapshot_test split_session (two carves on one device, F16/INT8/Q4_0 x resident/streamed K/V): the capture stays within its estimate; a one-GPU restore, another layer range and a damaged stage are refused before any write; A/B/A is byte-exact on every stage and the draft; the PLE window and the checkpoint parts survive a restore.

Measured on two RTX 3060 12 GB (x16 + x8), UD-Q4_K_XL, 262K context, 16 GiB cache: two chats of 55K and 47K tokens asked in turn; the follow-ups' time to the first token fell from 59.4 s / 52.2 s to 1.98 s / 1.85 s (restore 0.47-0.49 s), same answers.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.