Issues / #747

#747 Feature request: multi-conversation parking on multi-GPU layer-split

closed · @AdaxLabs · 1 コメント · GitHub で見る

Setup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux

本文

<h2><span>Feature request</span></h2><p><span>Please extend Strata's </span><strong><span>bounded multi-conversation cache / parking</span></strong><span> to work with the existing </span><strong><span>multi-GPU layer-split</span></strong><span> serving path.</span></p><p><span>This is the remaining gap after #57 / #189: same-conversation checkpoints work very well on a split, and bounded multi-conversation parking works very well on one GPU, but the two features cannot currently be combined. The current docs explicitly state that enabling parking with </span><code dir="ltr">--layer-split</code><span> is rejected before model loading.</span></p><p><span>We are </span><strong><span>not</span></strong><span> asking for concurrent inference. Requests can remain serial. The request is only to preserve and restore the complete state of several alternating conversations when one Strata model is split across multiple GPUs.</span></p><h2><span>Why this matters for agent workloads</span></h2><p><span>A typical controller/worker sequence is:</span></p><div><pre dir="ltr"><code>Parent A
  -&gt; Worker B
  -&gt; Parent A again</code></pre></div><p><span>On a single-GPU Strata server, bounded conversation parking can restore A after B and avoid rereading A's full history.</span></p><p><span>On a layer-split server, the parent can reuse its own prior turn if no unrelated conversation intervenes, but an intervening worker causes the parent to be read cold again because multi-conversation parking is unavailable.</span></p><p><span>That makes the current multi-GPU path excellent for one long conversation, but expensive for agent harnesses that alternate controller and worker conversations.</span></p><h2><span>Reproduction hardware / software</span></h2><p><span>Linux host:</span></p><ul><li><span>2 × NVIDIA RTX 5060 Ti 16 GB</span></li><li><span>64 GB system RAM</span></li><li><span>Qwen3.8-Flash-Next IQ3_XXS</span></li><li><span>Strata dual-GPU layer split, K=22</span></li><li><span>256K context</span></li><li><span>prompt cache enabled for the interactive arm</span></li></ul><p><span>The same machine was also tested in single-GPU mode as a control.</span></p><p><span>A separate RTX 3060 12 GB single-GPU host was used to verify that the current multi-conversation cache behaves as expected when no layer split is involved.</span></p>

</span></p><h2><span>Measured behavior</span></h2><p><span>

### 1. Dual-GPU layer split: same conversation works extremely well

With the cache disabled, the second ~86K-token turn took:

| profile | A2 wall | reused |
|---|---:|---:|
| cache off | 69.705 s | 0 |

With normal prompt reuse enabled on the same dual-GPU K22 topology:

| profile | A2 wall | reused |
|---|---:|---:|
| prompt cache on | **1.380 s** | **86,340 tokens** |

So the ordinary layer-split checkpoint/reuse path itself is working very well.

### 2. Dual-GPU layer split: A -> B -> A loses the parent

With the same dual-GPU interactive profile:

| turn | wall | prompt tokens | reused |
|---|---:|---:|---:|
| Parent A1 | 70.582 s | 86,343 | 0 |
| Worker B1 | 14.150 s | 14,533 | 0 |
| Parent A2 | 70.045 s | 86,376 | **0** |
| **Total** | **154.777 s** | | |

The returned parent is effectively cold again.

Attempting to enable the bounded multi-conversation cache on this K22 layer-split profile is rejected, matching the documented current limitation.

This contrast is the main point of the report:

- same conversation on dual GPU: **~69.7 s -> 1.38 s**
- Parent A -> Worker B -> Parent A on dual GPU: **Parent A returns cold at ~70.0 s**

### 3. Single-GPU control: multi-conversation parking restores A

On a single RTX 5060 Ti, parking works:

- A -> B -> A restore: **PASS**
- returned A reused: **86,336 tokens**
- returned A wall: **1.587 s**

However, this is not a practical substitute for the split on this hardware:

- cold single-GPU parent: about **267 s**
- Parent -> Worker -> Parent total: about **579 s**

So dropping to one GPU solely to gain parking gives up too much cold/prefill performance.

### 4. Independent single-GPU control on RTX 3060

On a separate RTX 3060 12 GB server:

- same-chat cached A2: **4.676 s**, ~99.98% reuse
- A -> B -> A with bounded parking: returned A **5.274 s**, **28,918 reused tokens**

This confirms that the multi-conversation feature itself solves the controller/worker pattern; the missing piece is compatibility with the layer-split state.</span></p><h2><span>Requested upstream capability</span></h2><p><span>Ideally the existing </span><code dir="ltr">--conversation-cache-mib</code><span> / </span><code dir="ltr">--conversation-cache-slots</code><span> feature would support a layer-split model without changing the client contract.</span></p><p><span>Conceptually:</span></p><ol start="1"><li><span>Capture the complete conversation state for </span><strong><span>every stage/device</span></strong><span> participating in the split.</span></li><li><span>Treat capture/restore as one atomic logical snapshot.</span></li><li><span>Match the same exact token/image/steering prefix rules already used by the current cache.</span></li><li><span>Restore all per-device session state before resuming generation.</span></li><li><span>Fail closed to an ordinary prompt reread if any stage is incompatible or cannot be restored.</span></li><li><span>Keep RAM/slot admission bounded as today.</span></li><li><span>Preserve the current serial request model; no cross-request batching or concurrent decode is required.</span></li></ol><p><span>I would prefer this to be implemented in upstream Strata rather than maintaining a private fork, because this touches the most correctness-sensitive session/KV/indexer/draft state and should evolve with the engine's own layer-split implementation.</span></p><h2><span>Suggested acceptance gate</span></h2><p><span>A useful minimum regression would be a real layer-split:</span></p><div><pre dir="ltr"><code>A1 -&gt; B1 -&gt; A2</code></pre></div><p><span>with:</span></p><ul><li><span>A2 restoring a substantial exact prefix from A1 instead of rereading from token 0;</span></li><li><span>exact/validated restored state across every GPU stage;</span></li><li><span>correct output versus an uninterrupted/cold reference;</span></li><li><span>bounded cache accounting including all per-device state;</span></li><li><span>safe fallback to a cold read on incompatibility/admission failure;</span></li><li><span>existing single-conversation layer-split checkpoints unchanged when multi-conversation parking is disabled.</span></li></ul><p><span>Longer-term coverage for KV streaming, images, steering and supported KV formats would be valuable, but the core ask is the multi-GPU snapshot/restore path.</span></p><h2><span>Relationship to existing work</span></h2><ul><li><span>#57 introduced the multi-conversation use case.</span></li><li><span>#189 shipped the bounded RAM conversation cache and explicitly listed whole-conversation multi-GPU parking as unsupported.</span></li><li><span>The current documentation still rejects enabled parking with </span><code dir="ltr">--layer-split</code><span>.</span></li></ul><p><span>Niko also noted during #57 that with a layer split each card owns session state for its layers, and that multi-GPU support could come later. This issue is intended as that focused follow-up rather than a competing cache design.</span></p><p><span>If useful, I can provide additional raw timing/monitor logs from the dual RTX 5060 Ti run or test an upstream branch on this hardware.</span></p><p><span>Testing and measurements were conducted as part of AdaxLabs' local multi-agent inference work.</span></p><p><span>Thanks for Strata — the existing same-conversation cache improvement is already dramatic on this setup; supporting parked controller/worker conversations on the split would remove the remaining major latency penalty for local agent workflows.</span></p>

関連リンク

インストール・モデル・リリースへの站内リンク。