Pull requests / #1164

#1164 conversation cache: a shared prefix is copied out of a parked conversation instead of taking it over, so the parked one keeps its history

open · @BlueKingMuch · 0 comments · View on GitHub

Setup & installNVIDIA / CUDAModels & quantsWindows

Description

Builds on #1163 (the first commit); this PR is the second.

An agent harness that runs subagents sends several conversations that start with the same system prompt and tool list. The conversation cache checkpoints that prefix, and when a new conversation finds it only inside a parked one (because another conversation is in the session at that moment), it moves that whole parked conversation into the session and overwrites it after the checkpoint. The parked conversation is gone, and its next turn reads its history again.

With this, the new conversation borrows: only the K/V up to the checkpoint and the checkpoint itself come over (`conversation_snapshot_restore_prefix`), and the parked conversation stays where it is. A match on a parked conversation's live state or its newest checkpoint is that conversation going on, and it is moved into the session as before (`ConversationCache::borrows`). While the outgoing conversation is parked, the one borrowed from is pinned, so parking cannot evict it; if the outgoing conversation fits only in its room, it is moved in as before and the outgoing one keeps its place. The engine log says "borrowed".

#1090 (prefix snapshots on disk) changes the same decision: a saved prefix wins over a parked conversation that matches no further, so that conversation is not taken out. With this PR it is not taken out either, and nothing is read from disk.

## Measured

Same setup as #1163, also on 0.1.40.1 against v0.1.40.1 (a layer split never borrows - its stages restore whole images only - a borrow clears a batch slot as a source, as a taken conversation does, and without `--mtp` it takes no draft K/V, as parking saves none): RTX 4080 SUPER 32 GB, Ryzen 7 5800X3D, 96 GB, Windows 11, IQ3_S, 262K context, `--conversation-cache-mib 2816 --conversation-cache-slots 3`, a fixed expert placement (`--adapt-every 1000000 --expert-cache 8031`), greedy, 64-token answers. A main agent P (30,000 tokens) and two subagents S1 and S2 with the same system prompt (15,000 each) take turns for two rounds, every return +2,300 tokens: P1 S1 P2 S2 S1b S2b P3 S1c S2c P4. Prompt tokens read, time to first token:

| | v0.1.40.1 | #1163 | #1163 + this PR |
|---|---|---|---|
| S2, new subagent while P is in the session | moves S1 in, S1 is overwritten | the same | borrows from S1 (32 ms), S1 stays parked |
| S1b, subagent 1 goes on | 14,878, 4.8 s | 14,878, 4.8 s | 2,369, 2.9 s |
| S2c, subagent 2 in the second round | 17,242, 5.4 s | 2,369, 2.9 s | 2,369, 2.9 s |
| all ten requests | 116,147 read, 49.9 s | 86,332 read, 44.6 s | 73,892 read, 42.9 s |
| evicted | 3 | 1 | 2 |

With both, every return resumes its conversation (2.8 to 3.0 s). The second eviction is S1 at P4, by 5 MB: three conversations of 0.8 to 1.1 GiB and the one being switched in fill 2816 MiB.

Why it builds on #1163: borrowing keeps one more conversation in the cache, and with the reserve unbounded that pushed the main agent out - this PR alone on v0.1.40.1, same run: P3 read 34,728 tokens again (12.1 s instead of 2.9 s), all ten requests 54.5 s instead of 42.9 s.

The answers: with the fixed placement all ten equal #1163 alone (S2 borrowed gives the same tokens as S2 moved in the old way), and the main agent's four answers equal a run without subagents.

Tests: `conversation_cache_test` (which matches borrow and which do not, pinning against slot and byte eviction and the drop of superseded copies, the refusal when the outgoing conversation fits only in the donor's room) and `conversation_file_test` on Windows; `conversation_snapshot_test` (on the GPU: the borrowed prefix and its checkpoint give the same bytes as restoring the whole conversation and rewinding to the checkpoint, and the parked image stays unchanged) on Windows with CUDA. The same run without `--mtp` borrows the same way.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.