Pull requests / #1503

#1503 conversation cache: never park one conversation with another's retained K/V

open · @sanastasiou · 0 commentaires · Sur GitHub

Setup & installNVIDIA / CUDAModels & quants

Description

## Bug

With `--batch` and conversation parking, one conversation can be parked with **another conversation's K/V** in its snapshot. Restoring it later serves the other conversation's content.

A restore keeps the restored K/V (`ConversationCache::retain`) so the next park of the main session copies only what changed. The reuse tracked a token *count* (`unchanged_tokens`), not *whose* tokens. `copy_from_slot` — a batch slot giving its conversation back to the main session — replaces the main session's K/V but did not drop the retained image. The next park then reused the earlier conversation's K/V for the new one's prefix.

## Seen in production

2-GPU layer split, `--batch 2`, `--conversation-cache-mib 24576`, 4 Claude Code clients:

```
conversation cache: restored 114705 tokens (live)          <- conversation P; its K/V retained
batch: slot 0 gave back 37564 tokens of this conversation   <- main session now holds Q
conversation cache: parked 37663 tokens ... reused_kv_bytes=573372672   <- Q parked with P's K/V
conversation cache: restored 37663 tokens (live)           <- Q's next turn
```

Q's reply then named a file path that existed only in P's context, and Q answered from P's history from then on. In five 4-conversation runs, 29 parks reused a retained image after a slot had replaced the main session. Most reused only the system prompt shared by all clients, so they went unnoticed.

## Fix

- `retain()` records the ids of the conversation the K/V belongs to.
- Parks call `take_reuse_for(live)`. It reuses only the prefix that `live` shares with those ids. With no shared prefix, or no recorded ids, nothing is reused.
- `copy_from_slot` drops the retained image before it replaces the main session's K/V.
- One stderr line is printed whenever a reuse is cut, so the guard is visible.
- The same edits are made in `sycl/src/program/generate.cpp`. I could not compile the SYCL build here (no oneAPI); the CUDA build compiles.

The reuse only ever skipped a copy. Cutting it costs at most one full copy of the parked conversation's K/V; the result is otherwise identical.

## Tests

`conversation_cache_test` gains four cases:
- another conversation never gets the retained K/V;
- a conversation that diverges reuses only the shared prefix;
- a rewrite's limit still holds, on every stage;
- an image of unknown identity is never reused.

The first case fails if the ids comparison is removed (checked by mutating it). 4204 checks pass.

Since the fix: a 4 × 400K-token compaction run (4 parallel Claude Code conversations, each compacting at about 430K) on the same 2-pair setup gave recall 5/5 in all four, and none of the transcripts mentions another conversation's working directory.

Sur le site

Liens install, modèles, releases.