Issues / #1620
#1620 serve: a cancelled request loses the conversation cache entry it restored (take() without a park), so the next request re-reads the prompt
open · @taxah92 · 0 Kommentare · Auf GitHub
Server & APINVIDIA / CUDAModels & quants
Beschreibung
## Summary A request restored from the conversation cache removes its entry from the cache (`ConversationCache::take()`, `include/strata/core/conversation_cache.hpp:251-256`), and the entry is put back only when the conversation is parked at the end of a **completed** turn (`put()`, line 323). A request that is cancelled (client disconnect, idle timeout, an explicit stop) never parks, so the conversation loses its slot; the next request for it matches only the root checkpoint and re-reads everything after it. Companion report: #1619 — with one request turn, a queued request receives no bytes at all, which is what makes clients with a stream idle timeout abort in the first place. ## Environment - Strata v0.1.41 (the same was visible on 0.1.40.3), source build sm_70, Tesla V100 SXM2 16 GB, Ubuntu 26.04 - Qwen3.8-Flash-Next IQ3_S, 1M context (YaRN), `--kv k8v4 --kv-resident 32768`, `--spec 3`, `--conversation-cache-mib 24576 --conversation-cache-slots 3 --prompt-cache 16 --prompt-cache-every 8192` - Three concurrent conversations (one parent session and two subagents), one request turn ## What we see Engine log, consecutive lines of one sequence: ``` conversation cache: restored 141020 tokens (live) in 218.4 ms; parked=3 bytes=13708493512 prompt 141160 tokens = 141020 reused + 0 of 140 read ... 0 generated <- cancelled conversation cache: restored 105295 tokens (live) in 164.4 ms; parked=2 bytes=9559936460 prompt 105674 tokens = 105295 reused + 374 of 379 read ... 0 generated <- cancelled prompt 141160 tokens = 12166 reused + 91392 of 128994 read ... 0 generated <- only the root matched prompt 105674 tokens = 12166 reused + 13056 of 93508 read ... 0 generated <- only the root matched prompt 141160 tokens = 12166 reused + 128994 read in 96446 ms, 585 generated <- 96 s re-read ``` `parked=2` shows the entry was taken and not returned. The 12 166-token prefix is what looks like the system prompt plus the tool definitions (the root checkpoint), so 93-129k tokens were re-read for each conversation. Totals from one day of logs: 28 events where only the root matched, ~740k tokens re-read ≈ 9-10 minutes of prefill, plus the same again for the aborted attempts. ## Suggested fix Keep the taken snapshot alive until the turn ends, or re-insert it when the request is cancelled before it parks. The image already exists — it is `take()`n and dropped. Parking what a cancelled turn has (its state is valid up to the point it stopped) would preserve the slot as well. At the very least, keep the entry instead of removing it: `best()` matches prefixes, so a stale live image still shortens the next read. Related: #175 (explicit `strata_cache_slot`) — this is the same "returning to the original chat rereads its large prompt" complaint, but on the cancel path rather than the interleaved-request path.
Mehr auf der Site
Links zu Install, Modellen, Releases.