Issues / #1620

#1620 serve: a cancelled request loses the conversation cache entry it restored (take() without a park), so the next request re-reads the prompt

open · @taxah92 · 0 comentarios · En GitHub

Server & APINVIDIA / CUDAModels & quants

Descripción

## Summary

A request restored from the conversation cache removes its entry from the cache (`ConversationCache::take()`, `include/strata/core/conversation_cache.hpp:251-256`), and the entry is put back only when the conversation is parked at the end of a **completed** turn (`put()`, line 323). A request that is cancelled (client disconnect, idle timeout, an explicit stop) never parks, so the conversation loses its slot; the next request for it matches only the root checkpoint and re-reads everything after it.

Companion report: #1619 — with one request turn, a queued request receives no bytes at all, which is what makes clients with a stream idle timeout abort in the first place.

## Environment

- Strata v0.1.41 (the same was visible on 0.1.40.3), source build sm_70, Tesla V100 SXM2 16 GB, Ubuntu 26.04
- Qwen3.8-Flash-Next IQ3_S, 1M context (YaRN), `--kv k8v4 --kv-resident 32768`, `--spec 3`, `--conversation-cache-mib 24576 --conversation-cache-slots 3 --prompt-cache 16 --prompt-cache-every 8192`
- Three concurrent conversations (one parent session and two subagents), one request turn

## What we see

Engine log, consecutive lines of one sequence:

```
conversation cache: restored 141020 tokens (live) in 218.4 ms; parked=3 bytes=13708493512
prompt 141160 tokens = 141020 reused + 0 of 140 read ... 0 generated          <- cancelled
conversation cache: restored 105295 tokens (live) in 164.4 ms; parked=2 bytes=9559936460
prompt 105674 tokens = 105295 reused + 374 of 379 read ... 0 generated        <- cancelled
prompt 141160 tokens = 12166 reused + 91392 of 128994 read ... 0 generated    <- only the root matched
prompt 105674 tokens = 12166 reused + 13056 of 93508 read ... 0 generated     <- only the root matched
prompt 141160 tokens = 12166 reused + 128994 read in 96446 ms, 585 generated  <- 96 s re-read
```

`parked=2` shows the entry was taken and not returned. The 12 166-token prefix is what looks like the system prompt plus the tool definitions (the root checkpoint), so 93-129k tokens were re-read for each conversation. Totals from one day of logs: 28 events where only the root matched, ~740k tokens re-read ≈ 9-10 minutes of prefill, plus the same again for the aborted attempts.

## Suggested fix

Keep the taken snapshot alive until the turn ends, or re-insert it when the request is cancelled before it parks. The image already exists — it is `take()`n and dropped. Parking what a cancelled turn has (its state is valid up to the point it stopped) would preserve the slot as well. At the very least, keep the entry instead of removing it: `best()` matches prefixes, so a stale live image still shortens the next read.

Related: #175 (explicit `strata_cache_slot`) — this is the same "returning to the original chat rereads its large prompt" complaint, but on the cancel path rather than the interleaved-request path.

En el sitio

Enlaces a install, modelos, releases.