Issues / #528
#528 Conversation cache: decode drops ~4-6x on continued conversations (0.1.36, Windows, RTX 5090, IQ3_XXS)
closed · @NikolaFC · 4 Kommentare · Auf GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsWindows
Beschreibung
**Summary** With the conversation cache enabled (`--conversation-cache-mib 8192 --conversation-cache-slots 4`), decode speed on continued conversations drops to roughly **18–31 tokens/s**. With the same engine build, model, flags otherwise, and the same conversation, but those two flags omitted, the conversation decodes at **87–116 tokens/s**. Fresh (non-restored) decoding runs at ~80–95 tokens/s in both configurations. So on 0.1.36 the feature currently costs ~4–6× decode throughput on exactly the workload it targets (multi-turn agent-style flows). **Environment** - Windows 11, RTX 5090 (32 GB), 96 GB system RAM - Strata 0.1.36 — Windows prebuilt binary - Model: `Qwen3.8-Flash-Next` GSQ-RCO IQ3_XXS native pack - Engine flags: `--kv int8 --kv-resident 32768 --max-context 262144 --spec 4 --prompt-cache 6` - Config **A**: plus `--conversation-cache-mib 8192 --conversation-cache-slots 4` - Config **B**: without those two **Measurements** A/B on the same night, same engine binary, same conversation (~100 K tokens), 300-token generations, reading `timings.predicted_per_second` from the OpenAI-compatible endpoint: | Config | Request | Decode (tokens/s) | |---|---|---| | A — cache on | continue ~100 K conversation (restored from parked snapshot) | 17.6, 20.4, 26.0, 31.1 | | B — cache off | continue the same conversation (prompt reuse) | 86.6, 100.5, 115.8 | | B — cache off | fresh ~100 K read | 80.4, 83.0, 94.5 | A longer session under config A (27 turns, prompt growing 71 K → 101 K tokens) showed turns 1–2 at 59–64 tokens/s, then turns 3–27 at 13–40 tokens/s (most 15–31). The GPU stayed at full boost the whole time (2.88–2.90 GHz core, 48–57 °C, well under power limits), so this is not throttling. KV streaming was at 97–98 % VRAM-resident and the expert cache at 93–98 % hit in both the fast and slow cases — it does not look like cache pressure. **What the engine log shows** After each turn the engine parks a snapshot (≈2.1–2.4 GB, ~0.6–0.9 s); each following turn restores it in a few hundred ms, and it is the restored state that decodes slowly. A from-scratch read of comparable size under the same config decodes fast — the cost tracks the restore path, not general memory pressure. **Reproduction** 1. Load the pack with config A. 2. Send a ~100 K-token prompt, generate 300 tokens — note `predicted_per_second`. 3. Continue the same conversation with small deltas a few times — decode settles at ~18–31 tokens/s. 4. Restart with config B and repeat — the same conversation decodes at ~87–116 tokens/s. Happy to share the small client script and the raw engine logs. **Question** Is decode on restored conversations currently expected to pay a reconstruction penalty of this size (e.g. retained-KV handling in the snapshot path), or is 4–6× outside the intent of the RAM tier? Happy to run any diagnostics on this machine. Related: #57 (multi-conversation cache tracker).
Mehr auf der Site
Links zu Install, Modellen, Releases.