Issues / #830
#830 Let a one-shot request skip the prompt-cache work (~45 ms per call that is never reused)
closed · @brenoperucchi · 2 comentários · No GitHub
Setup & installServer & APINVIDIA / CUDAModels & quantsWindows
Descrição
A follow-up to part 2 of #519. Your explanation there of the fixed cost of the batched read (with about 40 % of the experts outside VRAM, a short batch still has to bring in nearly all the missing ones) matches what I see on 0.1.39: over 35 warm batched reads of 212 to 558 tokens the time fits about 245 ms + 0.33 ms per token. Following your suggestion, I moved the checkpoint to the end of the fixed instruction with `--prompt-cache-root 64`, which saves 20 to 40 ms per call. What is left on the cache side still costs about 45 ms per call. ## Setup - RTX 5090 32 GB (no display), Ryzen 9 5950X (AVX2), 96 GB DDR4-3200, Windows 11 - Strata 0.1.39 release, Swift IQ3_XXS, `--expert-cache auto` (14,864 slots, 24.2 GiB), `--prefill auto:32768`, `--max-context 32768`, `--kv int8`, `--pcie-frac 0.20`, `--spec 4 --spec-min-p 0.70`, `--prompt-cache-root 256` for the numbers below (production now runs 64), `STRATA_IQ_MT_MIN=1` - The workload: a fixed system prompt of ~100 tokens, then a user message of 150 to 450 tokens of code, `max_tokens` 1, thinking off. Every call has a different user message, and the same server also takes interactive chats that do reuse their conversations. - The numbers below use the same shape built from public text: a 278-character system prompt (about 70 tokens) and excerpts of v0.1.32's `src/program/generate.cpp`, the same prompts in every run, 12 calls per size with the first 2 skipped, client on the same PC. I can share the script. ## What the prompt cache costs on these calls | User message | Prompt tokens read | `prompt_ms`, cache on | `prompt_ms`, `--prompt-cache 0` | |---|---|---|---| | ~150 tokens | 230 | 375 | 328 | | ~300 tokens | 379 | 434 | 388 | | ~450 tokens | 487 | 460 | 407 | That is about 46 ms of `prompt_ms` and about 40 ms of wall time per call. With the cache on, the prompt is read in two parts, split at the last `<|im_start|>`, and a checkpoint is taken there. That checkpoint ends after the user message, which is different on every call, so nothing reuses it. With `STRATA_TRACE=1`, a 486-token prompt shows `read 479 tokens (batched) in 403.5 ms`, `refilled 190 slots ... in 14.0 ms` and `read 6 tokens (windows) in 14.4 ms` against `prompt 486 tokens ... read in 460 ms`. The trace leaves about 27 ms outside the timed reads, so the checkpoint itself costs at most that and the rest comes from the split. `--prompt-cache 0` removes this for the whole server, and the interactive chats lose their reuse with it. `serve/server.py` doesn't read `cache_prompt`, so a request can't ask for it. ## Question Could a request opt out of both the split at the last turn and the checkpoint when it has nothing to keep, for example by honouring `cache_prompt: false`, or automatically when a one-turn prompt's only reusable part is already a root? I'd expect it to save up to what `--prompt-cache 0` saves on calls like these. Separately: for a short batch where most of the missing experts are needed anyway, would running the non-resident ones on the CPU, as the decode windows do with misses, ever beat streaming them? If there's a trace counter that splits the batched read into transfer and compute time, I'd run it here. I can share the traces and the script, and run a build with a change on this machine.
No site
Links install, modelos, releases.