Pull requests / #652
#652 serve: commit only the tokens a verify window hands out
closed · @anon761 · 0 评论 · 在 GitHub 查看
Server & APIMulti-GPUNVIDIA / CUDAModels & quants
描述
## What A verify window committed every accepted position (`a + 1`) into the session, but hands out only the outputs below `max_tokens` and through the first end of turn. When an answer ended *inside* a window - a `max_tokens` cap, or an end of turn followed by accepted drafts - the session held drafted tokens the client never saw. The follow-up that sends the answer back then no longer starts with the live session, and the engine falls back to the turn-boundary checkpoint: the whole answer and the new message are read again. The window now counts the outputs it hands out first and commits exactly their positions - the same state as an answer that ends on a window's last token. `consumed` (the live session's ids) follows the same count. Decode itself is unchanged: the drafts and the next window are as before, only the final window's commit can be shorter. ## Measured 2x RTX 3090, Unsloth UD-Q4_K_XL, layer split auto, 262K context, `--spec 4 --spec-min-p 0.5`. A coding conversation of 8 turns over the OpenAI endpoint (greedy); every turn sends the whole history back with the answers verbatim, `max_tokens` alternating between short caps and room for a natural end. "Resumed" = the engine reused at least what it held after the previous turn (prompt + answer less the uncommitted last token), read from `/metrics`. | | follow-ups resumed | the misses' prompt time | | --- | ---: | ---: | | `main` | 5 of 7 | 530-770 ms instead of ~220 ms | | this branch | 7 of 7 | - | | this branch, same conversation after an ~8K-token code block | 7 of 7 | - | Both misses on `main` were answers cut by `max_tokens` inside a window (`reused 30 (held 97)`, `reused 753 (held 849)`). Follow-up turns resume in 200-230 ms at 1K and at 9K tokens of context. One kind of miss remains and is not the engine's: Qwen's chat template strips trailing whitespace from an assistant message, so an answer cut right after a `\n` comes back without it, and the session (which holds that `\n`) cannot rewind to the shorter prompt. That happened once in ~40 sampled follow-ups.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。