Pull requests / #652

#652 serve: commit only the tokens a verify window hands out

closed · @anon761 · 0 评论 · 在 GitHub 查看

Server & APIMulti-GPUNVIDIA / CUDAModels & quants

描述

## What

A verify window committed every accepted position (`a + 1`) into the session, but hands out only the outputs below `max_tokens` and through the first end of turn. When an answer ended *inside* a window - a `max_tokens` cap, or an end of turn followed by accepted drafts - the session held drafted tokens the client never saw. The follow-up that sends the answer back then no longer starts with the live session, and the engine falls back to the turn-boundary checkpoint: the whole answer and the new message are read again.

The window now counts the outputs it hands out first and commits exactly their positions - the same state as an answer that ends on a window's last token. `consumed` (the live session's ids) follows the same count. Decode itself is unchanged: the drafts and the next window are as before, only the final window's commit can be shorter.

## Measured

2x RTX 3090, Unsloth UD-Q4_K_XL, layer split auto, 262K context, `--spec 4 --spec-min-p 0.5`. A coding conversation of 8 turns over the OpenAI endpoint (greedy); every turn sends the whole history back with the answers verbatim, `max_tokens` alternating between short caps and room for a natural end. "Resumed" = the engine reused at least what it held after the previous turn (prompt + answer less the uncommitted last token), read from `/metrics`.

| | follow-ups resumed | the misses' prompt time |
| --- | ---: | ---: |
| `main` | 5 of 7 | 530-770 ms instead of ~220 ms |
| this branch | 7 of 7 | - |
| this branch, same conversation after an ~8K-token code block | 7 of 7 | - |

Both misses on `main` were answers cut by `max_tokens` inside a window (`reused 30 (held 97)`, `reused 753 (held 849)`). Follow-up turns resume in 200-230 ms at 1K and at 9K tokens of context.

One kind of miss remains and is not the engine's: Qwen's chat template strips trailing whitespace from an assistant message, so an answer cut right after a `\n` comes back without it, and the session (which holds that `\n`) cannot rewind to the shorter prompt. That happened once in ~40 sampled follow-ups.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。