Pull requests / #1537

#1537 serve: keep a checkpoint where a request continues the live session after a long reply

open · @zsariboga · 0 评论 · 在 GitHub 查看

Setup & installServer & APINVIDIA / CUDAModels & quantsWindows

描述

## Problem

Some requests extend the live session: a thinking budget's wrap-up continuation, a forced call's opening, a reasoning-close retry. Such a request starts from the state after the reply generated so far, so it reuses everything. No checkpoint exists inside a reply, because GDN cannot roll back.

If a later prompt leaves the session right at that point, the engine goes back to the last checkpoint below, which is the turn's start, and reads the whole reply again. Two ways this happens:

- the client's history does not carry the text the continuation added;
- the client renders a call differently.

After a 12288-token budget, that is ~12K tokens, or 6–7 s on my laptop, every time it happens. In my engine log it happened 15 times at a 12288-token budget and 5 times at 4096.

## Change

Before reading, a request that continues the live session saves a conversation checkpoint at that point. It does this only when the newest checkpoint below is at least `STRATA_LIVE_CKPT_MIN` tokens back (default 1024; 0 = off).

It is one ordinary checkpoint, saved through the same `checkpoint_at()` the prompt path uses (~118 MB of host RAM; the LRU keeps `--prompt-cache` of them). It is skipped for `ckpt=0`, `STRATA_CKPT_REREAD`, a parked conversation, a slot or a prefix snapshot. The engine log says `checkpoint at the live end: N tokens`. Nothing else changes.

## Measured

Setup: RTX 5070 Ti Laptop 12 GB, Core Ultra 7 255HX, 64 GB DDR5-5600, Windows 11. Engine 0.1.40.3, Swift 1.5 abliterated IQ3_XXS.

The test is a tool turn after a budget wrap-up (budget 600; `STRATA_LIVE_CKPT_MIN=256` so the short test reaches it). My server keeps the wrap-up sentence out of the client's history, which is a local patch. The next turn therefore leaves the session exactly where the wrap-up starts:

| next turn | reused | read |
|---|---|---|
| before | 362 / 364 (turn start) | 1379 – 3635 |
| after | 969 / 971 (the checkpoint at the live end) | 583 / 1770 (only what follows) |

In four opencode agent tasks with the default threshold, the engine saved four live-end checkpoints, at 3134–4801 tokens past the one below. A 4096-token budget wrap-up in asmsum then read ~1 token again instead of 4795. The tasks passed: asmsum 4/4 twice, exprcalc 3/3 (20/20), asmbox 4/4.

For ids that diverge inside a reply's text, see the server PR, which reuses the reply's own token ids. The two PRs are independent. This one helps whenever the divergence is at the point where a continuation started.

Built on `main` (0.1.40.4) for sm_120.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。