Pull requests / #1536

#1536 serve: reuse a reply's own token ids on the next turn; keep tokens drained after a budget stop (depends on #1455)

open · @zsariboga · 0 评论 · 在 GitHub 查看

Setup & installServer & APINVIDIA / CUDAModels & quantsWindows

描述

**Depends on #1455** (W1nge's `PromptEncoder`). This branch is #1455's commit plus one commit (`serve: reuse a reply's own token ids on the next turn; keep tokens drained after a budget stop`). Only that last commit is new here. Once #1455 is merged, the diff is that commit alone.

## Problem

In agent sessions, the engine often reads a whole reply again on the next turn, even though the reply's state is still in the session.

After a request, the engine's session holds the prompt plus the reply's **own** token ids. The next turn's prompt renders that reply's text again, and the server encodes the text afresh. The two id sequences part at the first place where the model's ids differ from what BPE produces for the same text. The model writes some text in pieces that BPE would merge: Turkish letters, digits, code. Here is one real case, taken from a reply that contained the regex `[a-zçğıöşü]`:

```
session : ' using regex [a-zçğıöşü' || 'ü] after normalization... Nice approach'
prompt  : ' using regex [a-zçğıöşü' || 'ü] after normalization... Nice approach:'
```

The text is identical, but the ids are not. The session has no checkpoint inside a reply, because GDN state cannot roll back. So the engine returned to the turn's start and read the whole reply again: 364 tokens reused, 1903 read, where about 885 could have been reused.

A second, smaller case: when the thinking budget stops a reply, the STOP reaches the engine 1 to 3 tokens late, and the engine keeps those tokens in its session. The server dropped them. The wrap-up continuation therefore left the session, and the whole thinking was read again. In one measured case, the engine had generated 4098 tokens for a 4096 budget, and the continuation read 4128 tokens.

## Change

- **`PromptEncoder.remember_reply(prompt_ids, gen_ids)`.** After a reply, the server keeps the prompt plus the reply's own ids as one encoder entry.
  - The entry's cuts are the reply's token boundaries. A cut never falls inside a UTF-8 character, and never inside a special tag (`_straddles`).
  - The next prompt takes those ids up to where its text really differs, then encodes the rest afresh.
  - The ids always decode to the prompt's text. This is the one place where prompt ids may differ from full BPE: they are the model's own pieces, which the engine already holds.
- **Encoder limits** raised from 256K characters / 64K tokens to 1M characters / 262K tokens, because agent prompts pass 64K tokens. Those turns are where a re-read costs the most.
- **Thinking budget stop.** The tokens the engine sends while draining after the STOP are now kept (`engine.drained`). They join the thinking when they are plain text: no stop token, no `<`, no split character. Otherwise the old behaviour applies.
- **`STRATA_PROMPT_DUMP=<file>`** (diagnosis, off by default) writes one JSON line per request: the prompt ids, every token the engine sent back including drained ones, and the reused count. Consecutive lines show exactly where a prompt left the engine's state.

## Measured

Setup: RTX 5070 Ti Laptop 12 GB, Core Ultra 7 255HX, 64 GB DDR5-5600, Windows 11. Engine 0.1.40.3, Swift 1.5 abliterated IQ3_XXS, opencode as the client. My server carries some local patches; they are noted where they matter.

**Engine log before the change.** 1740 requests took 2836 s of prompt reading in total. About 400 s of that (~14%) was spent reading a reply again that the engine still held: the new prompt was at least as long as the previous session, but reuse stopped well short of it.

- About 150 s came from divergences inside ordinary replies, the case shown above.
- The rest came from thinking-budget turns. Part of that is a local patch of mine that keeps the wrap-up sentence out of the client's history; the companion engine PR covers that part.

**With the dump (`STRATA_PROMPT_DUMP`)**, an agent test that sends a tool turn after a budget wrap-up:

| next turn | reused | read |
|---|---|---|
| before | 362 / 364 (turn start) | 1379 – 3635 |
| after (this PR + the engine PR) | 969 / 971 (where the texts really differ) | 583 / 1770 (only what follows) |

Four opencode agent tasks with this branch's server: asmsum 4/4 twice, exprcalc 3/3 (20/20 hidden tests), asmbox 4/4. Within a session, the only divergence left was at a 4096-token budget wrap-up, and ~1 token was read again there. Before the change, the same task read 4128 + 4795 tokens again at that point.

Output is unchanged. The engine sees the same tokens it generated, and the text is identical.

## Tests

- `test_prompt_encoder.ReplyReuse` (4 cases):
  - the next turn keeps the reply's ids up to the difference;
  - without its prompt entry, nothing is kept;
  - a tag across a cut is encoded whole;
  - a split character is never a cut.
- `ThinkingBudget`, 2 cases:
  - drained plain-text tokens join the thinking, and the continuation is `first + thought[:20+3] + wrap-up`;
  - drained tokens containing a tag are left out.
- Full suite: `python -m unittest discover -s serve -p "test_*.py"`: 580 tests OK, 9 skipped.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。