Pull requests / #1013

#1013 serve: a continued batch request reports its own prompt's reuse; input_tokens never 0

closed · @CurtisThe · 0 comentarios · En GitHub

Setup & installServer & APINVIDIA / CUDAModels & quantsLinux

Descripción

With `"parallel": 2`, our Claude Code agents stopped with **"Prompt is too long · automatic compaction failed"** at 60–65k tokens of a 137k window. The engine never rejected a request: every turn was HTTP 200, with no SSE error event and no stream cut short.

**Cause.** A request that is STOPped on the solo path and continued in a batch slot (or sent back to the solo path) is sent again as its prompt plus what it generated. The DONE of that last leg counts this request's own output as `reused`. `min(reused, len(ids))` then turns it into the whole prompt, so `message_delta` says `input_tokens: 0` and `cache_read_input_tokens` = the whole prompt.

Claude Code keeps `message_start`'s `input_tokens` when `message_delta`'s is 0 and adds `cache_read_input_tokens`, so it counts the prompt twice and compacts or blocks at half the window.

**Measured** (RTX 5090 32 GB, Ryzen 9 9950X3D, 91 GB RAM, Qwen3.8-Flash-Next IQ3_S, v0.1.39 at 6f32ec0, Claude Code 2.1.290, a 137k window):

| v0.1.39 | agent runs | responses | `input_tokens: 0` | agent result |
|---|---|---|---|---|
| agent alone (parallel 1: 3 runs; parallel 2: 1 run) | 4 | 204 | 0 | no "Prompt is too long" |
| agent beside chats or a second agent (parallel 2 and 4) | 5 | 177 | 2–14 per run | 2 ended "Prompt is too long · compaction failed" at 60–65k |
| this patch, parallel 2, agent beside 8 chats | 1 | 54 | 0 (smallest 57) | passed, context up to 108k |

**Claude Code's side was measured against a stub** that sends these exact usage events:
- `input_tokens: 0` with a 64k prompt: auto-compaction at `preTokens: 128073`, the same as a real 128k prompt.
- `input_tokens: 1` with the rest as `cache_read_input_tokens`: no compaction.

**Change**
- `generate_batched` keeps the first leg's `reused`, which is the request's own prompt, across a continuation.
- `anthropic_events` clamps the reported reuse to `len(ids) - 1`. The last prompt token is always read, so `input_tokens` is never 0 and `input_tokens + cache_read_input_tokens` is still the prompt.
- `test_parallel.py`: the fake engine can report a held prefix (`--reuse`). The new test, `test_a_continued_request_reports_its_own_prompt_reuse`, fails without the fix (`input_tokens: 0`, `cache_read_input_tokens: 260`) and passes with it.
- All 269 tests under `serve/` pass (`python -m unittest discover -s serve -t . -p 'test_*.py'`, Linux, Python 3.12, `jsonschema` installed). No engine change.

The OpenAI path (`prompt_tokens_details.cached_tokens`, timings `cache_n`) uses the same `reused`, so it now also reports the first leg's figure for a continued request. I left its clamp alone.

**Related open PRs**
- #998 and #751 touch the same lines. #998 adds `--fail-window` to `test_parallel`'s fake engine where this adds `--reuse`. #751 adds a line to `generate_batched`'s cleanup block next to the one added here. In both cases the fix is to keep both changes; I can rebase onto whichever lands first.
- #753 (draft, waiting on #559) reworks per-request accounting in the batch path, including continuation segments. If it lands, it would replace the `reused0` part of this change. This PR is the minimal fix on current `main`.
- #615 applies the same rule (report the first segment's input/cache counts) to `reasoning_budget_tokens` continuations, which use a different code path.

I wrote this with an AI coding assistant and checked the change and the test runs myself.

En el sitio

Enlaces a install, modelos, releases.