Pull requests / #965

#965 serve: pin Claude Code's per-request billing stamp, so the conversation cache reuses agent prompts

closed · @signalnine · 0 コメント · GitHub で見る

BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentationLinux

本文

Claude Code sessions reuse 91.7% of their prompt tokens with this change, up from 49.8%, and a SWE-bench run drops from 80 s to 60 s per instance.

Claude Code starts its system prompt with a block like

```
x-anthropic-billing-header: cc_version=2.1.170.bf4; cc_entrypoint=sdk-cli; cch=b145e;
```

`cch` changes on every request and the 4th part of `cc_version` changes on every session. The template renders the tool list before the system text, so with Claude Code's ~22K tokens of tools every turn's prompt differed at token ~22,060. The engine's turn checkpoint never matched the next request, only the periodic 16,384-token one did, and every follow-up turn read 8-13K tokens again.

`anthropic_to_messages` now pins both stamps to `f`s, the way llama.cpp does (ggml-org/llama.cpp#21793). It only touches a header at the very start of the system text, only inside its first 160 characters, and only a `cch` value of at most 16 characters or a `cc_version` with more than three parts. The rest of the system prompt and the same text in a user message stay as they are.

I found it from the logs: every follow-up request logged exactly `16384 reused`. Rendering consecutive recorded `/v1/messages` bodies through `anthropic_to_messages` + `Service.encode_prompt` put the first differing token at 22,060-22,061 in every pair, right at the `cch=` value.

## Measured

Unsloth UD-IQ4_XS on one RTX 5090 (Ryzen 9 9900X, 128 GB, Linux), 128K context, int8 KV. Claude Code 2.1.170 (headless, `claude -p`) on 12 pinned SWE-bench Verified instances, one run each, same harness and engine, only this change between the two runs:

| | prompt tokens reused | tokens read | wall per instance |
| --- | ---: | ---: | ---: |
| main | 49.8% | 2.15M over 181 requests | 80 s |
| this PR | 91.7% | 0.48M over 194 requests | 60 s |

Follow-up turns read ~2.5K tokens instead of ~12.8K. Both runs edited the gold patch's file in 10-11 of 12 instances; with one run each, that difference is noise.

## Tests

- New `ClaudeCodeBillingStamp` in `serve/test_server.py`: both stamp forms (block-list and string system prompts), Claude Code's newer form without `cch`, a plain 3-part version kept, a header not at the start (or in a user message) left alone, the bounds, and two turns through `Service.encode_prompt` sharing the whole system prompt. 4 of the 7 fail on `main` and all pass with the change.
- `python -m unittest discover -s serve -p 'test_*.py' -t .`: 275 tests, plus one error that `main` has too (`test_responses.OverHttp.test_json_schema_text_format`).
- `docs/DETAILS.md`: one sentence in the conversation cache paragraph.

関連リンク

インストール・モデル・リリースへの站内リンク。