Pull requests / #1222
#1222 serve: render a Codex compaction with that conversation's tool prefix
closed · @softbearlolz · 0 comentários · No GitHub
Server & APINVIDIA / CUDAModels & quantsDocumentation
Descrição
**Stacked on #1221.** This branch carries its commits; only the last 3 commits are new. Codex's local compaction (`tools: []`) reuses 12,875 of 12,976 tokens (3 s) instead of reading 10,789 tokens from scratch (48 s) on main. Supersedes #924 (closed when main was rewritten). Codex CLI's local compaction (when context fills up or on `/compact`) sends the conversation again with `tools: []`. This server's chat template writes tool definitions at the front of the prompt, so a request without tools shares only initial system tokens with the prefix the engine is holding. The compaction prompt, the conversation at its longest, is then read again from the start. For a request marked `request_kind: "compaction"`, with the same `session_id` and `thread_id` as the conversation that last rendered a prompt, and with no tools of its own, render the prompt with those tools; if the prompt is too long for the context with the kept tools, the request is rendered as sent. The output parser and the response still see the tools Codex sent, which are none. Turn metadata is extracted using `_codex_turn_metadata(req)` from the base branch (#1221), reading only `client_metadata["x-codex-turn-metadata"]`. A thread-title turn does not update the stored tool entry. I found it from the logs: I noticed the server re-reading the whole prompt on every turn (the server log showed very few tokens reused), saved the requests with my own request-dump script and compared consecutive ones. That showed the compaction request (`request_kind: "compaction"`, `tools: []`) rendering without the conversation's tool definitions, so its prompt diverged from the cached one right after the first system tokens. ## Measured One RTX 2080 Ti, Qwen3.8-Flash-Next-IQ3_XXS, 262144 context, strata:v0.1.39-sm75-compact. Only this change between the two runs. | Arm | Turn 1 Prompt | Turn 1 Reused / Elapsed | Turn 2 Prompt | Turn 2 Reused / Elapsed | Turn 2 Dump | |---|---|---|---|---|---| | Main | 12,880 | 0 / 58 s | 10,789 | 0 (0.0%) / 48 s | miss-20261006T142850-996 (reused=0) | | Branch v2 | 12,880 | 7,572 / 24 s | 12,976 | 12,875 (99.2%) / 3 s | None (cache hit) | v2's Turn 1 reuse of 7,572 is left over from the previous set (Set 1) run in the same container, not part of this change; Turn 2 is the comparison. On main, Turn 2 stripped tools and had 0 reused tokens, taking 48 s and writing a request dump miss. On Branch v2, Turn 2 restored the tool prefix, reused 12,875 tokens, and finished in 3 s without triggering a dump. ## Tests Command: `python -m unittest discover -s serve -p 'test_*.py' -t .` - Upstream `main`: 452 tests (446 passed, 6 errors) - `codex/thread-title-v2`: 454 tests (448 passed, 6 errors) - `codex/compaction-tool-prefix-v2`: 465 tests (459 passed, 6 errors) The 6 errors are pre-existing on `main` (`test_detok.HeapBpe` x3, `test_server.LiteralControlTokens` x2, `test_server.LiteralThinkTags` x1). The 11 new tests (`CodexCompaction` x8, `CodexCompactionOverHttp` x3) exercise functions the base `codex/thread-title-v2` does not have (`prompt_tools`, `prompt_made`), so they cannot run there; the base's behaviour is what ## Measured shows. The commits add a 'Codex's compaction' paragraph to the Codex section of `docs/DETAILS.md` and update the 'stateless' paragraph to note the prompt-cache hint kept between requests, matched by the module docstring of `serve/responses.py`; the last commit replaces the paragraph's measurement with these A/B numbers.
No site
Links install, modelos, releases.