Pull requests / #924
#924 serve: render Codex's compaction with the conversation's tool prefix (Codex compatibility)
closed · @softbearlolz · 0 comentários · No GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentation
Descrição
Depends on #923 (merge that first; this branch will need a rebase after).
Related to #451. A Codex compatibility fix: Codex CLI 0.160 compacts with a request that other clients do not send, and on Strata's single prefix cache it misses the prefix of the conversation it compacts. Other clients are unaffected: the change applies only when Codex's own `x-codex-turn-metadata` says so.
Codex CLI's local compaction (the context fills up, or `/compact`) sends the conversation once more with `tools: []` (`codex-rs/core/src/compact.rs`: `Prompt { input, base_instructions, ..Default::default() }`; also reported for xAI as openai/codex#46773). This server's chat template writes the tool definitions at the top of the prompt, so that request shares only its first few tokens with the conversation the engine is holding, and the longest prompt is read again.
## What we saw
On a local Codex 0.160 session against Strata 0.1.39 (one RTX 2080 Ti, IQ3_XXS, 262144 context) a compaction was the point where a long conversation stopped matching. The template puts tools first, and `tools: []` changes that prefix. The engine keeps one live prefix (`generate.cpp`: the next prompt has to start with the tokens it holds). `--conversation-cache-mib` defaults to 0, so the previous conversation is not parked. A full reread of ~80k–120k tokens on this card is several minutes (a 119,871-token read took 555 s, about 209 tok/s).
`prompt_cache_key` is not the right identity. Codex gives sub-agents and ephemeral forks their parent's key, while `session_id` and `thread_id` in `client_metadata["x-codex-turn-metadata"]` are their own.
A thread-title turn (`thread_source` `thread_title`) is a different session Codex sends beside the user turn, also with `tools: []`. It must not be stored as the conversation's last prompt, or the one remembered entry would follow the title instead of the conversation.
## Change
For a request Codex marks `"request_kind": "compaction"`, with the same `session_id` and `thread_id` as the conversation that last rendered a prompt, and with no tools of its own, render the prompt with those tools. The output parser and the response still see the tools Codex sent, which are none. One entry, replaced on each such prompt. A compaction that would not fit with the kept tools, a different conversation, missing metadata (Codex before 0.140), or a restart is rendered as sent, as today. A `thread_title` turn does not update the entry.
## Evidence
Branch `codex/compaction-tool-prefix`: the change at `c273a1b` (with its `docs/DETAILS.md` section) and a rewrap of that section at `85396ef`, from `v0.1.39` (`6f32ec0`, which is current `main`).
Mock engine, same code and the IQ3_XXS tokenizer, scripted 85,000-token conversation: compaction reused 84,895 of 84,997 tokens and read 102. Without this it reused 41 of 80,683. `python3 -m unittest test_responses` in `serve/`: 31 tests OK, including A/B isolation, a sub-agent and a fork, a custom compact prompt only when `request_kind` says so, the compact wording inside a normal turn, and a thread-title turn leaving the kept tools in place.
On the same 2080 Ti, this build, one short conversation then a compaction with `tools: []`:
- the turn: `reading the prompt: 896 of 901 tokens, 7 s`, `done: 8 tokens in 8 s`
- the compaction: no second cold read; `done: 8 tokens in 1 s`; `cache_n` 896, `prompt_n` 20
The request dump used while investigating is not in this branch.
No site
Links install, modelos, releases.