Pull requests / #923
#923 serve: answer Codex's thread-title requests without running them (Codex compatibility)
closed · @softbearlolz · 0 commentaires · Sur GitHub
BenchmarksServer & APIDocumentation
Description
Related to #451. A Codex compatibility fix: Codex CLI 0.160 sends a request that other clients do not, and on Strata's single prefix cache it evicts the conversation. Other clients are unaffected: the change applies only when Codex's own `x-codex-turn-metadata` says so.
Codex 0.160's TUI (`tui/src/app/thread_title.rs`) sends a second `POST /v1/responses` beside the user turn to name the thread. The body has `thread_source` `thread_title`, an empty `tools` list, a different `session_id` than the conversation, and the full Codex instructions, so the prompt is about 4,000 tokens rather than a few dozen. The schema is a single `title` of at most 36 characters. `request_kind` is still `turn`.
## What we saw
This server has one prefix cache. A prompt matches only when it starts with the tokens the engine is holding (`generate.cpp`). The title request does not: its tools are empty, and the template writes tools at the front. It also reaches `prepare` before the real turn finishes tokenizing, so it takes the single slot first.
Captured from the live Codex session on 2026-10-05 (requests were logged only for this investigation; that logger is not in this branch):
- 16:29:42 the user line and the title were posted 33 ms apart. The title file is `thread_source=thread_title`, `tools: []`, session `01a10b2e`. The conversation file is session `01a109ab`, 8 tools, 687 input items.
- The engine then read the title with `cache_n` 0 (`prompt` 3,992, 185 generated; earlier the same day 4,149/222 and 4,050/999, the count moving with the user line and the reasoning length).
- The conversation read that followed started at token 0: 207,708 tokens. At ~200 tok/s on this 2080 Ti that is the long stall.
The title is not in the conversation JSONL. Its `session_id` is not the rollout's.
## Change
If `thread_source` is `thread_title`, do not call the engine. Return a completed response whose message is `{"title": "<the user line, cut to 36 characters>"}`. The conversation request is unchanged and keeps the cache.
The title is the user's line, not a model-written title. `docs/DETAILS.md` gets a row for it in the Responses API table.
Why not `--conversation-cache-mib`: with it (the default is 0) the conversation could be parked and restored, but the title request would still read ~4,000 tokens and take the slot first on every user line. This answers it without the engine.
## Evidence
Branch `codex/thread-title`: the change at `58fe594` and the docs row at `9d111b1`, from `v0.1.39` (`6f32ec0`, current `main`). No dump code.
`python3 -m unittest test_responses` in `serve/`: 21 tests OK. The new test posts a normal turn, then a title: the mock engine's prompt is unchanged, and the body is `{"title":"只回四個字然後停:傾印測試"}`. A 40-character line is cut to 36.
On the same 2080 Ti, this build:
- a turn with tools: `reading the prompt: 855 of 862 tokens, 6 s`, `done: 3 tokens in 7 s`, `cache_n` 0, answer `Hi.`
- the title: HTTP 200 in 0.002 s, body `{"title":"Say hi"}`. No `reading the prompt` line and no `done` line for it.
- the next turn of that same conversation, sending the assistant reply back: `done: 3 tokens in 1 s`, `cache_n` 864, `prompt_n` 18.
The request dump used to capture the original bodies is not in this branch.
Sur le site
Liens install, modèles, releases.