Pull requests / #1221

#1221 serve: answer Codex thread-title turns without reading them

closed · @softbearlolz · 0 コメント · GitHub で見る

Server & APINVIDIA / CUDAModels & quantsDocumentation

本文

Codex thread-title turns are answered without entering the engine: 0 tokens and 5.67 ms instead of reading a 4,020-token prompt and generating 498 tokens (17 s warm, 47 s cold) on main.

Supersedes #923 (closed when main was rewritten). Same kind of fix as #965: one client's request shape kept the single prefix cache from matching.

Codex 0.160 sends a second `POST /v1/responses` with `thread_source: "thread_title"` next to each user turn to name the thread. The request has `tools: []`, a different `session_id` than the conversation, and the full system instructions.

On Strata's single prefix cache, that request reaches prepare before the user turn finishes tokenizing, takes the slot, and replaces the conversation KV cache. The user turn must then be read again from token 0.

If `thread_source` is `thread_title`, do not call the engine. Return a completed response whose message content is the user prompt line cut to 36 characters. The conversation request is unchanged and keeps the cache.

To support this and downstream compaction handling cleanly, turn metadata extraction is handled by `_codex_turn_metadata(req)`, which parses `client_metadata["x-codex-turn-metadata"]` and returns None when absent or invalid.

I found it from the logs: I noticed the server re-reading the whole prompt on every turn (the server log showed very few tokens reused), saved the requests with my own request-dump script and compared consecutive ones. That showed a Codex thread_title request (`thread_source: "thread_title"`, `tools: []`, its own `session_id`) going into the engine between the user turns and taking the cache.

## Measured

One RTX 2080 Ti, Qwen3.8-Flash-Next-IQ3_XXS, 262144 context, strata:v0.1.39-sm75-compact. Only this change between the two runs.

| Arm | Engine invoked | Prompt tokens | Output tokens | Latency | Cache slot impact |
|---|---|---|---|---|---|
| Main (cold) | Yes | 4,020 | 655 | 47.4 s | Evicts conversation |
| Main (warm) | Yes | 4,020 (4,015 cached) | 498 | 17.4 s | Evicts conversation |
| Branch v2 | No | 0 | 0 | 5.67 ms | Conversation kept |

The title returned is the first user line cut to 36 characters, returned immediately without engine prefill or generation.

## Tests

Command: `python -m unittest discover -s serve -p 'test_*.py' -t .`
- Upstream `main`: 452 tests (446 passed, 6 errors)
- `codex/thread-title-v2`: 454 tests (448 passed, 6 errors)

The 6 errors are pre-existing on `main` (`test_detok.HeapBpe` x3, `test_server.LiteralControlTokens` x2, `test_server.LiteralThinkTags` x1). The 2 new tests (`CodexTurnMetadata.test_metadata_parsing`, `ThreadTitle.test_a_thread_title_is_answered_without_the_engine`) exercise functions `main` does not have (`_codex_turn_metadata`, `thread_title_events`), so they cannot run there; the base's behaviour is what ## Measured shows.

`docs/DETAILS.md` gets a row for thread-title turns in the Responses API table.

関連リンク

インストール・モデル・リリースへの站内リンク。