Pull requests / #1032

#1032 serve: add opt-in bounded Responses history and continuation

closed · @mdwsk88 · 0 コメント · GitHub で見る

Server & APIAMD / HIPModels & quantsSecurityDocumentationWindows

本文

Clients written against OpenAI's Responses API continue a conversation with `previous_response_id` and read or
delete results by ID (`responses.retrieve`, `responses.delete`). On Strata today these requests are refused
(`previous_response_id` is a 400, the ID routes are 404), so such a client has to be rewritten to resend the whole
conversation. I ran into this running Strata on Windows.

This adds optional, bounded storage to the existing Responses adapter: create a response, send only the next input
with `previous_response_id`, then retrieve or delete the stored result.

**Off by default, and off means unchanged.** Without `--responses-store-mib` (or `"responses_store_mib"` in the
config), status codes and response fields are the same as on main: `store` is not looked at (the answer says
`store: false`) and `previous_response_id` is refused as before (`test_disabled_keeps_the_stateless_behavior`).

**When on** (`--responses-store-mib 64`, the config key, or the web page's Model settings):
- `previous_response_id` continues a stored response through the existing generation path, including function-call
  outputs and streaming. Top-level `instructions`, tools and generation settings are not inherited, as in OpenAI's API.
- `GET /v1/responses/{id}` and `DELETE /v1/responses/{id}`; other routes under it (`input_items`, `cancel`) answer
  "not supported".
- Requests are stored unless they say `store: false`. Completed and incomplete responses are saved before their final
  event; failed, cancelled and in-progress ones are not.
- Bounded: the MiB budget counts the serialized UTF-8 JSON including replay history, at most 256 records, 3,600 s
  lifetime (`responses_store_ttl_s`), oldest evicted first. A record that cannot fit is HTTP 413 (`response.failed`
  once streaming has started); a replay that alone exceeds the budget is refused before the model runs.
- Each child stores its own history, so branching works and deleting a parent does not break its children.
- The same API key, Host, browser-Origin and CORS checks as the other routes (a DELETE has no body, so it does not
  need the JSON content type a POST does).

**The conversation cache.** A continuation reads token for token the same prompt as a client that sends the whole
conversation, and the previous turn's prompt is its start
(`test_a_continuation_reads_the_same_prompt_as_the_whole_history`), so the cache is reused as it is for Codex.

Process memory only, lost on restart, and separate from the engine's KV cache. Builds on the stateless
implementation from #451; the inference engine is unchanged.

**Measured** on Windows, RX 7900 XT (HIP), Qwen3.8-Flash-Next IQ2_XS, 32K context, 4 MiB store, with the first
revision of this commit: the OpenAI SDK flow create → continue → retrieve → stream → delete → 404 passed, as did
continuing a child after its parent was deleted, and `store: false`. A run of this revision, with the continuation's
`cached_tokens`, follows.

**Tests:** 292 run, 0 failed, 7 skipped (environment-dependent; macOS, mock engine). 24 are new: CRUD,
continuation, the replayed prompt, branching, tool results, expiry, byte/count limits, concurrency, validation,
failures, auth/CORS, the unchanged default, the Settings key and an optional OpenAI SDK round trip (`openai==3.24.0`).

    python -m unittest serve.test_responses serve.test_response_store serve.test_runconfig
    python -m unittest serve.test_server serve.test_security serve.test_structured serve.test_lifecycle serve.test_parallel serve.test_monitor serve.test_mcp serve.test_detok serve.test_winjob serve.test_vram

This is one self-contained commit on current main. If you would rather land your own version, cherry-picking it
keeps the tests and docs together.

関連リンク

インストール・モデル・リリースへの站内リンク。