Pull requests / #1032
#1032 serve: add opt-in bounded Responses history and continuation
closed · @mdwsk88 · 0 评论 · 在 GitHub 查看
Server & APIAMD / HIPModels & quantsSecurityDocumentationWindows
描述
Clients written against OpenAI's Responses API continue a conversation with `previous_response_id` and read or
delete results by ID (`responses.retrieve`, `responses.delete`). On Strata today these requests are refused
(`previous_response_id` is a 400, the ID routes are 404), so such a client has to be rewritten to resend the whole
conversation. I ran into this running Strata on Windows.
This adds optional, bounded storage to the existing Responses adapter: create a response, send only the next input
with `previous_response_id`, then retrieve or delete the stored result.
**Off by default, and off means unchanged.** Without `--responses-store-mib` (or `"responses_store_mib"` in the
config), status codes and response fields are the same as on main: `store` is not looked at (the answer says
`store: false`) and `previous_response_id` is refused as before (`test_disabled_keeps_the_stateless_behavior`).
**When on** (`--responses-store-mib 64`, the config key, or the web page's Model settings):
- `previous_response_id` continues a stored response through the existing generation path, including function-call
outputs and streaming. Top-level `instructions`, tools and generation settings are not inherited, as in OpenAI's API.
- `GET /v1/responses/{id}` and `DELETE /v1/responses/{id}`; other routes under it (`input_items`, `cancel`) answer
"not supported".
- Requests are stored unless they say `store: false`. Completed and incomplete responses are saved before their final
event; failed, cancelled and in-progress ones are not.
- Bounded: the MiB budget counts the serialized UTF-8 JSON including replay history, at most 256 records, 3,600 s
lifetime (`responses_store_ttl_s`), oldest evicted first. A record that cannot fit is HTTP 413 (`response.failed`
once streaming has started); a replay that alone exceeds the budget is refused before the model runs.
- Each child stores its own history, so branching works and deleting a parent does not break its children.
- The same API key, Host, browser-Origin and CORS checks as the other routes (a DELETE has no body, so it does not
need the JSON content type a POST does).
**The conversation cache.** A continuation reads token for token the same prompt as a client that sends the whole
conversation, and the previous turn's prompt is its start
(`test_a_continuation_reads_the_same_prompt_as_the_whole_history`), so the cache is reused as it is for Codex.
Process memory only, lost on restart, and separate from the engine's KV cache. Builds on the stateless
implementation from #451; the inference engine is unchanged.
**Measured** on Windows, RX 7900 XT (HIP), Qwen3.8-Flash-Next IQ2_XS, 32K context, 4 MiB store, with the first
revision of this commit: the OpenAI SDK flow create → continue → retrieve → stream → delete → 404 passed, as did
continuing a child after its parent was deleted, and `store: false`. A run of this revision, with the continuation's
`cached_tokens`, follows.
**Tests:** 292 run, 0 failed, 7 skipped (environment-dependent; macOS, mock engine). 24 are new: CRUD,
continuation, the replayed prompt, branching, tool results, expiry, byte/count limits, concurrency, validation,
failures, auth/CORS, the unchanged default, the Settings key and an optional OpenAI SDK round trip (`openai==3.24.0`).
python -m unittest serve.test_responses serve.test_response_store serve.test_runconfig
python -m unittest serve.test_server serve.test_security serve.test_structured serve.test_lifecycle serve.test_parallel serve.test_monitor serve.test_mcp serve.test_detok serve.test_winjob serve.test_vram
This is one self-contained commit on current main. If you would rather land your own version, cherry-picking it
keeps the tests and docs together.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。