Pull requests / #1319

#1319 serve: add opt-in bounded Responses history and continuation

open · @mdwsk88 · 0 commentaires · Sur GitHub

BenchmarksSetup & installServer & APIAMD / HIPModels & quantsWindows

Description

A local client using the OpenAI Python SDK cannot currently continue a Strata response with `previous_response_id` or retrieve/delete its result by ID. This adds opt-in bounded history to the existing Responses adapter so the client can send only new input on later turns.

Enable it with `--responses-store-mib 64`, `"responses_store_mib": 64` in the config, or Model settings. With storage disabled, POST /v1/responses keeps the existing stateless behavior, including accepting/ignoring `store: true`.

- Continue through the existing generation path, including function-call outputs and streaming.
- Add GET and DELETE /v1/responses/{id}, reusing the API-key, Host and browser-Origin checks.
- Save completed/incomplete responses before their final event; honor `store: false`.
- Keep independent child histories, so deleting a parent does not break surviving children. Top-level instructions, tools and generation settings must be sent again.

Storage is process memory only and is lost on restart. The serialized JSON budget includes replay history, with a maximum of 256 records and a default 3,600-second lifetime. Oldest records are evicted; oversized records are refused. Python bookkeeping and temporary copies need additional memory. This does not change the inference engine or its KV cache.

**Validation on current main:** rebased onto `d5ea713`, with one conflict in `serve/server.py` `main()`, beside the new `api_key_of` check; upstream's check is kept and the store setup follows it. On this base all 587 `serve` tests ran (macOS, mock engine): 0 failed, 8 environment-dependent skips.

**On the previous base `82f46a8`:** 476 tests exercised on Windows: 470 passed, 6 environment-dependent skips. The 54 Responses/config tests passed, as did the 422 service regression tests using a long-path TEMP/TMP directory. The report records a pre-existing short-path TEMP assertion failure, reproduced on unmodified upstream main.

Native IQ2_XS validation (on `82f46a8`; the feature code is unchanged since) also passed on the RX 7900 XT (HIP), 32K context and a 64 MiB store, using the installed engine. All 10 helper checks and the official OpenAI SDK 3.24.0 round trip passed, including streaming and a bodyless DELETE from an allowed CORS origin. A continuation reused **777/812 input tokens (95.69%)** and took **1.109 s** for the whole request; this was a functional check, not a throughput benchmark.

[Commands, source-tree identity and measured results](https://github.com/mdwsk88/Strata/tree/codex/pr1032-windows-results/bench/results/2026-10-07-responses-state-windows).

This replaces #1032 after main's history rewrite, following the [maintainer's instructions](https://github.com/Niko1221/Strata/pull/1032#issuecomment-6013973026). It contains one feature commit on `d5ea713` (first opened on `82f46a8`), with the existing implementation, tests and documentation carried forward.

Sur le site

Liens install, modèles, releases.