Pull requests / #630

#630 serve: add opt-in stateless Responses API

closed · @CC-David-CC · 0 Kommentare · Auf GitHub

Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindows

Beschreibung

Related to #451.

Adds an experimental stateless `POST /v1/responses` adapter for clients such as Codex. Text, reasoning, function calls and client-owned tool results pass through the existing Strata service. One request-local assembler produces final JSON and typed SSE with stable item IDs, indexes and sequence numbers.

This is a bounded compatibility profile. Unsupported behavior is rejected explicitly. This draft needs the accounting and replay-contract follow-ups below before it is ready to merge.

## Behavior and ownership

- Calls `Service.run()` directly, consuming semantic parser events. Existing FIFO admission, model execution, authentication, CORS, cancellation and engine-output draining remain in use.
- Supports content-bearing input history, per-request instructions, explicit `store:false`, non-strict functions including namespaces, matching `function_call_output` items, visible reasoning, separately generated summaries and authenticated Strata-issued replay tokens.
- Keeps response state request-local. The client supplies history; no state is inferred from GPU cache. No new serving framework or native/MTP changes are included.
- Includes a small existing `/load` and `/unload` fix: consume bounded request bodies before replying. The Windows connection-reset failure was also reproduced against the selected upstream base; the R4 report records it.

## Enablement

Disabled by default. Install `requirements-responses.txt`, supply a stable `STRATA_RESPONSES_REPLAY_KEY` in the server environment, and start the existing configured server with `--experimental-responses` or top-level config `"experimental_responses": true`. Use the existing API-key configuration. Disabled startup imports no optional crypto dependency and leaves `/v1/responses` unavailable.

Clients use the ordinary Responses endpoint with explicit `store:false`; no experimental request field is needed. The [guide](https://github.com/CC-David-CC/Strata-a5500/blob/46a44c51782eae4119b85b27c5b3639bc870e0a2/docs/RESPONSES.md) describes setup, continuation, the tested Codex profile and capability limits.

## Validation and its limits

- Recorded R4 CPU/MockEngine regression checkpoint: 235 passed, five optional detokenizer skips. The final environment-key check passed all 44 Responses tests. [R4 report and commands](https://github.com/CC-David-CC/Strata-a5500/blob/46a44c51782eae4119b85b27c5b3639bc870e0a2/docs/responses-evidence/R4/REPORT.md).
- Official `openai==3.23.0`: JSON, typed streaming and assembly, namespaced function/result replay, reasoning and authenticated replay. Process-restart checks verify replay with the same deployment key and no stored response files.
- Codex CLI 0.160.0 performed a real local read/edit/verify loop over four requests with three successful tool results. **The model output in that protocol test was scripted MockEngine output.** [Client receipt](https://github.com/CC-David-CC/Strata-a5500/blob/46a44c51782eae4119b85b27c5b3639bc870e0a2/docs/responses-evidence/R4/codex/codex-task.json).
- Additional native evidence exists on the GBNF follow-up: pinned local Codex received `Hello world` from a real Coder IQ1_M model on an RTX 4090 over LAN, including real reasoning and a generated summary. It did not call tools. This validates the combined follow-up configuration, not an isolated native test of this PR. [Native receipt](https://github.com/CC-David-CC/Strata-a5500/blob/c91260ccd3a929ae51696c9dff1f972d6d56adfe/docs/responses-evidence/native-codex/REPORT.md).

The source audit verifies no native/build changes against upstream `99f3dbd0b21d1401b3769e0c0d963913607f380b`. These results cover the recorded client/profile and do not establish compatibility with every Codex version or a complete native coding workflow.

## Excluded capabilities

No retained-response CRUD, `previous_response_id`, background execution/cancel endpoint, compaction, media, hosted tools or WebSockets. Strict/omitted-strict functions, JSON Schema/JSON-object output and custom/Lark tools are rejected. Raw GBNF belongs to a separate native contribution and is not needed to use this adapter. Reasoning summaries require another generation pass and count toward the output allowance.

## Before merge

- [ ] Correct original-prompt cache accounting when the shared service wraps a reasoning-budget continuation; coordinate with #615 and add a Responses regression. A fresh original prompt can currently be reported as fully cached because the last continuation's cache count is clamped to its length.
- [ ] Reconcile visible reasoning replay with the guide: removing only `encrypted_content` from an item that also has raw content and a summary currently rejects the nonempty summary. Full returned-item encrypted replay passes.
- [ ] Qualify an isolated native run on this Responses branch, followed by a genuine model-generated Codex read/edit/verify loop before claiming that workflow.
- [ ] Reduce repeated evidence/type-catalog material in the upstream diff while retaining the full public archive and representative regression fixtures.
- [ ] Put the replay-key step before server startup in the quickstart and provide one complete matching client configuration.

Reviewed head: `46a44c51782eae4119b85b27c5b3639bc870e0a2` on `CC-David-CC:work/responses-api`. Its R4 implementation is `96092670da0dc3c1cfcce99bb90a3e6ca25ae1d9`; the tip adds the checkpoint receipt.

Mehr auf der Site

Links zu Install, Modellen, Releases.