Pull requests / #1489
#1489 Responses persistence and disk-only cache: CUDA/HIP integration
open · @CC-David-CC · 0 commentaires · Sur GitHub
Server & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux
Description
## Summary
Validate #1269, #1271 and #1480 together with instruction/reasoning replay corrections, experimental Responses persistence and optional generated summaries.
The point is to test the boundaries between these features: can a chat resume after another chat, a summary pass or a restart without losing its history or restoring the wrong inference state? If KV files are evicted or damaged, the server must rebuild from durable history. This provides a tested core for downstream profile-based applications.
## What changed
Preserves the original commit histories. Live testing found and fixed non-MTP snapshot failures: capture and restore now omit the absent draft ring. Adds regression coverage, a live probe and integration notes. Persistence and summary generation remain separate opt-ins.
## Simple examples
Enable server options:
```text
--experimental-responses-persistence --responses-store-path ./history
```
Add engine options to your existing configuration:
```text
--conversation-cache-disk-only --conversation-cache-spill-dir ./kv --conversation-cache-disk-mib 2048
```
Keep those directories across restarts. Generated summaries remain off by default.
Use your configured model name and API key. Each reply is capped at 128 tokens:
```python
from openai import OpenAI
c = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="your-key")
def ask(text, parent=None):
return c.responses.create(
model="strata", input=text, previous_response_id=parent,
max_output_tokens=128, reasoning={"effort": "none"})
# Code chat, then an independent research chat.
code = ask("Write a Python function that squares a number.")
research = ask("Explain why a control group matters in an experiment.")
# Continue the code chat without resending its history.
reply = ask("Add a test for a negative number.", code.id)
print(reply.output_text)
print(reply.id) # Save this ID; use it as parent after a server restart.
```
The same continuation call works after restart. KV reuse accelerates compatible prefixes; a missing cache does not erase the conversation.
## Validation and merge notes
| Check | CUDA / RTX 4090 | HIP / RX 7900 XTX |
|---|---:|---:|
| Native cache tests | 8 passed | 7 passed |
| Live integration checks | 33 passed | 33 passed |
| Cached tokens after restart | 1429 | 1429 |
Each backend ran 389 API tests with four optional skips; a dependency-enabled follow-up passed 25 tests. Live coverage includes branching, deletion, summaries on/off, cancellation, corrupt-cache fallback and eviction.
Validated engine: e6da7e11. Scope: Linux, one GPU, 8K context, int8 KV, MTP off. Maximum context, multi-GPU and live MTP remain unvalidated. This is functional validation, not a speed comparison.
Merge the combined candidate once, or use resulting main if its components land separately. Keep downstream profile selection and immutable revision binding in the application.
Detailed results: https://github.com/CC-David-CC/Strata-a5500/blob/work/responses-disk-cache-core/docs/RESPONSES_DISK_INTEGRATION.md
Sur le site
Liens install, modèles, releases.