Pull requests / #727
#727 Add opt-in durable chat history, legacy import and browser compaction
closed · draft · @medking82 · 0 comentarios · En GitHub
Server & APIModels & quantsSecurityWindows
Descripción
This is a **Draft for integration discussion** because durable chat storage and compaction overlap [#696](https://github.com/Niko1221/Strata/pull/696), [#680](https://github.com/Niko1221/Strata/pull/680) and [#350](https://github.com/Niko1221/Strata/pull/350). It offers a bounded alternative focused on preserving full legacy branches and making retained history available through read-only MCP memory. Maintainers can assess whether to combine the pieces with those proposals. The existing web chat retains a single browser conversation. With an explicit `chat_archive_path` or `--chat-archive`, this adds a local SQLite archive, conversation and branch selection, Rename, import and full JSON backup. Without that option, existing single-chat/New chat/Undo behavior remains available. ## Behavior and preservation - Transactional, revision-checked saves prevent stale tabs from overwriting newer history. Import is atomic and existing destination IDs always win on reimport. Original legacy graphs, all branches and the selected `currNode` path are retained. - Read-only import supports same-origin `LlamaUi`/`LlamacppWebui` IndexedDB, exported JSON/JSONL and Strata backups. Baseline Strata browser records with name-only attachments retain their raw source and visibly mark missing content; malformed or unsupported legacy formats fail atomically. Remote image URLs are retained without fetching them merely to display history; historical tool calls remain archive data. - Browser Auto compact and Compact now summarize earlier rounds while keeping recent rounds verbatim and original messages in the archive. Restore full context uses those originals. Template-aware counting includes offered tools, image reserve and reply headroom. Raw API code/diff/review requests are not silently compacted. - Original persistence is required before summary generation. Storage failure and save conflict retain original context, and epoch guards prevent delayed saves or file-read callbacks from changing another selected draft. - Builtin `memory__search` and `memory__recall` reuse the existing MCP loop, preserve configured providers, and return bounded read-only text with chat/branch/message coordinates and continuation offsets. Literal search supports Chinese text; image bytes are omitted from tool results. - Private archive routes require the server's own host/port origin (HTTP or HTTPS) and configured API key and never inherit wildcard API CORS or foreign `trusted_origins`. A reverse proxy must preserve Host. No account system, filesystem coding agent, external storage service or new inference provider is added. ## Limits The 100 MB cap applies to the whole SQLite database, including duplicated search text, so usable payload capacity is lower; memory results are limited to 12,000 characters. Capacity errors return safe JSON and preserve committed history. Without a configured archive, oversized full-media browser saves preserve the text fallback and display an explicit warning. Retrieval is selective and can miss facts; persistence is not a disk-failure backup. API traffic is not automatically archived. Drafts persist on chat saves/actions; this does not add continuous autosave of unsent edits. Browser compaction is lossy, although full originals remain available. Oversized recent turns may still need shortening. This contribution is independent of document extraction and adaptive memory budgeting and leaves C++ kernels, model precision, context configuration and lazy vision unchanged. The `/strata` alias and retirement worker support the transition from an old root UI without deleting legacy browser databases or caches. ## Current upstream integration Head `23f4e24809ccf0fdbcf7a511484b1b23e36d4fc9` integrates upstream `main=6f32ec070f23ced9f50e704d854d775da52591ab` (0.1.39). The server import/POST conflicts retain both the archive routes and upstream Responses API. Upstream Monitor Conversation cache/MCP and About Model settings remain available. The PR remains Draft because the integration/overlap discussion above is still open. Template-aware counting now uses upstream `Service.encode_prompt`, matching actual inference when prompts quote thinking tags or use the trailing effort turn. Two HTTP regressions reproduced the previous undercounts (47 versus 60, and 191 versus 408) before the fix; both now match completion usage. Counting still does not load the model or invoke tools/vision. ## Current validation - Windows: 40 focused archive, memory, HTTP admission and template-count tests passed. - Windows: 240 affected server, MCP, security, lifecycle, Monitor, runconfig, Responses and structured-output tests passed. One existing upstream unclosed-socket ResourceWarning remains visible. - 32 Node tests passed, plus 46 import assertions, covering browser-source lifecycle, quota/reload fallback, conflicts, stale callbacks, legacy links and Rename preservation. - Browser script/Python syntax checks and the feature diff check against current upstream passed. A generic merge check against the historical branch HEAD reported inherited upstream SYCL whitespace; those files were preserved rather than cleaned up. - Isolated synthetic browser acceptance verified send, Rename, New chat retaining prior history, reopening and reload persistence, Monitor Conversation cache/MCP, and About Model settings. Console showed no warning/error entries. No production archive, GPU or real model was used. The earlier frozen archive fixture additionally verified all 16 originals and both branches survived compaction, Chinese Rename/reload, Restore, inactive branch selection and repeated legacy import. Those older results are retained separately; manual import, branch selection and mobile layouts were not repeated in this integration, and the Node import/library/context checks passed. One independent Native Review completed with Claude Opus 5.5 on frozen commit `bba531515cc7d8032fddf514b1bbf60c6fd15144`: three P2 and two P3 findings, no P0/P1. Its same-round Gemini Flash/high attempt failed with `native_output_truncated`; that failure is retained and is not treated as completed review evidence. All five findings were adjudicated and corrected in follow-up commit `ba290b6afe141af6271e5fcee075562794d8a5ab`, with its then-current deterministic validation rerun. That Native Review is historical evidence for its frozen input; the current bounded upstream integration has the separate checks above and was not submitted to another model review.
En el sitio
Enlaces a install, modelos, releases.