Pull requests / #1493
#1493 Experimental cooperative memory relief and resumable text requests
open · draft · @midhatn · 0 comments · View on GitHub
BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindows
Description
## Summary When Strata shares a machine with another application, resizing expert caches alone cannot always release memory promptly: a long prefill can retain borrowed storage, and fixed execution allocations remain after caches reach their floor. This draft adds cooperative control boundaries and an opt-in way to park one active text request, release the engine, and continue the response after capacity recovers. The target use case is an agent that needs to leave resources for its own compiler or other foreground work. The current result is bounded resource control and tested continuation, **not a general speedup or a never-OOM guarantee**. All new policies remain opt-in. The base is v0.1.40.3 (`d5ea713`). ## What changed - Ported **medking82's [PR #726](https://github.com/Niko1221/Strata/pull/726)** as the live-memory foundation, retaining its author credit and MIT notices. Its RAM blocks, GPU VMM cache resizing, MEMORY acknowledgements, resource presets and UI account for a substantial part of this diff. - Added safe yield points during prompt processing and decode. Readers, stagers and borrowed cache views drain before resize; prompt processing resumes from its committed position. Bounded background leases support waits/pacing, sleeping workers, heartbeats and cancellation. A STOP epoch closes the queued-request admission race. - Added native CUDA/DXGI capacity feedback, physical-RAM/commit-aware admission, conservative recovery and a bounded bulk Windows process sampler. Missing or stale telemetry cannot authorize growth. - Added opt-in active **text-request parking**. The request owner drains STOP through DONE, commits a private token journal, unloads the engine and waits for measured full-reload capacity. The HTTP parser, tool-call identity, FIFO ownership and cancellation remain with the same request. After reload it reprocesses the exact accepted prefix and emits only new output. Ambiguous completion and generation errors do not authorize replay. - Kept a separate experimental request-level routing-cost gate. It needs qualified matched measurements; uncalibrated use retains the configured split. It has no demonstrated local performance benefit. The journal is **not** a KV/recurrent-state snapshot. Windows persistence requires a protected current-user/SYSTEM DACL and a filesystem supporting persistent ACLs. Parking is disabled by default and requires an explicit private journal directory. ## Validation Tested on Windows, Ryzen 9 7940HS, RTX 4070 Laptop 8 GiB and 64 GiB DDR5-5600, with ISTA-DASLab IQ3_S, 65,536 configured context, high main-request reasoning and vision configured. The [sanitized evidence packet](https://github.com/midhatn/Strata/tree/codex/adaptive-background-draft/bench/results/2026-10-08-adaptive-background) includes settings, source/binary hashes, selected measurements, fixtures and failed gates. | Check | Observed result | |---|---| | Python suite | 851 run: 844 passed, 7 skipped | | Native checks | Five targeted executables passed, including real CUDA VMM resize and queued STOP coverage | | 60,270-token prompt with real 2 GiB external RAM allocation | Released 2,363 MiB of resident cache in 7.156 s, by position 768; correct answer, same process, recovery and fresh-conversation checks passed | | Repeated long prompt | 60,265 reused / 5 newly read tokens; correct result in 10.940 s versus 839.035 s cold. This demonstrates retained reuse, not an upstream A/B speedup | | Active park/reload/resume | Released 37.348 GiB available RAM and 6,802.7 MiB global GPU headroom; exact accepted prefix/seed/remaining budget verified; generated C++ passed 104 checks | | Cancellation while parked | Returned about 37 ms after cancellation in this trial, without reloading | | Ordinary HTTP compatibility | Vision answered the synthetic image correctly; actual function-call/result continuation returned 42 | | Actual visible Hermes compiler cycle | First attempt hit its six-turn cap after repairing code. Same-session continuation received a new successful four-job build with 2,096 passing cases; independent 104-case check also passed. Combined active agent time 514.401 s / 11 recorded API calls | The parking pressure trigger was deliberately injected; memory readings, process exit, release, admission and reload were real. No intentional physical exhaustion was used. Baseline and resumed answers had different lengths, so 64.781 s versus 116.098 s is not an isolated suspension-overhead estimate. Sampled native GPU headroom remained above 250 MiB; samples are not a continuous guarantee. The foreground pacing A/B **did not show a benefit**: median foreground time was 45.25 s off versus 45.81 s on. All four final generated-code tasks passed, but different output lengths prevent attributing shorter model elapsed times to pacing. An earlier incomplete-telemetry campaign and its output-limit failure are retained. The Windows sensor improved coverage and reduced median snapshot time to 6.603 ms; that is a sensor result, not a tok/s gain. There is no matched nonadaptive Hermes run and no claim of faster agent completion. ## Scope, overlap and remaining work This is a single-engine prototype. Active parking supports text, not image-request migration, crash/reconnect recovery, arbitrary tensor relocation or multi-GPU scheduling. It does not resize CPU worker counts or choose SSDs. Future output after re-prefill is not guaranteed bit-identical. The current controller can reclaim caches while idle, but does **not** add pressure-triggered full idle unloading. Upstream already has timed idle unload/autoload; integrating pressure admission before the next load and vision preparation remains work. Inference still needs a finite execution working set. Sudden external allocations, device failures and client timeouts remain possible. I checked overlap with [#1117](https://github.com/Niko1221/Strata/pull/1117), [#1324](https://github.com/Niko1221/Strata/pull/1324), [#1461](https://github.com/Niko1221/Strata/pull/1461), [#1471](https://github.com/Niko1221/Strata/pull/1471) and [#1480](https://github.com/Niko1221/Strata/pull/1480) (including its #1271/#1269 dependencies). Those patches are not imported or validated here. A shared allocator/policy owner is needed before combining independently acting controllers. I am opening this as a draft for scope and integration review; the allocator foundation, cooperative control and parking work can be separated if that fits the maintainers' direction better. ## Sources actually used The [implementation provenance](https://github.com/midhatn/Strata/blob/codex/adaptive-background-draft/docs/ADAPTIVE_PROVENANCE.md) maps every reused source to its concrete role and license boundary: - **Adapted code:** medking82's PR #726, original snapshot [`15a59d4`](https://github.com/medking82/Strata/commit/15a59d4785f492b4df3fb78862373a5383696450), plus the existing Strata engine, expert source/pool, prefill, FIFO, lifecycle and parsers. Original author and MIT notices are retained. - **Design inspiration only:** [ATSInfer, sections 4.3–4.4](https://arxiv.org/html/2607.10183v2) and [StarPU's data-aware performance models](https://starpu.gitlabpages.inria.fr/features.html), used in the optional measured routing gate. No external scheduler implementation, complete algorithm or benchmark claim was imported. - **API/ABI references:** Microsoft DXGI budget, memory-status, NT process-snapshot and ACL/filesystem contracts; NVIDIA CUDA memory reporting; psutil's pinned Windows ABI reference; Python atomic file replacement and flushing. Every concrete link and affected file is listed in the provenance document. The ctypes integration is original code, not copied sample implementations. Other surveyed papers, repositories and PRs are explicitly separated as related work. They are not credited as incorporated implementations merely because they were read. Prepared with Codex; the evidence and limitations above are specific to this submitted implementation.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.