Pull requests / #1493
#1493 Experimental memory relief with idle unloading and resumable text requests
open · draft · @midhatn · 0 commentaires · Sur GitHub
BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindows
Description
## Summary Resizing expert caches cannot always return enough memory to another application: long prefill can retain borrowed storage, and fixed allocations remain after caches reach their floor. This draft adds cooperative control, opt-in active text-request parking, and pressure-triggered unloading **between requests**, followed by guarded reload admission. An agent's compiler/tool interval is one use case: Strata can release its native and vision processes, then wait for capacity before the next call or image encoding. No Hermes-specific integration is required. This is bounded resource control and tested continuation, **not a general speedup or never-OOM guarantee**. All new policies remain opt-in. The current branch includes upstream **v0.1.40.4**, [`6674a006`](https://github.com/Niko1221/Strata/commit/6674a0065fb96bacde33e3eb10f91a1df86f95f2). Its Pascal-specific changes are not presented as a measured RTX 4070 improvement. Earlier evidence below retains its original v0.1.40.3 source/binary identity. ## What changed - Ported **medking82's [PR #726](https://github.com/Niko1221/Strata/pull/726)** as the live-memory foundation, retaining its author credit and MIT notices. Its RAM blocks, GPU VMM cache resizing, MEMORY acknowledgements, resource presets and UI account for a substantial part of this diff. - Added cooperative prompt/decode boundaries: readers, stagers and borrowed views drain before resizing; prefill resumes from its committed position. Bounded background leases support waits/pacing, sleeping workers, heartbeats and cancellation. A STOP epoch closes queued-request admission races. - Added CUDA/DXGI feedback, separate physical-RAM/commit admission and a bulk Windows process sampler. Stale readings cannot authorize growth/reload. Windows working-set peaks account for trimmed RSS without double-counting the startup arena restoration delta. - Added active **text-request parking**: drain STOP through DONE, privately journal exact tokens, release the engine and wait for full-reload capacity. The HTTP parser, tool-call identity, FIFO and cancellation remain owned; re-prefill emits only new output. Generation errors do not authorize replay. - Added **idle-pressure full unloading and load admission**, extending the existing lifecycle. The next caller supplies normal history; no response is replayed. Preparation reservations and renewed FIFO admission cover delayed image downloads, tool images, active-parking cancellation and manual/timed unload. Waiting heartbeats, cancellation/shutdown cleanup and bounded retries for recognized pre-READY capacity failures preserve ownership. - Stopped publishing stale process-allocation metrics after unload, using the reporting observation from **miskahm's [PR #1093](https://github.com/Niko1221/Strata/pull/1093)**. Internal capability/identity information remains available for safe reload; its frontend changes are not copied. - Kept the experimental request-level routing-cost gate separate. It requires qualified matched measurements; uncalibrated use retains the configured split. No local performance benefit has been demonstrated. The journal is **not** a KV/recurrent-state snapshot. Windows requires a private directory, protected current-user/SYSTEM DACL and persistent ACL support. Idle unloading needs no journal and rereads existing model files, without creating another SSD model copy. ## Validation Hardware: Windows, Ryzen 9 7940HS, RTX 4070 Laptop 8 GiB, 64 GiB DDR5-5600. Model/settings: ISTA-DASLab IQ3_S, 65,536 configured context, high main-request reasoning and vision configured. The [sanitized evidence packet](https://github.com/midhatn/Strata/tree/codex/adaptive-background-draft/bench/results/2026-10-08-adaptive-background) preserves revision identities, settings, fixtures, selected measurements and failed gates. Counts from different revisions are not combined. ### Current v0.1.40.4 idle-admission revision The [idle-pressure evidence packet](https://github.com/midhatn/Strata/tree/codex/adaptive-background-draft/bench/results/2026-10-08-idle-pressure) identifies the final server source (`a97599d3`) and native binary (`af83775b`). The longer-context results below belong to the earlier campaign. | Check | Result for the final submitted idle-admission revision | |---|---| | Python suite | 893 run: 886 passed, 7 skipped | | Native checks | Five targeted tests passed on the rebuilt v0.1.40.4 native binary; full SHA-256 in the packet | | Real idle pressure → unload → admitted reload | Original native and vision identities exited; both reloaded with new identities; arithmetic probe returned exactly 42; no owned process remained after cleanup | | Waiting client cancellation and heartbeat | Real capacity heartbeat; disconnect released preparation in 0.514 s without loading; separate held request remained blocked for 5.817 s before pressure release | | Actual Hermes compiler/tool feedback | Normal final assistant stop, real build and tool-only nonce verified; 2,096 build cases plus 104 independent cases; 190.442 s including controlled wait/reload, two main model calls | | HTTP vision and tool compatibility | Real red-square vision and function-call/result roundtrip passed with idle policy enabled; final tool answer 42 | | Sampled headroom | Hermes cycle: 5.623 GiB RAM / 294 MiB native VRAM; HTTP: 9.186 GiB / 342 MiB | Idle acceptance uses a real four-job C++ build, a **separate real 3 GiB RAM holder**, and an extended tool-response wait. The 2.296-second build finishes before unloading: this tests continuity during a controlled tool wait, **not compiler acceleration or memory-heavy compilation**. The final assistant must quote the passing build and a fresh tool-only nonce. Preliminary results with incomplete vision-PID history do not replace final-revision lifecycle checks. ### Earlier cooperative-control and active-parking revision These are the pre-v0.1.40.4 campaign results already recorded in this draft, **not rerun claims for the new idle path**. The real parking run used native SHA-256 `b17c3ef7a55ff1b00e217fb5a5251a23583734b2ce6e2d57d1423012215d83aa` and server SHA-256 `f714e58627696099ecda6969307c066fee597a6b28412d218805045a98ea0e45`. | Check | Earlier observed result | |---|---| | Python / native | 851 tests: 844 passed, 7 skipped; five targeted native executables passed | | 60,270-token prompt with real 2 GiB external RAM allocation | Returned 2,363 MiB resident cache in 7.156 s, by position 768; correct answer, same process, recovery and fresh-conversation checks passed | | Repeated long prompt | 60,265 reused / 5 newly read tokens; correct result in 10.940 s versus 839.035 s cold. Retained reuse, not an upstream A/B speedup | | Active park/resume | Available RAM increased 37.348 GiB and global GPU headroom 6,802.7 MiB; exact accepted prefix/seed/remaining budget verified; generated C++ passed 104 checks | | Cancellation while parked | About 37 ms after cancellation in that trial, without reloading | | HTTP | Vision answered the synthetic image correctly; actual function-call/result continuation returned 42 | | Actual Hermes repair/build cycle | First attempt hit its six-turn cap after repairing code; same-session continuation received a four-job build with 2,096 passing cases. Independent 104-case check passed. Combined agent time 514.401 s / 11 recorded API calls | The earlier parking trigger was injected; exit, memory release, capacity readings and reload were real. No physical exhaustion was induced. Output lengths differed, so 64.781 s baseline versus 116.098 s resumed is not an isolated overhead estimate. Sampled GPU headroom exceeded 250 MiB, without a continuous guarantee. **Pacing showed no foreground benefit:** median 45.25 s off versus 45.81 s on. All four generated-code tasks passed, but differing output lengths prevent causal model-time comparisons. Failed gates remain recorded. The sampler's median 6.603 ms snapshot is a sensor measurement, not a tok/s gain. No matched nonadaptive Hermes run supports faster agent-completion claims. ## Limits and next step This single-engine prototype retains a finite execution floor. Active parking is text-only; subsequent vision preparation is supported, active image migration is not. There is no crash/reconnect recovery, arbitrary tensor relocation, dynamic CPU worker resizing, multi-GPU scheduling or SSD selection. Future tokens may differ after re-prefill. Allocation races, device failures and client timeouts remain possible. Admission waits for the **full measured startup profile**. Separate static 32/20 GiB resident-budget tests on the earlier `c465e86d` server with the same v0.1.40.4 native binary reduced peak RSS from 39.184 to 27.033 GiB and private commit from 46.599 to 34.368 GiB; the runner sampled at least 324 MiB native free VRAM in both. This shows smaller-profile feasibility, **not automatic selection or an admission certificate**. Measured envelopes, model/resource identity separation and failure-path tests remain work. Native VRAM figures come from the runner guard; missing observer fields remain documented. Overlap checked: [#1117](https://github.com/Niko1221/Strata/pull/1117), [#1324](https://github.com/Niko1221/Strata/pull/1324), [#1461](https://github.com/Niko1221/Strata/pull/1461), [#1471](https://github.com/Niko1221/Strata/pull/1471), [#1480](https://github.com/Niko1221/Strata/pull/1480) and its #1271/#1269 dependencies. These are not imported/validated here; combining controllers needs shared allocator ownership. This remains a draft; the allocator foundation, cooperative control and parking changes can be split for integration review. ## Sources actually used The [implementation provenance](https://github.com/midhatn/Strata/blob/codex/adaptive-background-draft/docs/ADAPTIVE_PROVENANCE.md) maps sources to concrete roles and license boundaries: - **Adapted code:** medking82's PR #726, original snapshot [`15a59d4`](https://github.com/medking82/Strata/commit/15a59d4785f492b4df3fb78862373a5383696450), plus Strata's existing engine, expert pool/source, prefill, FIFO, lifecycle and parsers. Original author and MIT notices are retained. - **Reporting idea:** miskahm's PR #1093 observation about stale allocation counters after unload; original integration, not its UI implementation. - **Routing-gate design inspiration only:** [ATSInfer sections 4.3–4.4](https://arxiv.org/html/2607.10183v2) and [StarPU's data-aware performance models](https://starpu.gitlabpages.inria.fr/features.html). No external scheduler source, complete algorithm or benchmark claim is imported. - **API/ABI references:** Microsoft DXGI, physical-memory/commit, NT process snapshot and ACL/filesystem contracts; NVIDIA CUDA memory reporting; psutil Windows counters/ABI; Python atomic replacement/flushing. Concrete links and affected files are in the provenance document. New ctypes/journaling integration is original code. Other surveyed work is separated as related work, not credited as an incorporated implementation merely because it was read. Prepared with Codex; claims above concern this implementation's measured boundaries.
Sur le site
Liens install, modèles, releases.