Pull requests / #1493

#1493 Experimental memory relief with idle unloading and resumable text requests

open · draft · @midhatn · 0 commentaires · Sur GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindows

Description

## Summary

Resizing expert caches cannot always return enough memory to another application: long prefill can retain borrowed storage, and fixed allocations remain after caches reach their floor. This draft adds cooperative control, opt-in active text-request parking, and pressure-triggered unloading **between requests**, followed by guarded reload admission.

An agent's compiler/tool interval is one use case: Strata can release its native and vision processes, then wait for capacity before the next call or image encoding. No Hermes-specific integration is required. This is bounded resource control and tested continuation, **not a general speedup or never-OOM guarantee**. All new policies remain opt-in.

The current branch includes upstream **v0.1.40.4**, [`6674a006`](https://github.com/Niko1221/Strata/commit/6674a0065fb96bacde33e3eb10f91a1df86f95f2). Its Pascal-specific changes are not presented as a measured RTX 4070 improvement. Earlier evidence below retains its original v0.1.40.3 source/binary identity.

## What changed

- Ported **medking82's [PR #726](https://github.com/Niko1221/Strata/pull/726)** as the live-memory foundation, retaining its author credit and MIT notices. Its RAM blocks, GPU VMM cache resizing, MEMORY acknowledgements, resource presets and UI account for a substantial part of this diff.
- Added cooperative prompt/decode boundaries: readers, stagers and borrowed views drain before resizing; prefill resumes from its committed position. Bounded background leases support waits/pacing, sleeping workers, heartbeats and cancellation. A STOP epoch closes queued-request admission races.
- Added CUDA/DXGI feedback, separate physical-RAM/commit admission and a bulk Windows process sampler. Stale readings cannot authorize growth/reload. Windows working-set peaks account for trimmed RSS without double-counting the startup arena restoration delta.
- Added active **text-request parking**: drain STOP through DONE, privately journal exact tokens, release the engine and wait for full-reload capacity. The HTTP parser, tool-call identity, FIFO and cancellation remain owned; re-prefill emits only new output. Generation errors do not authorize replay.
- Added **idle-pressure full unloading and load admission**, extending the existing lifecycle. The next caller supplies normal history; no response is replayed. Preparation reservations and renewed FIFO admission cover delayed image downloads, tool images, active-parking cancellation and manual/timed unload. Waiting heartbeats, cancellation/shutdown cleanup and bounded retries for recognized pre-READY capacity failures preserve ownership.
- Stopped publishing stale process-allocation metrics after unload, using the reporting observation from **miskahm's [PR #1093](https://github.com/Niko1221/Strata/pull/1093)**. Internal capability/identity information remains available for safe reload; its frontend changes are not copied.
- Kept the experimental request-level routing-cost gate separate. It requires qualified matched measurements; uncalibrated use retains the configured split. No local performance benefit has been demonstrated.

The journal is **not** a KV/recurrent-state snapshot. Windows requires a private directory, protected current-user/SYSTEM DACL and persistent ACL support. Idle unloading needs no journal and rereads existing model files, without creating another SSD model copy.

## Validation

Hardware: Windows, Ryzen 9 7940HS, RTX 4070 Laptop 8 GiB, 64 GiB DDR5-5600. Model/settings: ISTA-DASLab IQ3_S, 65,536 configured context, high main-request reasoning and vision configured. The [sanitized evidence packet](https://github.com/midhatn/Strata/tree/codex/adaptive-background-draft/bench/results/2026-10-08-adaptive-background) preserves revision identities, settings, fixtures, selected measurements and failed gates. Counts from different revisions are not combined.

### Current v0.1.40.4 idle-admission revision

The [idle-pressure evidence packet](https://github.com/midhatn/Strata/tree/codex/adaptive-background-draft/bench/results/2026-10-08-idle-pressure) identifies the final server source (`a97599d3`) and native binary (`af83775b`). The longer-context results below belong to the earlier campaign.

| Check | Result for the final submitted idle-admission revision |
|---|---|
| Python suite | 893 run: 886 passed, 7 skipped |
| Native checks | Five targeted tests passed on the rebuilt v0.1.40.4 native binary; full SHA-256 in the packet |
| Real idle pressure → unload → admitted reload | Original native and vision identities exited; both reloaded with new identities; arithmetic probe returned exactly 42; no owned process remained after cleanup |
| Waiting client cancellation and heartbeat | Real capacity heartbeat; disconnect released preparation in 0.514 s without loading; separate held request remained blocked for 5.817 s before pressure release |
| Actual Hermes compiler/tool feedback | Normal final assistant stop, real build and tool-only nonce verified; 2,096 build cases plus 104 independent cases; 190.442 s including controlled wait/reload, two main model calls |
| HTTP vision and tool compatibility | Real red-square vision and function-call/result roundtrip passed with idle policy enabled; final tool answer 42 |
| Sampled headroom | Hermes cycle: 5.623 GiB RAM / 294 MiB native VRAM; HTTP: 9.186 GiB / 342 MiB |

Idle acceptance uses a real four-job C++ build, a **separate real 3 GiB RAM holder**, and an extended tool-response wait. The 2.296-second build finishes before unloading: this tests continuity during a controlled tool wait, **not compiler acceleration or memory-heavy compilation**. The final assistant must quote the passing build and a fresh tool-only nonce. Preliminary results with incomplete vision-PID history do not replace final-revision lifecycle checks.

### Earlier cooperative-control and active-parking revision

These are the pre-v0.1.40.4 campaign results already recorded in this draft, **not rerun claims for the new idle path**. The real parking run used native SHA-256 `b17c3ef7a55ff1b00e217fb5a5251a23583734b2ce6e2d57d1423012215d83aa` and server SHA-256 `f714e58627696099ecda6969307c066fee597a6b28412d218805045a98ea0e45`.

| Check | Earlier observed result |
|---|---|
| Python / native | 851 tests: 844 passed, 7 skipped; five targeted native executables passed |
| 60,270-token prompt with real 2 GiB external RAM allocation | Returned 2,363 MiB resident cache in 7.156 s, by position 768; correct answer, same process, recovery and fresh-conversation checks passed |
| Repeated long prompt | 60,265 reused / 5 newly read tokens; correct result in 10.940 s versus 839.035 s cold. Retained reuse, not an upstream A/B speedup |
| Active park/resume | Available RAM increased 37.348 GiB and global GPU headroom 6,802.7 MiB; exact accepted prefix/seed/remaining budget verified; generated C++ passed 104 checks |
| Cancellation while parked | About 37 ms after cancellation in that trial, without reloading |
| HTTP | Vision answered the synthetic image correctly; actual function-call/result continuation returned 42 |
| Actual Hermes repair/build cycle | First attempt hit its six-turn cap after repairing code; same-session continuation received a four-job build with 2,096 passing cases. Independent 104-case check passed. Combined agent time 514.401 s / 11 recorded API calls |

The earlier parking trigger was injected; exit, memory release, capacity readings and reload were real. No physical exhaustion was induced. Output lengths differed, so 64.781 s baseline versus 116.098 s resumed is not an isolated overhead estimate. Sampled GPU headroom exceeded 250 MiB, without a continuous guarantee.

**Pacing showed no foreground benefit:** median 45.25 s off versus 45.81 s on. All four generated-code tasks passed, but differing output lengths prevent causal model-time comparisons. Failed gates remain recorded. The sampler's median 6.603 ms snapshot is a sensor measurement, not a tok/s gain. No matched nonadaptive Hermes run supports faster agent-completion claims.

## Limits and next step

This single-engine prototype retains a finite execution floor. Active parking is text-only; subsequent vision preparation is supported, active image migration is not. There is no crash/reconnect recovery, arbitrary tensor relocation, dynamic CPU worker resizing, multi-GPU scheduling or SSD selection. Future tokens may differ after re-prefill. Allocation races, device failures and client timeouts remain possible.

Admission waits for the **full measured startup profile**. Separate static 32/20 GiB resident-budget tests on the earlier `c465e86d` server with the same v0.1.40.4 native binary reduced peak RSS from 39.184 to 27.033 GiB and private commit from 46.599 to 34.368 GiB; the runner sampled at least 324 MiB native free VRAM in both. This shows smaller-profile feasibility, **not automatic selection or an admission certificate**. Measured envelopes, model/resource identity separation and failure-path tests remain work. Native VRAM figures come from the runner guard; missing observer fields remain documented.

Overlap checked: [#1117](https://github.com/Niko1221/Strata/pull/1117), [#1324](https://github.com/Niko1221/Strata/pull/1324), [#1461](https://github.com/Niko1221/Strata/pull/1461), [#1471](https://github.com/Niko1221/Strata/pull/1471), [#1480](https://github.com/Niko1221/Strata/pull/1480) and its #1271/#1269 dependencies. These are not imported/validated here; combining controllers needs shared allocator ownership. This remains a draft; the allocator foundation, cooperative control and parking changes can be split for integration review.

## Sources actually used

The [implementation provenance](https://github.com/midhatn/Strata/blob/codex/adaptive-background-draft/docs/ADAPTIVE_PROVENANCE.md) maps sources to concrete roles and license boundaries:

- **Adapted code:** medking82's PR #726, original snapshot [`15a59d4`](https://github.com/medking82/Strata/commit/15a59d4785f492b4df3fb78862373a5383696450), plus Strata's existing engine, expert pool/source, prefill, FIFO, lifecycle and parsers. Original author and MIT notices are retained.
- **Reporting idea:** miskahm's PR #1093 observation about stale allocation counters after unload; original integration, not its UI implementation.
- **Routing-gate design inspiration only:** [ATSInfer sections 4.3–4.4](https://arxiv.org/html/2607.10183v2) and [StarPU's data-aware performance models](https://starpu.gitlabpages.inria.fr/features.html). No external scheduler source, complete algorithm or benchmark claim is imported.
- **API/ABI references:** Microsoft DXGI, physical-memory/commit, NT process snapshot and ACL/filesystem contracts; NVIDIA CUDA memory reporting; psutil Windows counters/ABI; Python atomic replacement/flushing. Concrete links and affected files are in the provenance document. New ctypes/journaling integration is original code.

Other surveyed work is separated as related work, not credited as an incorporated implementation merely because it was read. Prepared with Codex; claims above concern this implementation's measured boundaries.

Sur le site

Liens install, modèles, releases.