Pull requests / #726

#726 Add opt-in live cache resizing and desktop resource presets

closed · @medking82 · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

描述

An opt-in memory policy resizes expert RAM/VRAM caches while the model remains loaded, preserving the native process, generation session and KV state. This contribution also adds Automatic, Full, Daily, Busy and Off resource controls to the existing full Monitor. The policy and resource presets remain disabled by default; request-boundary reload mode and the original percentage policy remain available.



## Desktop resource presets



- Full targets 2 GiB free RAM and a 256 MiB VRAM reserve; Daily targets 4 GiB / 700 MiB; Busy targets 8 GiB / 1,536 MiB. These are reserves, not actual cache allocations or throughput guarantees.

- Automatic starts in Daily, selects Daily while Codex/ChatGPT is present, selects Busy for sustained external application CPU/RSS load, and selects Full after confirmed absence and low external load. Conservative transitions require 8 seconds; relaxed transitions require 60 seconds of valid advancing readings. Manual selection persists until Automatic is restored; Off restores the configured percentage policy.

- The optional sampler uses process identity, CPU counters and RSS in the existing telemetry loop, excludes the server/native process family, and exposes only sanitized preset status. Absent or disabled presets do not sample processes. Missing, stale, incomplete or replayed readings cannot earn a relaxed transition.

- Stable fixed-preset growth requires 30 seconds of fresh headroom and proposes at most +2 GiB RAM / 128 MiB VRAM reserve reduction per step. The configured sustained-pressure window still wins. Retargeting preserves actual allocation, pending acknowledgements and the native-error growth retry gate; no selection unloads or restarts the model.



`GET/POST /v1/resources` retains API-key admission, own-page JSON origin checks, bounded request bodies and no wildcard CORS for resource controls. Only a complete validated `enabled`/`selection` block is saved, atomically, before live retargeting; the existing backup is preserved. A failed save cannot change the live selection. The UI distinguishes saved targets from actual, pending, limited and failed native allocations, and guards stale metrics responses after a selection change.



## Existing live-policy behavior retained



Matching applied acknowledgements own actual cache sizes. Sustained pressure never grows the other cache; stale capabilities, failed writes and partial/error acknowledgements cannot fabricate successful recovery. Live mode rejects parallel batch capacity owners and enabled `--vram-elastic`; those upstream features remain available with the policy disabled. Safe reader boundaries, guarded prefill borrowing, STOP/EOS and complete continuation reuse remain documented in [LIVE_MEMORY.md](https://github.com/medking82/Strata/blob/15a59d4785f492b4df3fb78862373a5383696450/docs/LIVE_MEMORY.md).



Optional `recovery_seconds` defaults to 0; nonzero values must be 30–3,600 seconds. Under the original percentage policy, only matching applied RAM-pressure relief arms recovery toward the actual pre-pressure ceiling, at most +2 GiB per fresh qualified window. Limited acknowledgements stop fast recovery, and native errors clear the episode without bypassing the growth retry deadline. Ordinary percentage-policy expansion retains its 120-second debounce and 600-second cooldown. Fixed presets use their separate bounded 30-second path.



The omitted `min_ram_headroom_gib` default remains 5.5 GiB; explicit values from 2 GiB are accepted. The pressure default remains 60 seconds; explicit values from 2 seconds are accepted. Percentage RAM targets retain both their percentage and absolute floors. An explicit VRAM floor of 0 retains percentage headroom and native admission/workspace guards, including live startup's 256 MiB minimum. The policy and presets do not change the model, quantization, context/KV allocation or sampling settings.



## Upstream compatibility

Based on upstream main `1735d6471df29b42c26170efaac1f1446a58640f`. The integration preserves upstream exchange rotation, resident layer-split support, lookup-chain drafting, fatal-engine recovery and slot-save configuration. Live RAM blocks keep their own ownership path, including the public resident accessor. Live mode explicitly rejects `--adapt-async 1`, whose split exchange path assumes a fixed resident arena; that upstream mode remains available without live resizing. Chained drafts remain bounded by the request's remaining output allowance, preserving the emitted continuation prefix.

## Validation

Contribution head: `15a59d4785f492b4df3fb78862373a5383696450`.

- Full Windows CPU server suite: 558 tests, 6 fixture skips, no failures/errors (166.017 seconds). The existing unclosed-file ResourceWarning remains visible. Focused memory/presets/fatal-recovery/slots: 163 tests, 1 skip, passed. The first attempt used a gateway environment without Jinja2; the checks above used the existing Strata environment.
- Parallel setup: 3 tests passed. Browser resource controls: 12 Node tests passed.
- MSVC 19.44.35229 / CUDA 13.0.48: changed `generate.cpp`, `expert_source.cpp` and `live_memory_test.cpp` translation units compiled.
- Eight CPU CTests passed (5.05 seconds), including 10,816 serve-window checks, 12,304 exchange-storage exchanges and 907,257 coupled-draft checks. The continuation fixture now covers chained drafts and EOS.
- Contribution whitespace checks relative to upstream main passed. Two whitespace lines in an inherited upstream benchmark are unchanged.
- No full GPU link, model load or new throughput benchmark was performed for this refresh. The new live-resident accessor fixture was compile-checked; its GPU-dependent full test was not run. Existing daily runtime source/configuration and native artifacts were not changed.

## Earlier daily runtime observation

The following measurements are historical evidence from the separately deployed daily source, not performance validation of the refreshed contribution head.




Functional GPU/browser acceptance used separately deployed daily source `f1183b586ba1a25bb88485b50471b6ee0f2f907b`, with its separate sidebar/context features and unchanged native artifact. Eight model runs retained the same native PID/create-time and 64K context without native error; HTTPS Monitor selections Daily/Full/Busy and Automatic saved and read back correctly, and Automatic was restored. With Thinking Off, approximately 6,765 input tokens and two 512-output runs per preset, output-weighted decode rates were Full 63.4, Daily 62.0 and Busy 53.7 tokens/s. These capped Off samples have no matched pre-change baseline and do not establish a controller speedup or whole-request High throughput.



The uncapped High-thinking Auto/Busy request completed at natural EOS: 33,159 input / 5,741 generated tokens, averaging 32.2 decode tokens/s over the whole request (208.502 seconds wall time). The requested 50+ whole High target was not achieved. Native resizing cost, cache geometry/warming, file misses, output length, speculative acceptance and desktop load remain relevant; successful functional selection is not a universal performance claim. Earlier real-request limitations and measurements remain documented with their original inputs. Only aggregate observations are published; no private conversation, credentials or runtime configuration is included. The daily branch must not be replaced directly by this contribution branch.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。