Pull requests / #1324

#1324 file tier: an LRU part of the RAM budget (--resident-lru-gib), elastic under memory pressure (Windows)

open · @malloc32 · 0 评论 · 在 GitHub 查看

BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

描述

**AI-written.** This PR was prepared by an AI (Claude Opus 5.5, via Claude Code) at the submitter's request: design, code and every measurement were done by the AI on the submitter's machine, and the submitter reviewed it before opening.

**Depends on #1323** (the #833 port is this branch's first commit); review that one first.

**Update 2026-10-07:** rebased onto 0.1.40.2 (`e8ca9af`) together with #1323. New commit: on Windows the elastic watcher keeps commit available as well as RAM (every LRU buffer is committed memory; here the commit charge reached 112-120 of 121 GiB while serving with 8-9 GB of RAM free). The release's I/O prefetch flag (`stage_pf_`) is cleared when the LRU frees a buffer. Measured on 0.1.40.2 with the production config: decode 72.9 tok/s against 71.4 for the 0.1.40.1-based build, prompts unchanged.

## What

`--resident-budget-gib` fills the RAM once from the global expert profile, so a conversation on another topic keeps reading the drive for experts that are cold in the profile but hot for it.

- `--resident-lru-gib L`: L GiB of the budget hold the experts the decode reads from the files, least recently used out first (the unbuffered reads' stage pool). The prompt path copies from it but never adds to it or refreshes it: a prompt chunk sweeps every cold expert once and would push out what the decode reuses.
- `STRATA_LRU_KEEP_FREE_GIB=G`: the LRU is elastic. A watcher frees the coldest buffers when the available RAM (on Windows also the available commit) falls below G (back to the OS at once, never paged) and grows the pool back a GiB at a time when RAM is free again. It never shrinks below 1 GiB.
- `STRATA_TIER_TRACE=<file>`: one line per expert served outside the GPU caches. It is how the policy was chosen: replaying a traced session against same-size policies, the decode-fed LRU cut the decode's drive reads by ~45%.

The serve log reports the LRU's hits and, when elastic, what it holds and what it gave back. Without the flag and the variable nothing changes.

## Measured

Full tables and method in `bench/results/2026-10-07-windows-ram-tier-lru/`. IQ3_S, 2x RTX 5060 Ti, i7-12700K, 64 GB:

- 10 GiB by profile + 4 GiB LRU: decode 47.0 tok/s, against 41.9 for 14 GiB by profile alone (the same RAM).
- 12 + 18 GiB elastic: decode ~64-65 tok/s, prompts ~950 / ~1,540 tok/s (8K / 32K).
- Another program taking 20 GiB gets it in ~6 s, unpaged; decode drops to ~50 and is back at 60-70 tok/s ~20 s after the RAM is freed.

Tried and not in this PR: offering idle buffers to Windows (`OfferVirtualMemory`) halved decode speed; a routing lookahead reading the predicted blobs tripled the bytes read.

## Not tested

The elastic watcher on Linux (it uses `MemAvailable`; only Windows was measured). HIP.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。