Issues / #1571
#1571 [Bug]: Windows --resident-experts: building the RAM copy keeps experts.bin's mapped pages in the working set until the last layer, available RAM drops to ~0 on a 48 GB PC
open · @LaciciBachuchu · 0 Kommentare · Auf GitHub
Setup & installNVIDIA / CUDAModels & quantsWindowsLinux
Beschreibung
### What happened
On Windows with 48 GB of RAM, `--resident-experts` drives the PC's available RAM to almost zero for about 20 seconds at the end of every load, while the RAM copy (cache complement) is built. The engine's own sizing is fine (it leaves the 4 GiB headroom against the RAM it finds free); the extra comes from the `experts.bin` pages the copy reads through the file mapping, which stay in the process working set until the copy is finished.
From the source (v0.1.40.3, `src/core/expert_source.cpp`, `FileExpertSource::pin_cache_complement`):
- mapped reads are used when `direct_` is empty (line 2465), which is the case in the resident mode: only the RAM-budget path calls `set_unbuffered` first (`program/generate.cpp:4062-4064`);
- on Linux each copied layer is released right away (`madvise(MADV_DONTNEED)` + `posix_fadvise`, around line 2510, inside `#if !defined(_WIN32)`);
- on Windows the mapped pages are only released after the whole copy, with `VirtualUnlock` on the full mapping (line 2543).
So on Windows the source pages of all 48 layers pile up in the working set on top of the page-locked copy. With ~30 GB available before the copy, that is ~25 GB pinned plus up to ~25 GB of mapped pages. Windows does not count those clean file pages as available until it trims them, and it only trims once memory is nearly gone.
Measured (engine log + `\Memory\Available MBytes` sampled every ~3 s, strata.exe working set):
```
FileExpertSource: available RAM 19.00 GiB, 30.18 GiB after the mapped experts left the process working set (#467)
FileExpertSource: allocating 25.46 GiB page-locked cache complement
FileExpertSource: before complement allocation: requested 25.46 GiB, available RAM 30.18 GiB, available commit 54.48 GiB (limit 95.71 GiB), ...
FileExpertSource: copied cache complement through layer 8/48 (5.76 GiB, 6 s)
FileExpertSource: copied cache complement through layer 16/48 (9.54 GiB, 11 s)
FileExpertSource: copied cache complement through layer 24/48 (13.28 GiB, 15 s)
time available MB strata.exe working set MB
14:37:57 20512 13149 (complement allocated)
14:38:00 4781 27843
14:38:03 1908 31128
14:38:07 219 27443
14:38:11 131 19602
14:38:16 766 19962
```
A second load (started from our launcher) showed the same shape: available RAM 5.6 -> 1.9 -> 0.1 -> 0.9 -> 0.6 -> 0.6 GB over ~20 s, then 6.9 and 13.6 GB once the load finished. A third load, while the desktop and a browser were active, ended with an NVIDIA driver reset (System log: nvlddmkm event 153) about 83 s into the load, right where this happens. That one is a correlation, not a proof, but it is the only nvlddmkm event on this machine in months; ECC counters did not change.
What confirms the cause: trimming strata.exe's working set from outside during the load (`SetProcessWorkingSetSize(handle, -1, -1)` every 0.4 s, which only moves unlocked pages to the standby list) kept available RAM at 9.7 GB or more for the whole load, with the same load time (~70 s) and the same `3007 of the prompt path's 3007 lendable slots keep their experts in RAM too`. It is not a workaround I can keep, though: an antivirus heuristic flagged the helper that opens another process and changes its working set.
The workaround we use now is `--resident-budget-gib 27` with `"env": {"STRATA_UNBUFFERED_LOAD": "1"}`. The copy then goes through `read_direct` and available RAM never drops below ~4.6 GB, but a budget turns the lend region off (`lend_from_slot = -1`, around line 2147: "A budget turns `lend` off"), so every prompt re-reads the lent experts from the drive. The same model, cache and args otherwise:
| prompt | resident mode (lend kept in RAM) | budget mode (lend read from the drive) |
|---|---|---|
| 10K-token agent system prompt | 4.0 s | 19.6 s (15.5 GB read from the file) |
| 30K | 8.4 s | 27 s |
| 60K | 17 s | 48 s |
| 200K | 58 s | 137 s |
| 250K | 72 s | 158-169 s |
Decode speed is the same in both modes.
### Suggested fix
1. Windows: in the complement copy loop, release each layer's mapped range after it is copied: `VirtualUnlock` on the layer's span, as `FileExpertSource::release(layer, expert)` already does per expert and as the Linux branch does with `madvise` per layer. Alternatively, use the `read_direct` path for the resident copy too.
2. Optional: let `--resident-budget-gib` keep the lend region in RAM when the budget covers it, so the RAM-budget mode does not pay the drive reads per prompt.
### Related observations (same PC, not separate bugs)
- **Commit limit:** with Windows' system-managed page file (~14 GB here), the resident copy got 0 GiB: `resident complement 22.30 GiB exceeds available commit (2.28 GiB) minus the 4 GiB safety headroom`. The ~22 GB of VRAM counts against commit under WDDM, and the desktop apps held ~40 GB of commit. The engine fell back to plain mapped mode with only a log line, and setup did not warn. A fixed 48 GB page file fixed it. A setup tip, for example "RAM + page file must cover the GPU's VRAM plus the RAM copy", would have saved some time.
- **Setup's mode choice:** setup picked the normal (non-low-RAM) mode from total RAM (47.7 GB is at least 35.5 + 10 for IQ2_XS), but with the desktop open only ~20-30 GB is available, so the normal mode cannot hold the arena. Passing `--low-ram resident` was needed. Perhaps setup could look at available RAM too.
- **`tools/iq_pack.py --experts-bin`:** the `np.memmap` read of the 39 GB GGUF grew the packer's working set to ~18 GB (available RAM 21 -> 6.5 GB) before we capped its working set at 3 GB, after which it finished normally in ~2.5 min. This is the same Windows behaviour (mapped pages stay in the working set) in the setup step.
### Strata version, GPU, OS
0.1.40.3 (ready-made engine, `strata-windows-x64.zip`) · RTX 4090 24 GB (ECC on, 22.5 GB usable), driver 616.92, PCIe 4.0 x16 · i7-14700K (AVX2) · 48 GB DDR5-5600 · Windows 11 Pro (26200) · NVMe
Model: Qwen3.8-Flash-Next GSQ-RCO IQ2_XS. Engine args:
```
--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <data>\mtp\rt --max-context 262144 --kv int8
--kv-resident 32768 --resident-experts --vram-reserve-mib 1024 --pool-workers 19 --pcie-frac 0
--conversation-cache-mib 4096 --conversation-cache-slots 4
```
(`--pcie-frac 0` and `--pool-workers 19` were faster than the probe's 0.55 and setup's 13 on this CPU: +16-32% and +3-10% decode, the same direction as #308.)
Thanks for Strata. With the lend region in RAM it reads 200K prompts faster than the dense 27B we use on the same card.
Mehr auf der Site
Links zu Install, Modellen, Releases.