Pull requests / #286
#286 Low RAM on Windows: a pinned tier of the most-read experts, the rest read unbuffered from experts.bin
closed · @sergqwer · 0 commentaires · Sur GitHub
BenchmarksSetup & installNVIDIA / CUDAWindowsLinux
Description
## Summary Stacked on #285 (it uses its deferred per-layer registration). Only the last commit is new here. The resident arena holds every expert, including the VRAM cache's share of them a second time. A PC whose RAM cannot hold them beside the system has setup's low-RAM mode, `--mmap-experts`, which leaves the experts to the OS file cache. On Windows that cache fills RAM and gets trimmed, and decode falls off. On the NVFP4 fork, copying blobs through the mapped file made Windows trim and decode fell from 72 to 33 tok/s. `TieredExpertSource` keeps host copies only where the engine reads them: - **Pinned tier.** A pinned, registered tier (large pages where allowed) holds as many experts as fit in the free RAM minus `STRATA_RAM_RESERVE_GIB` (default 6). They go in by rank: first the experts outside VRAM, then the VRAM cache's tail that the prompt path borrows. The GPU copies them by DMA like the full arena's. - **Everything else** is read from `experts.bin` unbuffered, so the file cache neither grows nor gets trimmed. A decode token's misses are prefetched per layer (`begin_layer`), the prompt path's stager reads them with `read_blob`, and lent slots are refilled from the file. - **The adaptive tier stays in step.** An expert leaving VRAM without a host copy gets a tier slot (copied back from its VRAM slot first); one entering VRAM gives its slot back. `--low-ram`, `--no-low-ram` and `--ram-budget GIB` control it. It also switches on by itself when the pack has `experts.bin` (setup writes it only for its low-RAM mode) and the experts plus 16 GiB exceed the RAM installed. It needs one GPU, an `--expert-profile`, and no remote experts. setup's low-RAM mode passes `--low-ram` on Windows instead of `--mmap-experts`. The loader is Windows-only for now, so Linux keeps `--mmap-experts`. ## Measured Measured on current main (0.1.29, d6708a4) with ISTA-DASLab's IQ2_XS, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB DDR5, Windows 11. KL is first-token KL divergence (`STRATA_DUMP_FIRST_LOGITS`, #276); two runs of main itself differ by 0 on 5 of 6 short prompts and by 0.00013 on the sixth. - **Correct to the token.** With the same expert cache (17,806 slots), no PCIe share and a static cache, the full arena and the tier (`--ram-budget 8`) generate the same 128 tokens. `STRATA_TIER_VERIFY` compared 5,460 host copies with the file, and later 5,604 including the ones copied back from VRAM by the adaptive tier: 0 differ. - **An 8 GiB tier on IQ2_XS** (1,338 of the 6,800 experts outside VRAM come from the file): 137.7 tok/s decode vs 149.6 with the full arena (static cache, no PCIe share, 128 tokens). - **The NVFP4 fork** (63 GiB of experts, RAM emulated by locking 60 of 128 GB): | 64 GB of RAM with | decode | 32K prompt, first token | | --- | ---: | ---: | | RTX 5090 (32 GB) | 112-118 tok/s (same as with 96+ GB) | 9.3 s (96+ GB: 5.7 s) | | a 24 GB card | 86-90 tok/s (96+ GB: 95) | | | a 16 GB card | 54-56 tok/s (96+ GB: 67) | | With 96 GB or more nothing changes: first-token KL to the arena is exactly 0. ## Switch `--no-low-ram` gives the full arena; `--mmap-experts` still selects the mapped file. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
Sur le site
Liens install, modèles, releases.