Pull requests / #286

#286 Low RAM on Windows: a pinned tier of the most-read experts, the rest read unbuffered from experts.bin

closed · @sergqwer · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installNVIDIA / CUDAWindowsLinux

描述

## Summary

Stacked on #285 (it uses its deferred per-layer registration). Only the last commit is new here.

The resident arena holds every expert, including the VRAM cache's share of them a second time. A PC whose RAM cannot
hold them beside the system has setup's low-RAM mode, `--mmap-experts`, which leaves the experts to the OS file
cache. On Windows that cache fills RAM and gets trimmed, and decode falls off. On the NVFP4 fork, copying blobs
through the mapped file made Windows trim and decode fell from 72 to 33 tok/s.

`TieredExpertSource` keeps host copies only where the engine reads them:

- **Pinned tier.** A pinned, registered tier (large pages where allowed) holds as many experts as fit in the free RAM
  minus `STRATA_RAM_RESERVE_GIB` (default 6). They go in by rank: first the experts outside VRAM, then the VRAM
  cache's tail that the prompt path borrows. The GPU copies them by DMA like the full arena's.
- **Everything else** is read from `experts.bin` unbuffered, so the file cache neither grows nor gets trimmed. A
  decode token's misses are prefetched per layer (`begin_layer`), the prompt path's stager reads them with
  `read_blob`, and lent slots are refilled from the file.
- **The adaptive tier stays in step.** An expert leaving VRAM without a host copy gets a tier slot (copied back from
  its VRAM slot first); one entering VRAM gives its slot back.

`--low-ram`, `--no-low-ram` and `--ram-budget GIB` control it. It also switches on by itself when the pack has
`experts.bin` (setup writes it only for its low-RAM mode) and the experts plus 16 GiB exceed the RAM installed. It
needs one GPU, an `--expert-profile`, and no remote experts. setup's low-RAM mode passes `--low-ram` on Windows
instead of `--mmap-experts`. The loader is Windows-only for now, so Linux keeps `--mmap-experts`.

## Measured

Measured on current main (0.1.29, d6708a4) with ISTA-DASLab's IQ2_XS, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB DDR5, Windows 11. KL is first-token KL divergence (`STRATA_DUMP_FIRST_LOGITS`, #276); two runs of main itself differ by 0 on 5 of 6 short prompts and by 0.00013 on the sixth.

- **Correct to the token.** With the same expert cache (17,806 slots), no PCIe share and a static cache, the full
  arena and the tier (`--ram-budget 8`) generate the same 128 tokens. `STRATA_TIER_VERIFY` compared 5,460 host
  copies with the file, and later 5,604 including the ones copied back from VRAM by the adaptive tier: 0 differ.
- **An 8 GiB tier on IQ2_XS** (1,338 of the 6,800 experts outside VRAM come from the file): 137.7 tok/s decode vs
  149.6 with the full arena (static cache, no PCIe share, 128 tokens).
- **The NVFP4 fork** (63 GiB of experts, RAM emulated by locking 60 of 128 GB):

| 64 GB of RAM with | decode | 32K prompt, first token |
| --- | ---: | ---: |
| RTX 5090 (32 GB) | 112-118 tok/s (same as with 96+ GB) | 9.3 s (96+ GB: 5.7 s) |
| a 24 GB card | 86-90 tok/s (96+ GB: 95) | |
| a 16 GB card | 54-56 tok/s (96+ GB: 67) | |

With 96 GB or more nothing changes: first-token KL to the arena is exactly 0.

## Switch

`--no-low-ram` gives the full arena; `--mmap-experts` still selects the mapped file.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。