Pull requests / #640
#640 STRATA_ARENA_MMAP=1: the expert arena as a read-only mapped file (Linux), for small-RAM machines
closed · @JeanP00l · 0 comentários · No GitHub
Multi-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
Descrição
## What `STRATA_ARENA_MMAP=1` (opt-in, Linux; nothing changes without it) maps the host expert arena read-only from the pack's `experts.bin` instead of reading it into pinned, locked RAM. This is for machines whose GPUs already hold most experts and whose RAM is small. `--mmap-experts` refuses native packs (the Coder), so this path is for them. Without it, the Coder (25 GB arena) on 2x 16 GB GPUs with 32 GB of RAM left ~1-3 GB available at 128K context. In long sessions the kernel swapped, and re-reading a multi-turn checkpoint took 25-80 s. - **First start:** writes the arena to `experts.bin` (padded by one blob, so a whole-slot copy may start at any expert). Later starts map it. - **Page cache:** keeps what the CPU pool and the prompt path read, and gives the rest back under pressure. A dropped page is re-read from the file, so no result depends on what is resident. The GPUs get no mapped alias: run with `--pcie-frac 0`. - **After the VRAM caches fill:** pages of experts a GPU holds are released (`MADV_DONTNEED`; `STRATA_ARENA_RELEASE=0` keeps them), and the rest are prefetched (`MADV_WILLNEED`). An adaptive swap releases the expert that went to VRAM and prefetches the one that came back. - **HIP:** ROCclr copies a pageable source of >= 1 MiB by locking its pages in place, and keeps them locked. That undid every release. The engine raises `GPU_PINNED_MIN_XFER_SIZE` (an explicit setting wins). It locks an adaptive swap's source pages itself, only until the copy lands. The ranges are merged first, because a copy from a range only partly registered fails with `invalid argument`. - **`STRATA_RSS_TRACE=1`:** prints `RssFile` at the startup steps. `madvise` reports success on locked pages and frees nothing, so `RssFile` is the honest check. - **Other platforms:** Windows compiles it out. The upstream HIP compat gets `cudaHostRegisterReadOnly`. ## Measured 2x AMD Instinct MI50 16 GB + 32 GB DDR4 (gfx906, see #638), Coder IQ1_M, 128K context, layer split: | | pinned arena | `STRATA_ARENA_MMAP=1` | |---|---|---| | RAM available while serving | ~1 GB | 25 GB | | decode, ms per verify window | 49.16 | 49.06 | ## Status Re-run on 2x MI50 with this branch on current `main`, results in the comment below. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
No site
Links install, modelos, releases.