Pull requests / #640

#640 STRATA_ARENA_MMAP=1: the expert arena as a read-only mapped file (Linux), for small-RAM machines

closed · @JeanP00l · 0 Kommentare · Auf GitHub

Multi-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Beschreibung

## What

`STRATA_ARENA_MMAP=1` (opt-in, Linux; nothing changes without it) maps the host expert arena read-only from the pack's `experts.bin` instead of reading it into pinned, locked RAM. This is for machines whose GPUs already hold most experts and whose RAM is small. `--mmap-experts` refuses native packs (the Coder), so this path is for them.

Without it, the Coder (25 GB arena) on 2x 16 GB GPUs with 32 GB of RAM left ~1-3 GB available at 128K context. In long sessions the kernel swapped, and re-reading a multi-turn checkpoint took 25-80 s.

- **First start:** writes the arena to `experts.bin` (padded by one blob, so a whole-slot copy may start at any expert). Later starts map it.
- **Page cache:** keeps what the CPU pool and the prompt path read, and gives the rest back under pressure. A dropped page is re-read from the file, so no result depends on what is resident. The GPUs get no mapped alias: run with `--pcie-frac 0`.
- **After the VRAM caches fill:** pages of experts a GPU holds are released (`MADV_DONTNEED`; `STRATA_ARENA_RELEASE=0` keeps them), and the rest are prefetched (`MADV_WILLNEED`). An adaptive swap releases the expert that went to VRAM and prefetches the one that came back.
- **HIP:** ROCclr copies a pageable source of >= 1 MiB by locking its pages in place, and keeps them locked. That undid every release. The engine raises `GPU_PINNED_MIN_XFER_SIZE` (an explicit setting wins). It locks an adaptive swap's source pages itself, only until the copy lands. The ranges are merged first, because a copy from a range only partly registered fails with `invalid argument`.
- **`STRATA_RSS_TRACE=1`:** prints `RssFile` at the startup steps. `madvise` reports success on locked pages and frees nothing, so `RssFile` is the honest check.
- **Other platforms:** Windows compiles it out. The upstream HIP compat gets `cudaHostRegisterReadOnly`.

## Measured

2x AMD Instinct MI50 16 GB + 32 GB DDR4 (gfx906, see #638), Coder IQ1_M, 128K context, layer split:

| | pinned arena | `STRATA_ARENA_MMAP=1` |
|---|---|---|
| RAM available while serving | ~1 GB | 25 GB |
| decode, ms per verify window | 49.16 | 49.06 |

## Status

Re-run on 2x MI50 with this branch on current `main`, results in the comment below.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Mehr auf der Site

Links zu Install, Modellen, Releases.