Pull requests / #1002
#1002 arena: STRATA_ARENA_MMAP on Windows (mapped experts.bin, streamed first fill)
closed · @mirifiuto135-debug · 0 评论 · 在 GitHub 查看
BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
描述
## What `STRATA_ARENA_MMAP=1` (the mapped expert arena from #640) is POSIX-only. This ports it to Windows, in `src/core/expert_source.cpp` only: - open: `experts.bin` mapped read-only with `CreateFileMapping` / `MapViewOfFile` - prefetch: `PrefetchVirtualMemory`; release: `VirtualUnlock` on the unlocked pages (drops them from the working set, they stay on the standby list) - close: `UnmapViewOfFile` - first start on Windows fills `experts.bin` through a **writable mapping** with `load_experts_gguf`, so the arena is never held in RAM (written to `.tmp`, then renamed). The POSIX first start is unchanged. One doc line in `docs/DETAILS.md` says "Linux and Windows" instead of "Linux". ## Why On a 128 GB Windows box the Q6 pack's arena (119.5 GiB) could not be served with a layer split: `--resident-budget-gib` is refused together with a split, and without the mapped arena the whole arena has to sit in RAM. With this change it maps from disk. ## Measured Windows 11, 3x RTX 5060 Ti 16 GB, 128 GB RAM, Qwen3.8 Flash-Next UD-Q6_K_XL converted pack, layer split auto, `--pcie-frac 0`, MTP on: - first fill of 119.5 GiB: 75 s - warm decode 20–24 tok/s, prefill 270–350 tok/s, flat from 8k to 199k prompt tokens - 262144 context loads; ~13 GB per card; 14–19 GiB of RAM free while serving - before this change on the same box: 17–18 tok/s on one card with the RAM budget, 3.7–6.2 tok/s on three cards ## Not in this PR Nothing else changes; AMD and Linux paths are untouched. I have not tested on a 32 GB RAM Windows machine. ## Disclosure The code was written with Claude Code (Anthropic) under my direction; I built and tested it on the hardware above. Happy to adjust anything.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。