Pull requests / #818
#818 Windows: opt-in release of mapped expert pages after GPU uploads
closed · @jordicor · 0 评论 · 在 GitHub 查看
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
描述
On a 32 GB Windows 11 laptop with an RTX 4080 Laptop and a Thunderbolt RTX 3090, file-backed expert pages remained in the engine's working set after their GPU copies. With IQ2_XS and the existing layer split, available RAM stayed near 0.5 GiB during the measured workload. This adds **opt-in Windows mapped-page release** via `STRATA_ARENA_RELEASE=1`. Unset or `0` preserves the current behavior. - Implement `FileExpertSource::release()` for `experts.bin` and direct GGUF file views using Windows working-set trimming. Preserve shared boundary pages and never trim private/pinned staging buffers. - Release pages after prefill's lent expert slots have finished refilling, in both serving and one-shot generation. Existing completed-upload release calls also use the implementation. - Add focused C++ coverage, document the switch and tradeoff, and include the benchmark scripts and scoped measurements. ### Measured result The same prototype executable was run with explicit release OFF and ON: one warmup and three fresh requests each at 1K, 4K and 16K prompt tokens, 256 output tokens, context 32768, int8 KV, MTP, `--mmap-experts --layer-split 12 --trim-stage-weights`. | Metric | OFF | ON | | --- | ---: | ---: | | Available physical RAM, median | 0.512 GiB | 12.187 GiB | | Available physical RAM, minimum | 0.277 GiB | 6.110 GiB | | Engine working set, median | 23.379 GiB | 11.946 GiB | | System commit, median | 51.880 GiB | 51.973 GiB | | Physical SSD reads, total | 16.750 GiB | 16.809 GiB | Prefill throughput was **3.7-4.4% lower**. This improves physical-memory availability on this machine; it does not reduce commit, demonstrate fewer SSD reads, or establish a general speed improvement. Runs were sequential, the OS file cache was retained, and generated code-explanation responses differed. The submitted patch defaults to OFF; the measured prototype defaulted to ON, but both benchmark arms set the switch explicitly. Full configuration, individual timings, hashes, reproduction commands and limitations: [measurement report](https://github.com/jordicor/Strata/blob/codex/windows-mapped-expert-release/bench/results/2026-10-04-windows-mapped-release/README.md). ### Validation - Fresh Windows Release build with MSVC 2019 and CUDA 13.1, SM 86/89, native experts enabled. - `file_expert_source_test` and `expert_layout_test` pass in separate processes with the switch unset, `0`, and `1`: page alignment and working-set behavior, unchanged bytes after repeated reads, invalid indices, closed sources, and three GGUF roles across two shards. - Eight benchmark-harness tests pass. - Final rebuilt engine completes fresh 1K/4K inference requests on both GPUs with release enabled. The prototype also completed 30 functional requests without runtime errors. All 24 objective responses/statuses matched the official engine at the same placement: 21 passes and three unchanged Markdown-fenced-JSON failures; six prose responses were unscored. These limited checks do not prove general output parity. Related prior work: #467 and #640. The existing multi-GPU implementation is reused; this PR contributes Windows mapped-page release.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。