Pull requests / #818

#818 Windows: opt-in release of mapped expert pages after GPU uploads

closed · @jordicor · 0 comentarios · En GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows

Descripción

On a 32 GB Windows 11 laptop with an RTX 4080 Laptop and a Thunderbolt RTX 3090, file-backed expert pages remained in the engine's working set after their GPU copies. With IQ2_XS and the existing layer split, available RAM stayed near 0.5 GiB during the measured workload.

This adds **opt-in Windows mapped-page release** via `STRATA_ARENA_RELEASE=1`. Unset or `0` preserves the current behavior.

- Implement `FileExpertSource::release()` for `experts.bin` and direct GGUF file views using Windows working-set trimming. Preserve shared boundary pages and never trim private/pinned staging buffers.
- Release pages after prefill's lent expert slots have finished refilling, in both serving and one-shot generation. Existing completed-upload release calls also use the implementation.
- Add focused C++ coverage, document the switch and tradeoff, and include the benchmark scripts and scoped measurements.

### Measured result

The same prototype executable was run with explicit release OFF and ON: one warmup and three fresh requests each at 1K, 4K and 16K prompt tokens, 256 output tokens, context 32768, int8 KV, MTP, `--mmap-experts --layer-split 12 --trim-stage-weights`.

| Metric | OFF | ON |
| --- | ---: | ---: |
| Available physical RAM, median | 0.512 GiB | 12.187 GiB |
| Available physical RAM, minimum | 0.277 GiB | 6.110 GiB |
| Engine working set, median | 23.379 GiB | 11.946 GiB |
| System commit, median | 51.880 GiB | 51.973 GiB |
| Physical SSD reads, total | 16.750 GiB | 16.809 GiB |

Prefill throughput was **3.7-4.4% lower**. This improves physical-memory availability on this machine; it does not reduce commit, demonstrate fewer SSD reads, or establish a general speed improvement. Runs were sequential, the OS file cache was retained, and generated code-explanation responses differed. The submitted patch defaults to OFF; the measured prototype defaulted to ON, but both benchmark arms set the switch explicitly.

Full configuration, individual timings, hashes, reproduction commands and limitations: [measurement report](https://github.com/jordicor/Strata/blob/codex/windows-mapped-expert-release/bench/results/2026-10-04-windows-mapped-release/README.md).

### Validation

- Fresh Windows Release build with MSVC 2019 and CUDA 13.1, SM 86/89, native experts enabled.
- `file_expert_source_test` and `expert_layout_test` pass in separate processes with the switch unset, `0`, and `1`: page alignment and working-set behavior, unchanged bytes after repeated reads, invalid indices, closed sources, and three GGUF roles across two shards.
- Eight benchmark-harness tests pass.
- Final rebuilt engine completes fresh 1K/4K inference requests on both GPUs with release enabled.

The prototype also completed 30 functional requests without runtime errors. All 24 objective responses/statuses matched the official engine at the same placement: 21 passes and three unchanged Markdown-fenced-JSON failures; six prose responses were unscored. These limited checks do not prove general output parity.

Related prior work: #467 and #640. The existing multi-GPU implementation is reused; this PR contributes Windows mapped-page release.

En el sitio

Enlaces a install, modelos, releases.