Pull requests / #1520
#1520 prefill: release copied Windows expert pages after host staging
open · @W1nge · 0 comments · View on GitHub
BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Description
## Summary With `STRATA_FILE_RELEASE=1` on Windows, GPU uploads/refills release their expert file pages, but prefill's host staging copies leave them in the working set. Long prompts can consequently exhaust available RAM even when the CPU expert complement already has its own resident copy. ## What changed Keep the source and expert ID in stable-pointer staging jobs too, and call the existing `release()` after `memcpy` / `copy_blob()` completes, before publishing the staging buffer as ready. DMA then reads the independent staging buffer. This covers both the full expert stream and routed-only staging. The existing Windows opt-in and full-interior-page handling remain in charge; pinned buffers, resident copies, shared boundary pages, and default settings are unchanged. No new tuning option or GPU-model check. ## Extra Notes Validation on Windows, i7-13850HX / 32 GB RAM, RTX 2080 Ti 22 GiB + P100 16 GiB, driver 537.13, Qwen3.8-Flash-Next IQ3_XXS: - Isolated 0.1.40.3 comparison, two starts per arm with release enabled in both: median available RAM **0.435–0.476 → 11.613–11.820 GiB** during the measured long-prompt interval. Each start ran 18 mixed requests, including 1.9K/5.1K prompts. Both candidate runs' 36 responses matched the control byte for byte. A separate 12-response coding comparison also matched exactly. - The reversed timing comparison was essentially flat, so this is a memory-pressure improvement, not a stable speedup claim. Measured system-level physical SSD reads increased **54.1–61.6 → 63.9–64.4 GiB**; trimming can lose warm-cache benefits. That tradeoff is why the existing opt-in remains appropriate. - MSVC / CUDA 12.4.131 build passed for sm_60 and sm_75. Existing `file_expert_source_test --rotation-gpu` passed with release off and on. Final integrated 0.1.40.4 validation passed 12 coding, 18 mixed, and 4 routed-staging responses. The last four also matched the earlier 0.1.40.3 outputs. - Both GPUs passed `mmvq_multi_parity` and real-model `native_expert_parity`. The extra `pdl_parity` diagnostic stopped at `cudaGraphGetEdges_v2` on the retained driver; an API probe records legacy graph-query success and v2 `invalid argument` (driver interface 12020 / runtime 12040). It is not counted as a passing test. No HIP/Linux hardware validation. Only these three source/documentation files are in this PR (+18/−12). Machine configuration, benchmark archives, and the upstream Pascal change are excluded. Full measurements, failed diagnostic, output changes after the separate upstream update, and reproduction records: [report and evidence](https://github.com/W1nge/Strata/blob/ea07f9d16506b96cce26d4eb444a208048e4ba6b/docs/performance/PART2-PREFILL-STAGING-2026-10-08.md).
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.