Pull requests / #42
#42 pinned: MEM_LARGE_PAGES on Windows never worked - enable the privilege, round the size
closed · merged 2026-09-28 · @pipeob0 · 0 comentários · No GitHub
NVIDIA / CUDAModels & quantsWindows
Descrição
### Why The pinned expert arena asks for `MEM_LARGE_PAGES` but on Windows it could never get them, for two independent reasons, and the log blamed a privilege the user had often already granted. **1. The privilege is never enabled in the process token.** Having *Lock pages in memory* assigned to the account is not enough — `VirtualAlloc` returns `ERROR_PRIVILEGE_NOT_HELD` (1314) unless the process calls `AdjustTokenPrivileges` for itself first. The code only had a comment about the privilege. This is what made the message misleading: granting the privilege, rebooting, and starting the engine printed the same "needs SeLockMemoryPrivilege" line, because a *different* reason was now the blocker. **2. The size was passed unrounded.** `VirtualAlloc` with `MEM_LARGE_PAGES` requires a size that is an exact multiple of `GetLargePageMinimum` (2 MiB). The arena size is not, so the call failed with `ERROR_INVALID_PARAMETER` (87) — a parameter error that reads exactly like the privilege problem in the log. Rounded up: the slack is under 2 MiB and the tail is never touched. Both fixed, on a desktop account with the privilege granted: ``` expert arena: cudaHostRegister PORTABLE ok; large pages (2097152 B) ``` ### The refusal note now distinguishes three different failures `1450` (`ERROR_NO_SYSTEM_RESOURCES`: the large-page pool does not have that many 2 MiB pages free *right now*), `87` (size not a multiple) and `1314` (the privilege really is off) all printed one identical line. Seen for real while A/B-ing this: a run started minutes after an engine process was killed printed 1450, and the next run succeeded — the killed process had not released its pinned pages yet. Note that this is a *state* symptom (a dying engine still holding pinned memory) reported as a privilege problem, which is the kind of thing worth being able to tell apart when reading a stall report (#31). The second commit adds the requested byte count and stops asserting the cause. ### Testing - `STRATA_NO_LARGEPAGES=1` skips the attempt entirely, so the same boot can be A/B'd. - End-to-end on a 5700X3D + RTX 5060 Ti, native `swift-iq3_xxs` pack, medians of interleaved runs: **neutral** here (48.8 ms/round without, 49.7 with). Reported as such: the reason to land this is that the arena ends up being what the code says it is, and that the log tells the truth. Boxes with a much bigger arena, or a heavier expert pool, have more TLB pressure to save than this one. No ABI or behaviour change for anyone who cannot get large pages: same 4 KiB fallback as before.
No site
Links install, modelos, releases.