Pull requests / #42

#42 pinned: MEM_LARGE_PAGES on Windows never worked - enable the privilege, round the size

closed · merged 2026-09-28 · @pipeob0 · 0 commentaires · Sur GitHub

NVIDIA / CUDAModels & quantsWindows

Description

### Why

The pinned expert arena asks for `MEM_LARGE_PAGES` but on Windows it could never get them, for two
independent reasons, and the log blamed a privilege the user had often already granted.

**1. The privilege is never enabled in the process token.** Having *Lock pages in memory* assigned to the
account is not enough — `VirtualAlloc` returns `ERROR_PRIVILEGE_NOT_HELD` (1314) unless the process calls
`AdjustTokenPrivileges` for itself first. The code only had a comment about the privilege. This is what made
the message misleading: granting the privilege, rebooting, and starting the engine printed the same
"needs SeLockMemoryPrivilege" line, because a *different* reason was now the blocker.

**2. The size was passed unrounded.** `VirtualAlloc` with `MEM_LARGE_PAGES` requires a size that is an exact
multiple of `GetLargePageMinimum` (2 MiB). The arena size is not, so the call failed with
`ERROR_INVALID_PARAMETER` (87) — a parameter error that reads exactly like the privilege problem in the log.
Rounded up: the slack is under 2 MiB and the tail is never touched.

Both fixed, on a desktop account with the privilege granted:

```
expert arena: cudaHostRegister PORTABLE ok; large pages (2097152 B)
```

### The refusal note now distinguishes three different failures

`1450` (`ERROR_NO_SYSTEM_RESOURCES`: the large-page pool does not have that many 2 MiB pages free *right
now*), `87` (size not a multiple) and `1314` (the privilege really is off) all printed one identical line.
Seen for real while A/B-ing this: a run started minutes after an engine process was killed printed 1450, and
the next run succeeded — the killed process had not released its pinned pages yet. Note that this is a
*state* symptom (a dying engine still holding pinned memory) reported as a privilege problem, which is the
kind of thing worth being able to tell apart when reading a stall report (#31).

The second commit adds the requested byte count and stops asserting the cause.

### Testing

- `STRATA_NO_LARGEPAGES=1` skips the attempt entirely, so the same boot can be A/B'd.
- End-to-end on a 5700X3D + RTX 5060 Ti, native `swift-iq3_xxs` pack, medians of interleaved runs:
  **neutral** here (48.8 ms/round without, 49.7 with). Reported as such: the reason to land this is that the
  arena ends up being what the code says it is, and that the log tells the truth. Boxes with a much bigger
  arena, or a heavier expert pool, have more TLB pressure to save than this one.

No ABI or behaviour change for anyone who cannot get large pages: same 4 KiB fallback as before.

Sur le site

Liens install, modèles, releases.