Issues / #1412

#1412 Windows: `SeLockMemoryPrivilege` moves the expert arena to 2 MiB large pages — but the token is issued at logon (1314 ≠ 1450)

open · @shahrokhzargarpour · 1 comentários · No GitHub

BenchmarksNVIDIA / CUDAWindows

Descrição

## What changes

On Windows, if the account running the engine holds the **"Lock pages in memory"** user right
(`SeLockMemoryPrivilege`), Strata backs the expert arena with **2 MiB large pages** instead of 4 KB pages.
Two things change in the startup log:

1. The `large pages refused …` warning disappears and becomes a positive line.
2. The `working-set minimum + VirtualLock` line for the (~46,700 MiB) arena disappears — large pages are
   non-pageable by construction, so the engine no longer needs to pin the working set to keep the arena resident.

This is an **observation of startup-log behavior** (allocator / working-set), not a throughput A/B.

## Environment (where this was observed)

- Engine: Strata **v0.1.40.1** (`82f46a8c`), Windows 11, WDDM driver model.
- GPUs: **RTX 4080 SUPER 16 GB + RTX 3090 24 GB**, layer split, expert weights in system RAM.
- Model: Qwen3.8-Flash-Next **UD-IQ4_XS** (MoE), `--expert-cache auto`.
- Flags: `--batch 4 --batch-mtp --spec 4`, KV int8 with `--kv-resident 98304`,
  `--remote-expert-opt --pcie-frac 0.55 --pool-workers 8`.

## How to enable it (step by step)

1. Grant **Local Policies → User Rights Assignment → "Lock pages in memory"** (`SeLockMemoryPrivilege`) to
   the account that runs the engine (via `secpol.msc` or your provisioning tool). It is a **user right**, not
   an "elevated process" privilege — the engine process does not need to run as administrator.
2. **Sign out and back in.** The privilege is baked into the access token **at logon**; restarting the engine
   (or restarting the scheduled Task) is **not enough**. This was the part that cost the most time.
3. Confirm the right is present in the account that will run the engine:
   ```
   whoami /priv
   ```
   `SeLockMemoryPrivilege` must be listed for that account.
4. Start the engine as usual and read the arena line.

## Evidence (verbatim startup lines)

Before (privilege not in the token) — the arena falls back to 4 KB pages:

```
large pages refused for 59523465216 B (GetLargePageMinimum=2097152, VirtualAlloc error 1314); using 4 KB pages
```

After (privilege present, fresh logon) — the arena is large-page backed:

```
expert arena: large pages are resident without a lock; … large pages (2097152 B)
```

and the ~46,700 MiB `working-set minimum + VirtualLock` line no longer appears.

## 1314 is not 1450

Two different failures show up as "large pages refused" and they need different fixes:

- **error 1314** = `ERROR_PRIVILEGE_NOT_HELD`: the account/token does not hold `SeLockMemoryPrivilege`.
  Fix: grant the right and **re-logon** (steps 1–2).
- **error 1450** = `ERROR_NO_SYSTEM_RESOURCES`: the privilege is fine, but the system cannot satisfy a
  contiguous large-page allocation — fragmentation. Observed here after killing and relaunching several
  engines in a row; a **clean reboot** made it succeed every time.

## Not measured

**The effect on tok/s / throughput was not measured.** Both lines above are startup-log observations, not a
controlled comparison. No A/B was run. Do not read this as a performance claim.

## Minimal repro

1. Run once without the right → capture `large pages refused … error 1314`.
2. Grant `SeLockMemoryPrivilege` to the account, **sign out / sign in**, verify `whoami /priv`.
3. Run again with the same args → capture `large pages are resident without a lock … large pages (2097152 B)`.

No site

Links install, modelos, releases.