Pull requests / #380

#380 HIP: count the desktop's VRAM on Windows (WDDM budget, STRATA_WDDM_BUDGET)

closed · @BlueKingMuch · 0 コメント · GitHub で見る

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

本文

On Windows, `hipMemGetInfo` reports the card's size minus this process's own allocations: ROCclr's PAL backend asks Windows for this process's usage only, so what the desktop and other programs hold on the card is not subtracted. On a card that also drives the desktop the engine counts the ~930 MiB those hold as free, `--expert-cache auto` fills the card past what Windows lets the process keep in VRAM, and WDDM moves part of the engine out to system RAM. On an RX 6800 with setup's configuration this PR takes decode from 30.5 to 41.4 tok/s (medians of 24 requests each).

## What changes

- Windows gives every process a video memory budget that does account for the others (DXGI `QueryVideoMemoryInfo`). On Windows HIP builds, `cudaMemGetInfo`, which the HIP shim maps to `hipMemGetInfo`, now lowers `hipMemGetInfo`'s free figure by what the budget withholds: the card's size minus the budget (799 MiB on this card while the process is small, about 1.8 GiB once it holds 11 GiB).
- Everything sized from that figure stays within the budget. For `--expert-cache auto` the existing check after the slots are written does the rest ("only 8 MiB free once the slots are written ... shrinking"), since the budget itself drops once the process holds most of the card (probe below).
- Lowering HIP's own figure, rather than taking the budget minus DXGI's usage, keeps memory the HIP runtime has freed but still caches counted as free, the way HIP counts it (in the probe below DXGI still counts 1946 MiB as used after everything was freed).
- `dxgi.dll` is loaded at run time and the card's adapter is found by its LUID, so nothing new is linked. Without `dxgi.dll`, the adapter, or a budget below the card's size, `hipMemGetInfo`'s figure stands.
- The first query logs it once: `strata: Windows budgets 15569 of this card's 16368 MiB for this process; free VRAM is counted within that (STRATA_WDDM_BUDGET=0: off)`. `STRATA_WDDM_BUDGET=0` turns it off.
- CUDA and Linux builds are unchanged: the code is under `#if defined(STRATA_USE_HIP) && defined(_WIN32)`, and the shim header is only used by HIP builds.

It is on by default, since without it setup's own configuration overfills a card that drives a display; if you prefer it opt-in, that is the `off` lambda in `mem_get_info`.

No kernel changes. What changes is the free-VRAM figure and with it how many experts sit in VRAM, as with any other `--expert-cache` size.

## Measured

RX 6800 16 GB (gfx1030, drives the desktop), Ryzen 7 5800X3D, PCIe 4.0 x16, Windows 11 (build 26300), HIP SDK 7.2, Visual Studio 2026 Build Tools. A local test build of 0.1.31 (9259cad) with #356, #311 (rebased onto 0.1.31), this commit, #373, #377, an earlier version of #376 (the same Release flags), and a scratch reserve for verify windows (branch `win-scratch` on my fork; not proposed, since 0.1.31 did not hang without it). Every comparison in the table is within that build: `STRATA_WDDM_BUDGET=0` against unset.

Qwen3.8-Flash-Next IQ3_XXS with the configuration setup wrote: `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 131072 --kv int8 --kv-resident 32768`, MTP draft head, `pcie_frac` 0.55 from the probe. Two engine starts per variant, each with four passes over `tools/calibrate.py`'s three prompts, 128 tokens each; the card's memory from Windows' GPU counters after loading.

| | budget off (`STRATA_WDDM_BUDGET=0`) | budget on (default) |
|---|---|---|
| free VRAM the cache sizing starts from | 10.69 GiB | 9.91 GiB |
| expert cache after the shrink check | 9.83 GiB (6072 slots) | 8.05 GiB (4960 slots) |
| the card in use, all programs | 16132 of 16368 MiB | 15230 MiB |
| the engine's dedicated VRAM | 15198 MiB | 14296 MiB |
| its shared GPU memory, above the budget-on start's | +929 MiB | (reference) |
| decode, median of 24 requests | 30.5 tok/s | 41.4 tok/s (+36 %) |
| decode per prompt (median of 8 each) | 30.7 / 29.4 / 30.0 | 47.1 / 40.1 / 41.4 |
| decode, slowest / fastest request | 26.8 / 31.9 | 37.8 / 47.7 |

With the budget off the expert cache is 1.78 GiB larger; 902 MiB more of the engine is in VRAM (dedicated) and 929 MiB more in system memory (shared). The two starts of each variant gave the same sizes and memory figures, and their decode medians are within 0.4 tok/s of each other. The prompt path borrows the same 4.44 GiB of cache slots in both and the prompt chunk stays 8192 tokens; long prompts were not measured.

`--vram-reserve-mib` alone does not help: on a 0.1.30 build, `--vram-reserve-mib 1500` still left the card at 16037 MiB, since the reserve is taken from the same figure.

The budget on this card, from a standalone probe that allocates in steps of 1 GiB (256 MiB past 14 GiB), in MiB:

| held by the probe | 0 | 10240 | 11264 | 14336 | 16128 | all freed |
|---|---|---|---|---|---|---|
| `hipMemGetInfo` free | 16227 | 5974 | 4950 | 1878 | 86 | 16214 |
| DXGI budget | 15569 | 15569 | 14515 | 14543 | 14538 | 15569 |
| DXGI usage | 141 | 10394 | 11418 | 14490 | 16282 | 1946 |

The allocations past the budget all succeed; whether and how much of that Windows then moves out depends on what else wants the card (in the engine runs above, 929 MiB).

To reproduce: `STRATA_WDDM_BUDGET=0` against unset, comparing the engine's dedicated and shared GPU memory between the two (Task Manager's details view or the `GPU Process Memory` counters). Shared memory is large in both (the registered expert arena counts there); the difference is what moved.

Tested on this card only, with the desktop on it. On a card without a display the budget should sit closer to the card's size, so less would change there; not tested, nor with more than one GPU.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_015Ld5YDzuKW1hgMVwwgRZDP

関連リンク

インストール・モデル・リリースへの站内リンク。