Pull requests / #279

#279 Windows: leave 1500 MiB of VRAM unused by default (decode stalled with 700)

closed · @sergqwer · 0 comments · View on GitHub

BenchmarksSetup & installNVIDIA / CUDAWindows

Description

## Summary

On Windows the GPU usually also drives the desktop. With the default `--vram-reserve-mib 700`, only ~470-570 MiB
stayed free once the expert cache was written. Whenever the desktop's programs took more than that, WDDM stalled the
GPU. It did not happen in every run, which is what made it hard to see: decode speed simply differed from run to run.

The default is now 1500 MiB on Windows (700 elsewhere). setup's vision reserve no longer lowers it on Windows, and
DETAILS.md gets a troubleshooting row for older engines.

## Measured

Measured on current main (0.1.29, d6708a4) with ISTA-DASLab's IQ2_XS, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB DDR5, Windows 11. KL is first-token KL divergence (`STRATA_DUMP_FIRST_LOGITS`, #276); two runs of main itself differ by 0 on 5 of 6 short prompts and by 0.00013 on the sixth. 400-token answers, 4 runs each, interleaved:

| reserve | decode | verify window wait | free after load | expert slots |
| --- | ---: | ---: | ---: | ---: |
| 700 MiB | 112, 82, 77, 78 tok/s | 14.6, 22.0, 22.1, 21.6 ms/round | 467-573 MiB | ~17,850 |
| 1500 MiB | 139, 139, 132, 132 tok/s | 10.5, 10.5, 11.0, 11.4 ms/round | 1,259-1,375 MiB | ~17,300 |

An NVFP4 pack showed the same (700: 58-90 tok/s in half the runs, 104-116 in the others; 1500: 108-116 in every run).
The CPU pool, the draft acceptance and the expert cache are the same in slow and fast runs; only the GPU's own time
per round grows.

## Switch

`--vram-reserve-mib 700` brings the old default back.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.