Pull requests / #1354

#1354 docs: the expert cache is a budget, and on Windows an over-sized one pages instead of failing

open · @1314521gjy · 0 comments · View on GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows

Description

Rebuilt on `82f46a8`: the earlier PR with this change (#799) was closed automatically when `main` was force-pushed during the history cleanup, so this is the same content cherry-picked onto the new main, plus the issue #781 data that arrived afterwards.

**What it documents.** `--expert-cache N` is compared only with the free VRAM read *before* the slots are written; `auto` checks again once they exist and shrinks, and that second check is auto-only. Under WDDM an over-committed allocation is not resident until it is touched and is put in system memory instead of failing, so the engine starts normally and only the speed shows it: 0 MiB free at 1,048,576 ran at 13.7 tok/s where 217 MiB ran at 102, and 453 slots (0.86 GiB) separate them, with the slower run showing the *higher* cache hit rate (so it is not misses). The startup line is the tell, and the two ways out are `auto` and a larger `--vram-reserve-mib` (deducted before the cache is sized).

**Added since the closed PR:** the 16 GB RTX 5080 reproduction from #781 (enkynakamura) - reserve 770 -> 0 grew the cache from 2,782 to 3,115 slots while decode fell from 40.9 to 16.6 tok/s with the hit rate flat - and the 2x RTX 5090 control from #781 (gucasbrg), where the same reserve change at healthy margins left decode unchanged (157.8 against 157.7 tok/s at 1,048,576). Together they separate the two regimes: the reserve is a safety knob, not a speed knob, unless the card is already at 0 MiB free.

Docs only, no engine changes, no new flags. The section points at issue #781 for the measurements, and the RTX 4080 SUPER configuration behind the numbers is in `bench/results/2026-10-04-community-rtx-4080s-iq3s` (PR #1351).

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.