Pull requests / #799
#799 docs: the expert cache size is a budget, and on Windows an over-sized one pages instead of failing
closed · @1314521gjy · 0 评论 · 在 GitHub 查看
BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows
描述
Documentation only, no engine change. `--expert-cache N` is `N` times the largest expert blob, and the engine compares it only with the free VRAM it reads **before** the slots are written. `--expert-cache auto` checks again **after** they are written and shrinks if they do not fit; that second check is auto-only. On Windows an allocation is not resident until it is touched, the free figure read before can be about a gigabyte too high, and the driver's default sysmem fallback moves an over-committed allocation to system memory instead of failing - so the start looks normal and only the speed shows it. Measured on an RTX 4080 SUPER 32 GB, IQ3_S, 0.1.39, `--max-context 1048576` (yarn factor 4), fresh 43,969 / 45,670 / 47,956-token prompts, 200 generated each, one engine start per row: | expert cache | slots | VRAM free at start | decode tok/s | | --- | ---: | --- | --- | | `--expert-cache 8900` | 11,631 | 0 MiB (`LOW`) | 13.7 / 14.7 | | `auto` | 11,178 | 217 MiB (`LOW`) | 102.1 / 122.3 | | `auto` + `--vram-reserve-mib 1500` | 10,766 | 1,075 MiB | 88.9 / 96.5 / 101.4 | | `--expert-cache 7000` | 9,148 | 4,130 MiB | 94.1 / 97.3 | | the same 11,631-slot cache at `--max-context 524288` | 11,631 | 299 MiB | 113.2 / 107.9 | 453 slots (0.86 GiB) separate 13.7 from 102 tok/s, and the slower run had the higher cache hit rate, so it is not misses. The page says what to do (leave it on `auto`, size it smaller, or raise `--vram-reserve-mib`, which is deducted before the cache is sized). The mechanism was suggested by the maintainer in #781; the measurements and the write-up are ours, including a **negative** result: the `\GPU Process Memory(*)\Shared Usage` / `Dedicated Usage` counters did **not** separate the slow run from the fast ones on that machine (they track the pinned expert arena and pinned K/V), which the page records rather than hides. The full report is PR #780; the same numbers with the profiler tables are in `bench/results/2026-10-04-community-rtx-4080s-iq3s/README.md`.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。