Pull requests / #799

#799 docs: the expert cache size is a budget, and on Windows an over-sized one pages instead of failing

closed · @1314521gjy · 0 Kommentare · Auf GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows

Beschreibung

Documentation only, no engine change.

`--expert-cache N` is `N` times the largest expert blob, and the engine compares it only with the free VRAM it reads **before** the slots are written. `--expert-cache auto` checks again **after** they are written and shrinks if they do not fit; that second check is auto-only. On Windows an allocation is not resident until it is touched, the free figure read before can be about a gigabyte too high, and the driver's default sysmem fallback moves an over-committed allocation to system memory instead of failing - so the start looks normal and only the speed shows it.

Measured on an RTX 4080 SUPER 32 GB, IQ3_S, 0.1.39, `--max-context 1048576` (yarn factor 4), fresh 43,969 / 45,670 / 47,956-token prompts, 200 generated each, one engine start per row:

| expert cache | slots | VRAM free at start | decode tok/s |
| --- | ---: | --- | --- |
| `--expert-cache 8900` | 11,631 | 0 MiB (`LOW`) | 13.7 / 14.7 |
| `auto` | 11,178 | 217 MiB (`LOW`) | 102.1 / 122.3 |
| `auto` + `--vram-reserve-mib 1500` | 10,766 | 1,075 MiB | 88.9 / 96.5 / 101.4 |
| `--expert-cache 7000` | 9,148 | 4,130 MiB | 94.1 / 97.3 |
| the same 11,631-slot cache at `--max-context 524288` | 11,631 | 299 MiB | 113.2 / 107.9 |

453 slots (0.86 GiB) separate 13.7 from 102 tok/s, and the slower run had the higher cache hit rate, so it is not misses. The page says what to do (leave it on `auto`, size it smaller, or raise `--vram-reserve-mib`, which is deducted before the cache is sized).

The mechanism was suggested by the maintainer in #781; the measurements and the write-up are ours, including a **negative** result: the `\GPU Process Memory(*)\Shared Usage` / `Dedicated Usage` counters did **not** separate the slow run from the fast ones on that machine (they track the pinned expert arena and pinned K/V), which the page records rather than hides.

The full report is PR #780; the same numbers with the profiler tables are in `bench/results/2026-10-04-community-rtx-4080s-iq3s/README.md`.

Mehr auf der Site

Links zu Install, Modellen, Releases.