Pull requests / #153
#153 Expert cache keeps ~4 GiB more on WDDM (+10% decode on a 32 GB card); Ctrl+C and client hang-ups on Windows
closed · @borexola · 0 comentários · No GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsWindows
Descrição
## Changes - **Expert cache:** when VRAM reads 0 MiB free after the cache is written, give back 1 GiB per retry instead of a quarter of the cache. On a 32 GB 5090 card the quarter left ~4 GiB unused. - **Ctrl+C:** now stops the server on Windows. A second Ctrl+C kills the engine. The run script pauses only on errors. - **Client hang-ups:** a cancelled request no longer prints a stack trace. ## Results RTX 5090, IQ3_S, 256K context, same prompt: | | main | this PR | |------------------|-------------|-------------| | Experts on GPU | 9,343 | 11,843 | | Cache hit rate | 95.2% | 97.5% | | Decode speed | 142.5 tok/s | 156.7 tok/s | ## Tested - Serve tests: same results as main. - Ctrl+C: main keeps running; this PR stops in 0.6 s.
No site
Links install, modelos, releases.