Pull requests / #153

#153 Expert cache keeps ~4 GiB more on WDDM (+10% decode on a 32 GB card); Ctrl+C and client hang-ups on Windows

closed · @borexola · 0 commentaires · Sur GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsWindows

Description

## Changes
- **Expert cache:** when VRAM reads 0 MiB free after the cache is written, give back 1 GiB per retry instead of a quarter of the cache. On a 32 GB 5090 card the quarter left ~4 GiB unused.
- **Ctrl+C:** now stops the server on Windows. A second Ctrl+C kills the engine. The run script pauses only on errors.
- **Client hang-ups:** a cancelled request no longer prints a stack trace.

## Results
RTX 5090, IQ3_S, 256K context, same prompt:

|                  | main        | this PR     |
|------------------|-------------|-------------|
| Experts on GPU   | 9,343       | 11,843      |
| Cache hit rate   | 95.2%       | 97.5%       |
| Decode speed     | 142.5 tok/s | 156.7 tok/s |

## Tested
- Serve tests: same results as main.
- Ctrl+C: main keeps running; this PR stops in 0.6 s.

Sur le site

Liens install, modèles, releases.