Pull requests / #563

#563 Give the expert cache's VRAM back without unloading the model (#533)

closed · @Mirtraxxx · 0 Kommentare · Auf GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Beschreibung

Addresses #533 for the expert cache, the largest part of the VRAM: the engine gives the cache's VRAM back to other programs and takes it again, while the model, the context and the conversation cache stay loaded.

## What it does

- **Engine, `--expert-cache-release`** (off by default). The expert cache is allocated with CUDA's virtual memory API (`cuMemAddressReserve` + `cuMemCreate`/`cuMemMap`) instead of `cudaMalloc`. Two new serve-loop commands:
  - `RELEASE`: every resident expert is marked not resident (and remembered), then the physical memory is unmapped and freed. The address range stays reserved, so the captured graphs and the experts' pointers stay valid.
  - `REFILL`: new physical memory is mapped at the same address, and the remembered experts are copied back into the same slots from the host copy. The residency table is restored as it was.
  - A `GEN`/`GENI` that arrives while the cache is released refills it first, so a client needs nothing new.
- **Server:**
  - `POST /cache/release` and `POST /cache/refill` answer `409` while a request is running and `400` when the engine can't (e.g. started without the option).
  - `--cache-release-idle SECONDS` / `"cache_release_idle_s"` gives the cache back after that many seconds without requests. It adds the engine option itself, and it works like `--idle-unload` but keeps the model loaded.
- **Docs:** a paragraph and a table row in DETAILS.md's "Sharing the GPU with other programs".

## Measured

RTX 3090 24 GB, Windows 11, driver 616.56, a Qwen3.8-Flash-Next fine-tune at IQ2_XS, `--expert-cache auto`, MTP on:

| | |
| --- | --- |
| VRAM given back | 14.5 GiB (10,800 experts) in ~50 ms |
| refill | 1.7-2.5 s |
| byte check (`STRATA_VERIFY_REFILL=1` reads every refilled slot back) | 0 of 10,800 slots differed from their host copy, over 3 release/refill cycles |
| a ComfyUI video render while the model stays loaded | 176 s with the cache in VRAM, 128 s with it given back |
| decoding after a refill | ~72 tok/s (a 400-token answer) |

`serve/test_lifecycle.py` gets three tests: the endpoints, the idle wait, and the engine flag being added. `test_server`, `test_lifecycle`, `test_monitor` and `test_structured` pass (140 tests).

## Limits

- Only the expert cache is given back. The dense weights, the KV cache and the engine's buffers stay; on the 3090 the card still held about 7 GiB, the desktop included. A full `/unload` is still the way to free everything.
- CUDA only. On HIP, when the device has no VMM support, with a layer split, and in the resident RAM mode (some experts only in VRAM), the engine prints why and keeps the cache as before. **Not built or tested on HIP or Linux here**; the HIP paths are `#if`-guarded, and `CUDA::cuda_driver` is only linked in the CUDA build.
- Refill when the VRAM is actually free. On Windows, a refill into a card another program still fills makes WDDM move memory to shared system RAM, and decoding ran at ~10-17 tok/s in the same setup until the other program let go. The engine can't tell that this happened, so deciding when to refill stays with the caller (or with the next request).

The idea and the use case are AXKore's in #533 (a CAD/games workstation). Mine is an LLM sharing a 3090 with ComfyUI.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Mehr auf der Site

Links zu Install, Modellen, Releases.