Pull requests / #208

#208 server: optional idle unload, /unload and /load, a free-VRAM guard (share the GPU with other programs)

closed · @bytethecookie · 0 评论 · 在 GitHub 查看

Setup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

描述

## What

Optional support for sharing the GPU with other programs: Strata can unload the model when it is idle and load it again on the next request. Everything is **off by default** — without the new options the server behaves exactly as before.

| Option | Config key | What it does |
| --- | --- | --- |
| `--idle-unload SECONDS` | `idle_unload_s` | unload after that long without requests; the next request loads it again |
| `--min-free-vram-mib N` | `min_free_vram_mib` | load an unloaded model only when that much VRAM is free (waits up to 15 s), else `503` "the GPU is in use by another program" |
| `--before-load CMD` | `before_load` (string or list) | a command run before the model is loaded again, e.g. to unload another server's model |

Plus `POST /unload` (`409` while a request is running) and `POST /load`; `/health` gains `"loaded"`, `/v1/models` lists the model with status `unloaded` (like llama.cpp's router, building on #122's endpoints), `/props` sets `is_sleeping`, and `/metrics` reports the state `unloaded`.

## Why

Strata fills all free VRAM and holds it until the window is closed. On a PC that also games, runs ComfyUI or another model server, that means stopping and starting Strata by hand. With these options it can stay up permanently next to them.

## How

- Unloading ends the engine process (and the image encoder, when images are on), so VRAM and pinned RAM go straight back. Loading reuses the existing `restart()` path from #27 — the "engine died, start it again" check in `run()` now goes through one `ensure_loaded()` that also handles a deliberate unload.
- The image encoder is started again **before** the engine, as at a normal start, so a GPU encoder takes its VRAM before the engine sizes its expert cache. Its cache of encoded images is kept.
- Requests load the model before the answer starts, so a busy GPU is a clean JSON `503` rather than an error mid-stream.
- Free VRAM is read through the existing `_Nvml` class in `telemetry.py`; if it can't be read, nothing is refused.
- Unload only happens between requests (it takes the request FIFO non-blocking and checks `busy`/`queued`).

## Measured

RTX 5060 Ti 16 GB, i5-12400F, 48 GB DDR4, Linux (Docker), Q2_0 with `--mmap-experts`, files in the OS cache:

- `POST /unload` returns after ~0.33 s with the VRAM free (15.7 GB -> 0.76 GB used)
- a text request to an unloaded model answers after 4.6 s
- a request with a picture (encoder on the CPU) answers after 14.7 s
- with 8 GB of VRAM held by another process and `--min-free-vram-mib 11000`: `503` after the 15 s wait, model stays unloaded

I run it this way next to llama.cpp's router and ComfyUI: `--idle-unload 120 --min-free-vram-mib 11000 --before-load "<script that unloads llama-server's models>"`.

## Tests

`python -m unittest serve.test_server`: the 39 existing tests pass unchanged, plus 9 new ones in `SharingTheGpu` (unload then reload on the next request, `/load`, refused while busy, idle unload, the VRAM guard refusing and then loading, unreadable VRAM never refusing, `before_load` running first, the encoder unloading and starting first, off by default).

## Notes for review

- `POST /unload` and `/load` follow the API-key check like the other POST routes but not the `_own_page` check that `/settings` has, since other programs on the machine are the intended callers. The worst a cross-site request could do is unload an idle model.
- Not tested on Windows or with the GPU image encoder or multi-GPU — I only have the setup above. `terminate()` on Windows is a hard kill, which should be fine for a process between requests, but it is untested.
- Docs: a short section in `docs/DETAILS.md`. Setup does not ask about any of this; it seemed better left as an advanced option.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。