Pull requests / #805

#805 Lazy vision encoder: per-GPU elastic expert caches

closed · @jmnargi · 0 コメント · GitHub で見る

Server & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

本文

## What

The image encoder (strata-vision, ~1.74 GiB across the GPUs) no longer has to sit in VRAM 24/7. With `"lazy": true` in the config's `"vision"` section (plus `"vram_elastic": true`), it starts on the **first image request** and unloads after `"idle_s"` seconds without one — its VRAM stays with the expert caches the rest of the time.

## Verified on the 2xV100 service

| State | Resident experts (both GPUs) |
| --- | --- |
| Before (resident vision) | 12,688 |
| Lazy boot, vision idle | **14,005** (+1,317) |
| During an image request (vision up) | 12,427 |
| 60 s idle after it (vision unloaded, caches regrown) | **14,005** |

Image request answered cold in 5.9 s (zero-warm "Red square shape"): per-GPU hand-back → spawn → encode → reply. `/metrics` `expert_slots`/`expert_cache_mib` aggregate across GPUs at every stage.

## Engine changes (src/program/generate.cpp, src/core/expert_cache.*)

- `--vram-elastic` now segments **every layer-split stage cache** (it was disabled with a split: "a layer split has a cache per GPU"). Also themed: per-GPU reserves, so the caches can return VRAM to the image encoder (or any other program) and take it back.
- The `VRAM` command takes **one reserve per GPU**: `VRAM <r0> [<r1> ...]` (CUDA0's, then each stage's; missing values keep the startup reserve). Reply fields are comma lists, one value per cache.
- Shrink evicts only the layers each cache owns — `host_res` is one global table whose slot values index whichever cache owns the layer, so an eviction loop over all rows would hit stage-owned slots with the wrong cache's `live` count.
- `ExpertCache::keep_bytes_for_free`: shrink keeps whole segments, so a release request smaller than the tail segment used to round up past the arena and free **nothing** (GPU1's 412 MiB tail); now the command releases the right segment.
- Bare `VRAM` reverts the caches to their **full start-up size** (the reserve is only the sizing target), so the vision idle-unload path regrows completely — including a short tail segment the conservative grow would otherwise leave unmapped forever.
- `relend()` reshapes the prompt-path loan of **every** part against its own cache (was: the drive's part only).

## Server changes (serve/server.py)

- `"vision": {"lazy": true}` skips the boot spawn; `prepare()` starts the encoder inside the request FIFO after the per-GPU `VRAM` hand-back (the engine is idle; text requests never touch the vision path).
- Vision idle thread: unloads the encoder after `"idle_s"`s without image requests and issues a bare `VRAM`, regrowing both caches to full. Defaults: `"vram_mib"` `[1700, 600]` for a layer split, `[1800]` for one GPU (measured with the shipped BF16 mmproj).
- `POST /v1/vram` and the served `vram()` accept a per-GPU list.
- `/metrics` aggregation fix: after a VRAM reply the engine info now sums `expert_slots`/`expert_cache_mib` across GPUs (the engine's own INFO line reports the total; a per-GPU reply had reduced it to CUDA0's value).

## Docs

`docs/DETAILS.md`: #533 (per-GPU reserves, bare = full start-up size), new "Lazy vision" section; `--vram-elastic` help.

Notes/limits: the lazy vision hand-back requires one-request-at-a-time (no `"parallel"`) — documented; `VRAM` still refuses with `--peer-device`, the helper caches or the resident low-RAM mode.

関連リンク

インストール・モデル・リリースへの站内リンク。