Pull requests / #1105

#1105 serve: the Monitor's per-GPU cards, each card's free VRAM, and the vision footprint

open · @gopinath87607 · 0 Kommentare · Auf GitHub

Server & APIMulti-GPUNVIDIA / CUDAModels & quantsLinux

Beschreibung

The split is invisible in the UI: the metric cards show pooled totals, so a four-GPU run
reads the same as one card holding 27 GiB. Each GPU now gets its own card, built from what
the engine alone knows - its layer range, its KV and prompt buffers, its expert cache -
paired with that card's VRAM from NVML.

This replaces #581, which GitHub closed automatically when `main` was force-pushed on
Oct 6. It is the same change, rebased onto the new `main` (`82f46a8`) as one commit; the
only real work was keeping upstream's newer `Vision` (the #480 work-dir fix and the `popen`
helper) alongside the `device` parameter and footprint probe added here.

The engine reports its own allocation as seven parallel CSV fields on the INFO line
(gpu_dev, gpu_layers, gpu_kv_mib, gpu_buf_mib, gpu_experts, gpu_expert_mib,
gpu_draft_mib). They are single comma-separated tokens because the server's INFO parser
splits on whitespace. A remote-expert GPU runs no stage, so its layer span is `-`; the MTP
head is reported on the one card that holds it and nowhere else, since no totals line
accounts for it. No VRAM figure comes from the engine: it knows only its own allocations,
and NVML already reports used/free/total.

For that last part NVML's `mem_free` was being read and thrown away, so telemetry now
keeps it, and each card's name is asked for once rather than once a second.

The image encoder is a separate process that reports no memory of its own - not to stdout,
not to the log - so its footprint is measured as what that card's free VRAM lost while it
started. `strata-vision` runs a 2048px warm-up before printing READY, so every buffer it
will ever need is allocated by the time the reading is taken; a delta outside 64..8192 MiB
is noise rather than this process and is discarded, keeping the last good number. It is
shown as approximate (`~`) for that reason.

Also folds in two fixes found on the way: `free_vram_mib` is one helper taking an index
rather than two copies of the same NVML call; and the vision encoder is pinned to the
config's first GPU, which is the card the Monitor has to show it on. (An image source that
fails now answers 400 rather than an empty reply - that one is an image-API fix, so it is
on serve-tool-result-images, not here.)

An engine too old to report the seven fields gets no panel: the grid stays hidden and the
metric cards' existing per-GPU sub-lines stand in, unchanged.

## Testing

- `python3 -m unittest serve.test_server` - **213 tests pass** on the rebased branch
  (Xeon E5-2680 v4, 2x RTX 3060 12 GB + 2x RTX 5060 8 GB, Linux, Python 3.14). New tests:
  VisionFootprint (the delta is what the card lost / an implausible delta is not reported /
  a bad restart keeps the last good reading), VisionMetrics (no encoder, a double without
  `alive`, a stopped encoder holds nothing, a running one reports its card and size), and
  GpuDraftField (the lists survive as single tokens - the INFO parser must not keep only
  the first number - and a lone `-` stays a string).
- The branch builds clean on the new `main`: `cmake --build`, CUDA 13.2, arches 86;120,
  gcc 15.2, exit 0.
- Served check: a lazy server from this branch serves the page with `id="gpu-grid"`,
  `web/app.js` carrying `renderGpuPanel`, and `web/app.css` with the card styles.

**Not tested, and worth knowing before a merge:** the per-GPU card rendering in `app.js`
has no test - the same gap the original PR noted. And the seven fields' values have not
been read back from a booted engine on a real split; that needs the four cards, which were
busy. Everything above is compile-, unit- and serve-level.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Mehr auf der Site

Links zu Install, Modellen, Releases.