Pull requests / #581

#581 serve: the Monitor's per-GPU cards, each card's free VRAM, and the vi…

closed · @gopinath87607 · 0 commentaires · Sur GitHub

Server & APINVIDIA / CUDAModels & quants

Description

## What this is

The split is invisible in the UI: the metric cards show pooled totals, so a four-GPU run reads the
same as one card holding 27 GiB. Each GPU now gets its own card, built from what the engine alone
knows — its layer range, its KV and prompt buffers, its expert cache — paired with that card's VRAM
from NVML.

## The engine side

The engine reports its own allocation as seven parallel CSV fields on the `INFO` line:

    gpu_dev, gpu_layers, gpu_kv_mib, gpu_buf_mib, gpu_experts, gpu_expert_mib, gpu_draft_mib

They are single comma-separated tokens because the server's `INFO` parser splits on whitespace. A
remote-expert GPU runs no stage, so its layer span is `-`; the MTP head is reported on the one card
engine: it knows only its own allocations, and NVML already reports used/free/total.

Every device the engine uses gets exactly one card, and the cards' expert counts and bytes sum to the
`expert_slots` / `expert_cache_mib` totals above them — they are built from the same accessors. A
helper device is never also a stage device: `remote_dev` skips the split's own cards and device 0.

## The server side

NVML's `mem_free` was being read and thrown away, so telemetry now keeps it, and each card's name is
asked for once rather than once a second.

## The vision footprint

The image encoder is a separate process that reports no memory of its own — not to stdout, not to
the log — so its footprint is measured as what that card's free VRAM lost while it started.
`strata-vision` runs a 2048px warm-up before printing `READY`, so every buffer it will ever need is
allocated by the time the reading is taken; a delta outside 64–8192 MiB is noise rather than this
process and is discarded, keeping the last good number.

The encoder is also pinned to the card the config names (`vision.cuda_device`, else the engine's
first), which is the card the Monitor has to show it on.

## Folded in: two fixes found on the way

* `free_vram_mib` is one helper taking an index rather than two copies of the same NVML call.
* The vision encoder is pinned as described above.

A third fix that rode in the same commit on the old branch — an image source that 404s, refuses the
connection or names an unknown host answering with no reply at all instead of a 400 — is **not** here:
that is an image-API change, not a Monitor one, so it lives on `serve-tool-result-images`.

## Testing

`serve`'s 127 tests pass on `v0.1.38` (`python3 -m unittest serve.test_server`). The new ones are the
server half: the seven CSV fields surviving the `INFO` parser as single tokens, a `"-,835,-,-"` draft
list staying strings, the `vision` block on `/metrics`, and the vision footprint's noise rejection (an
implausible delta, a bad restart, no encoder, a stopped encoder). The card rendering is `app.js`,
which this repo does not test.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Sur le site

Liens install, modèles, releases.