Pull requests / #1105
#1105 serve: the Monitor's per-GPU cards, each card's free VRAM, and the vision footprint
open · @gopinath87607 · 0 评论 · 在 GitHub 查看
Server & APIMulti-GPUNVIDIA / CUDAModels & quantsLinux
描述
The split is invisible in the UI: the metric cards show pooled totals, so a four-GPU run reads the same as one card holding 27 GiB. Each GPU now gets its own card, built from what the engine alone knows - its layer range, its KV and prompt buffers, its expert cache - paired with that card's VRAM from NVML. This replaces #581, which GitHub closed automatically when `main` was force-pushed on Oct 6. It is the same change, rebased onto the new `main` (`82f46a8`) as one commit; the only real work was keeping upstream's newer `Vision` (the #480 work-dir fix and the `popen` helper) alongside the `device` parameter and footprint probe added here. The engine reports its own allocation as seven parallel CSV fields on the INFO line (gpu_dev, gpu_layers, gpu_kv_mib, gpu_buf_mib, gpu_experts, gpu_expert_mib, gpu_draft_mib). They are single comma-separated tokens because the server's INFO parser splits on whitespace. A remote-expert GPU runs no stage, so its layer span is `-`; the MTP head is reported on the one card that holds it and nowhere else, since no totals line accounts for it. No VRAM figure comes from the engine: it knows only its own allocations, and NVML already reports used/free/total. For that last part NVML's `mem_free` was being read and thrown away, so telemetry now keeps it, and each card's name is asked for once rather than once a second. The image encoder is a separate process that reports no memory of its own - not to stdout, not to the log - so its footprint is measured as what that card's free VRAM lost while it started. `strata-vision` runs a 2048px warm-up before printing READY, so every buffer it will ever need is allocated by the time the reading is taken; a delta outside 64..8192 MiB is noise rather than this process and is discarded, keeping the last good number. It is shown as approximate (`~`) for that reason. Also folds in two fixes found on the way: `free_vram_mib` is one helper taking an index rather than two copies of the same NVML call; and the vision encoder is pinned to the config's first GPU, which is the card the Monitor has to show it on. (An image source that fails now answers 400 rather than an empty reply - that one is an image-API fix, so it is on serve-tool-result-images, not here.) An engine too old to report the seven fields gets no panel: the grid stays hidden and the metric cards' existing per-GPU sub-lines stand in, unchanged. ## Testing - `python3 -m unittest serve.test_server` - **213 tests pass** on the rebased branch (Xeon E5-2680 v4, 2x RTX 3060 12 GB + 2x RTX 5060 8 GB, Linux, Python 3.14). New tests: VisionFootprint (the delta is what the card lost / an implausible delta is not reported / a bad restart keeps the last good reading), VisionMetrics (no encoder, a double without `alive`, a stopped encoder holds nothing, a running one reports its card and size), and GpuDraftField (the lists survive as single tokens - the INFO parser must not keep only the first number - and a lone `-` stays a string). - The branch builds clean on the new `main`: `cmake --build`, CUDA 13.2, arches 86;120, gcc 15.2, exit 0. - Served check: a lazy server from this branch serves the page with `id="gpu-grid"`, `web/app.js` carrying `renderGpuPanel`, and `web/app.css` with the card styles. **Not tested, and worth knowing before a merge:** the per-GPU card rendering in `app.js` has no test - the same gap the original PR noted. And the seven fields' values have not been read back from a booted engine on a real split; that needs the four cards, which were busy. Everything above is compile-, unit- and serve-level. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。