Pull requests / #1114

#1114 Monitor: per-GPU view, more cards, and a layout you can rearrange

open · @Efs-O · 0 commentaires · Sur GitHub

Server & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

Description

## Multi-GPU
- A GPU selector in the Model state card: All / GPU 0 / GPU 1 … (a dropdown above 6 cards). It switches the GPU load, VRAM, temperature, power and PCIe cards to one card's values and history.
- NVIDIA cards the engine does not use (a vision-encoder card, for example) are listed and marked "not in use". All keeps the engine-only totals.
- PCIe shows in/out rates and link gen/width per card.
- A single-card machine looks as before: no selector.

## New cards and table columns
- Speculation (draft acceptance), Prompt reuse, Prefill time, VRAM free (the tightest card; warns under 1 GiB, danger under 0.5 GiB), all-GPU power, session time.
- Recent requests gains Prefill and Drafts columns, plus a Clear button that empties the table on this page only (nothing is sent to the server; a refresh brings the rows back).
- Temperature graphs scale to their own range, so a 25-30 °C swing is visible instead of a flat line.
- Missing values say "No requests yet" or "Not reported by this engine" instead of a bare dash.

## Optional sensors
- **CPU temperature and power** from LibreHardwareMonitor's local web server (`127.0.0.1:8085` by default, `STRATA_LHM_URL` to change it, empty to disable). This adds a CPU temp card, the package power under the CPU load (the CPU card, renamed to match "GPU load") and a measured "GPUs + CPU" power card. It reads the first CPU's package sensors by name ("CPU Package" on Intel; "Core (Tctl/Tdie)" and "Package" on AMD), because LHM's sensor numbers differ between CPU models. On Windows without LHM running, the CPU temp card says "Needs LibreHardwareMonitor" and its tooltip explains how to turn on LHM's web server. On other systems the card is hidden.
- **Forge card, off by default.** With `STRATA_FORGE_URL=http://127.0.0.1:8799` set, it shows the active chat of Forge (a VS Code coding assistant for local models): turns, tool calls and tokens, read from Forge's local, read-only `/stats`. Disclosure: Forge is my own project and I run it on top of Strata. If you'd rather not have it in Strata, say so and I'll drop it from the PR.

Both are localhost-only. LHM is polled every 2 s and backs off for 30 s after a failure. Forge is never contacted unless `STRATA_FORGE_URL` is set.

## Layout
- The cards sit in one 4-column grid (2 columns on narrow screens).
- Drag a card to reorder it: with the mouse, with press-and-hold on touch, or with Alt+←/→ from the keyboard (the new position is announced to screen readers). Dropping outside the cards puts it back.
- The × on a card hides it.
- Order and hidden cards are saved in the browser (`localStorage`). A "Reset layout · 2 icons hidden" link above the cards brings everything back. It only appears when something has changed.

## /metrics
Every new field is optional, and single-card and AMD output is unchanged:
- `hardware.gpus[*]`: `power_limit`, `pcie_rx_mb`, `pcie_tx_mb`, `pcie_gen`, `pcie_gen_max`, `pcie_width`
- `history.gpus`: per-card history
- `hardware.other_gpus` and `hardware.all_gpu_power`
- `hardware.cpu_temp` and `hardware.cpu_power` (LHM)
- `hardware_static.os`
- `forge` at the top level (`null` when the card is off)

They are documented in `docs/DETAILS.md`.

## Testing
- `python -m unittest serve.test_monitor serve.test_server serve.test_forge_stats serve.test_lhm` on current `main`: 238 pass. The AMD sensor names are covered by a test fixture; I have no AMD machine to try it on.
- Checked in Edge at 1600 and 390 px wide, in light and dark: drag with the mouse, emulated touch and the keyboard; hide, reload, reset.
- Larger machines simulated by rewriting `/metrics` in the browser: 6 cards (buttons, wrapping to two rows at 390 px), 5 + 1 not in use, and 7 + 1 (dropdown). The per-card lines switch to summaries above 3 cards, with every card in the tooltip.
- Run on Windows with an i9-9900KF, two RTX 5060 Ti running the model and an RTX 3060 for the vision encoder, with LHM running.
- Rebased onto current `main` as one commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Sur le site

Liens install, modèles, releases.