Pull requests / #105
#105 The Monitor's GPU readings on ROCm: an amdgpu sysfs backend, and the disk rate without psutil
closed · @xyzzing · 0 comments · View on GitHub
Setup & installServer & APIAMD / HIPWindowsLinux
Description
## What
The Monitor's hardware half is NVML-only, so it is empty on a ROCm box. Measured on this machine (Fedora 44, ROCm 7.1.1, **Radeon RX 7900 XTX / gfx1100**) with `libnvidia-ml.so.1` absent:
- `GET /metrics → hardware` carried **no `gpu_*` key at all** — GPU load, VRAM, GPU temp, power and PCIe all showed `–`
- `hardware_static.gpu_name` was `null`, so **About → GPU** printed *"not readable (NVML)"*
- with `psutil` not installed (the venv this port runs in), the disk card showed *"needs psutil (setup installs it)"* and the physical-core count was missing (`"psutil": false`)
## What changed
**`_AmdSysfs`** — the same fields from amdgpu's sysfs, used only when `_Nvml` fails to load:
- the card is found by scanning `/sys/class/drm/card*/device` for `DRIVER=amdgpu` (the index is not fixed: this box has the AMD card at `card1`, not `card0`)
- `gpu_busy_percent` → util; `mem_info_vram_used` / `mem_info_vram_total` → VRAM
- the hwmon sensor **labelled `edge`** → temperature (what NVML calls the GPU temperature; junction and memory sensors are 55 °C and 52 °C at the same moment on this card)
- `power1_average` / `power1_cap` (µW) → power and its limit
- `current_link_speed` / `max_link_speed` (`"16.0 GT/s PCIe"` → Gen4) and `current_link_width` → the link
Every reading is a small file read, so a sample cannot block the server, and anything the kernel does not expose stays `None` rather than being invented. amdgpu has **no PCIe throughput counter** (`rocm-smi --showbw --json` answers *"No JSON data to report"* here), so the link is reported and the MB/s series stays empty — `serve/web` already renders "Gen4 x16" from those fields, so no client change was needed.
**Without psutil** (Linux): `/proc/diskstats` for the read/write rate — the source psutil itself reads — counting whole disks only (never partitions, and never `loop`/`zram`/`dm-*`, which would count a device-mapper layer twice), and `/proc/cpuinfo`'s `(physical id, core id)` pairs for the physical-core count.
## Measured after the change
```
hardware : gpu_util 12%, gpu_mem_used 23.0/24.0 GB, gpu_temp 49.0 °C,
gpu_power 35.0 W of 339 W, gpu_pcie_gen 4, gpu_pcie_width 16, disk_read_mb live
static : gpu_name "Navi 31 [Radeon RX 7900 XT/XTX/GRE] (gfx1100)", cores 12, threads 24, psutil false
```
With `psutil` forced absent, every GPU series, the disk rate and the 12 cores still come through.
## Compatibility
NVML still wins whenever it loads and the `psutil` path is untouched, so an NVIDIA install behaves exactly as before; the sysfs and `/proc` paths are only ever reached when the previous backend is unavailable (the `/proc` reads return `None` off Linux, which is what the existing Windows fallbacks already expect).
## Tests
New `serve/test_telemetry.py` — no GPU, no pack, no driver:
- the `GT/s` → generation mapping, and the whole-disk filter (`sda` yes, `sda1`/`dm-0`/`zram0`/`loop0` no)
- **no amdgpu card → nothing at all**: `ok()` false, `read() == {}`, `name() == "AMD GPU"` (rather than a traceback)
- a card present → sane values (`mem_used <= mem_total`, `0 <= util <= 100`, temperature in range when the sensor exists)
- `Telemetry.sample()` still merges the server's own `extra()` and still carries every series the Monitor plots
```
python -m unittest serve.test_telemetry -v # 5 tests, OK
python -m unittest serve.test_server serve.test_mcp serve.test_telemetry # 55 tests, OK
```
Independent of #104 — no shared file. Found while porting this engine to HIP/gfx1100, where the Monitor was the last thing still speaking NVIDIA.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.