Pull requests / #1590
#1590 serve: keep the given card order when the reorder starves the last stage (#1576)
open · @nekomario28 · 0 comments · View on GitHub
BenchmarksServer & APIAMD / HIP
Description
Fixes #1576. #1352's auto card order puts the faster card last because the last stage runs the head, the draft layer and the verify. On a VRAM-asymmetric pair the same rule lands the heaviest stage on the card least able to hold it — the reporter's fast 12 GB 4070 Ti got 31 layers + head + verify and a 3.1 GiB expert cache (decode ~58 → ~44 tok/s, the engine's own \"card this small\" warning). ## Change `ordered_gpus` skips the reorder when it would place the last stage on a card with less **total VRAM** than the card the config put there. VRAM comes from NVML (`mem_total`) or amdgpu sysfs; when any card is unreadable the reorder stays (previous behavior). `scores`/`vrams` are injectable for tests. This is deliberately conservative — it only keeps the given order, it does not invent a third arrangement. The reporter's other idea (reorder but let the split compensate) needs engine-side split changes, which this does not touch. ## Test plan - `serve.test_server.GpuChoice.test_card_order` extended: the reporter's shape (fast+small / slow+big → keep given), faster-and-bigger still reorders, equal VRAM reorders, a card whose VRAM cannot be read reorders, 3-card case keeps. - Real probe verified locally on an AMD box: `gpu_vram_mib([0], hip=True)` → `16368` MiB. Not tested: a real two-GPU run — no such rig here; the policy is exercised through the injected-table path the existing tests already use.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.