Pull requests / #1184
#1184 AMD HIP: the primary ordinal, and the helper caches beside the resident RAM mode
open · @tuandat3019 · 0 comments · View on GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
Description
Three small changes for a two-card AMD PC whose faster card does not come first in the runtime's
enumeration. On Windows the HIP runtime orders the devices itself - `HIP_VISIBLE_DEVICES` only
filters, it cannot reorder - so with both cards visible an engine run could only put the model on the
first one. On the measured PC the RX 6600 (gfx1032) is numbered 0 and the RX 6800 (gfx1030) 1;
before this, the only way to run the model on the 6800 was to hide the 6600 from the engine.
## What changes
1. `cmake/hip_backend.cmake`: gfx1032 (RX 6600) joins the unvalidated arch list; it shares the
gfx1030 / gfx1031 / gfx1034 dp4a path. `docs/AMD_HIP.md` gains a line with what was run.
2. `STRATA_PRIMARY_DEVICE=N`: picks the visible ordinal the engine's own session runs on. It is set
before any context or allocation exists (every other device switch in the engine saves/restores
the current device, so that one call is where the primary is picked). The helper caches
(`--expert-cache-device1..3`) take the visible devices other than the primary, the HIP arch check
covers the primary, the primary's session is created on the chosen device, and the peer tier's
`cudaDeviceCanAccessPeer` / `cudaDeviceEnablePeerAccess` queries use the same ordinal. A layer
split still assumes device 0 and is refused with this set. Unset or 0 is the current behavior.
3. The helper-GPU caches are allowed beside the resident RAM mode (`--resident-experts`, the low-RAM
mode) - the combination was refused before. The experts a helper holds are left out of the RAM
copy as a layer split's later stages are, and the two adaptive tiers already keep each other's
experts out (`helper_holds` in the primary's candidates, `RemoteExpertOpt::adapt`'s resident
table). `RemoteExperts::finish` also spins on its stream like `PeerExperts::finish`: on HIP the
CUDA branch's `cudaInitDevice(ScheduleSpin)` does not exist, so the device kept the default
policy and the blocking sync's wake-up sat on the pool's critical path for every layer.
`docs/SECOND_GPU.md` documents the combination and the sizing lesson below.
## Measured (RX 6800 + RX 6600, Windows 11)
- Hardware: RX 6800 16 GB (gfx1030) + RX 6600 8 GB (gfx1032), both PCIe 4.0 (the engine's probe
reads ~13 GB/s host->device); i5-11400F (6 cores / 12 threads); 32 GB of RAM; NVMe.
- Software: Windows 11, driver 32.0.21045.5002, ROCm 10.2.0a20260930 nightly (TheRock wheels),
engine 0.1.40 + this change, built with `tools\hip\build_windows.bat` and
`STRATA_HIP_ARCHS=gfx1030;gfx1032`.
- Model: Swift 1.5 (Qwen3.8-Flash-Next) IQ3_XXS, 332 experts, native pack; 160K context; INT8 KV
with 32,768 resident cells; MTP `--spec 4 --spec-min-p 0.6`.
- Method: 8K-token request, 256-token answers, `--resident-experts` on all arms, one warmup then 8
measured runs at a fixed seed (temperature 0.6); after the warmup the engine reuses ~99.9% of the
prompt, so these are warm-decode numbers. Decode tok/s are the engine's own per-request figures;
one machine, one model.
| arm | decode tok/s (mean of 8) | CPU experts per layer-window |
| --- | ---: | ---: |
| RX 6800 alone, resident mode | 37.5 | ~2.9 |
| + RX 6600 helper, 3400 slots | 41.4 | ~0.7 |
| + RX 6600 helper, 3600 slots | 43.3 | ~0.5 |
| + RX 6600 helper, 3800 slots | 44.2 | ~0.4 |
- The helper's cache takes the profile's next-ranked experts after the 6800's own cache; its rows
(about 10 per layer) cost about 0.3-0.6 ms per layer on the CPU pool's critical path, well under
the CPU rows and page faults they replace.
- The RAM copy shrinks to 10.5 GiB from 19.7 (the helper's experts are left out), so the combination
also relieves RAM pressure on a 32 GB PC.
## The sizing lesson (also in docs/SECOND_GPU.md)
`--expert-cache-device1 auto` reads `cudaMemGetInfo`, which on Windows reports the *process's* WDDM
budget: it cannot see the desktop, the vision encoder (a separate process), a streaming host or
other apps on that card. An auto-sized helper on such a card over-commits; WDDM backs the excess
with system RAM ("shared GPU memory"), and that memory comes out of the machine's RAM - on this 32 GB
PC it pushed the resident complement towards the pagefile and made the helper read its own cache over
PCIe. An explicit count (3800 here, of ~4100 that fit an empty card) kept the cache in real VRAM.
One more measurement note: when a helper swaps experts during a long session, a re-admitted expert
that the helper held at startup but the RAM copy left out is read from the model files again. With
the default `--adapt-swaps 96` that measured 0.5-1.0 GB of file reads per 8K-token request here
(averaging ~10-30 MB/s of SSD traffic through decode); `--adapt-swaps 8` cut it ~5x (70-400 MB per
request) with the same decode throughput (44.0 vs 44.2 tok/s), and `--adapt-swaps 0` removed it
entirely at a cost of ~7 tok/s.
## Not covered
- Three or four helper cards, and the layer split together with `STRATA_PRIMARY_DEVICE` (refused).
- gfx1032 as a primary card; it was only run as the helper.
- NVIDIA: `CUDA_VISIBLE_DEVICES` with `CUDA_DEVICE_ORDER=PCI_BUS_ID` already reorders, this is not
needed there.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.