Pull requests / #854

#854 expert plan: a helper GPU's experts are not part of the PCIe share (--remote-expert-opt decode 40 -> 66 tok/s on 2x RX 6900 XT)

closed · @xjc10 · 0 comentarios · En GitHub

BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Descripción

## The bug

The verify window's GPU plan (`expert_source.cpp`) counts the layer's misses and moves the last `pcie_num/256` of them to
the primary over PCIe. The count and the pick leave out experts resident on the primary and those of the peer tier
(`--peer-device`), but not those a helper GPU (`RemoteExperts`, `--expert-cache-deviceN`) holds: the plan runs before
`RemoteExperts::begin()` claims its rows, and `begin()` only takes rows the plan left at kind -1. So with a PCIe share
above 0, part of the helper's experts is read over the primary's link and computed there, while the helper sits idle for
them.

setup recommends `--remote-expert-opt` for 2+ GPUs (#578), and the PCIe share is probed per link (0.39 on PCIe 4.0 x8 here),
so this hits the recommended two-card configuration.

## The fix

`RemoteExperts::holds(layer, expert)` (the helper cache's `slot_of`, as `RemoteExpertOpt::owns` already reads it), and the
plan treats such an expert as it treats the peer tier's: not a miss, never in the PCIe share. Its rows stay at kind -1
for `begin()`. The PCIe share now only takes experts no GPU holds, i.e. ones the CPU would otherwise compute. No default
changes; without helpers nothing changes.

## Measurements

2x RX 6900 XT (PCIe 4.0 x8 each, probe 14.1 GB/s -> pcie_frac 0.39), ROCm 10.0, Qwen3.8-Flash-Next GSQ-RCO IQ3_S,
`--expert-cache auto --expert-cache-device1 auto --remote-expert-opt`, `--spec 4`, two ~5K-token prompts with ~450-token
answers, `STRATA_DECODE_TIMING=1`:

| | ms / window | GPU-reach wait | PCIe entries / layer-window | CPU entries / layer-window | decode tok/s |
| --- | ---: | ---: | ---: | ---: | ---: |
| main (probed 0.39) | 59.6 / 58.5 | 43.4 | 4.00 | 0.7 | 40.6 / 39.3 |
| **this change (probed 0.39)** | 37.2 / 33.8 | 19.5 / 18.6 | 0.26 / 0.06 | 1.17 / 0.52 | **65.6 / 67.7** |
| this change, `--pcie-frac 0` | 36.5 / 35.1 | 18.8 | 0 | 1.39 / 0.76 | 66.9 / 67.1 |

With the fix the probed share costs nothing against 0 and takes a little off the CPU. Until then, `--pcie-frac 0` is the
workaround (main with it: 34.1-34.6 ms per window, 60.5 / 67.7 tok/s).

## Not tested

CUDA (no CUDA here; the change is backend-independent host code). Three or four helper GPUs (`remote_count` > 1 is
covered by the same loop). The prompt path is not affected (it does not use this plan).

## Validation and docs

- **ctest** (`-DSTRATA_BUILD_TESTS=ON`, the same options as the engine build, gfx1030, ROCm 10.0): 60 of 65 pass, the same 60 as `6f32ec0` without this change. Skipped: `hip_prompt_attn_wmma` (no matrix cores on gfx1030), `hip_prefill_hipblaslt_gemm` (no hipBLASLt table). Failed for reasons outside the engine, identically with and without the change: `ple_parity` (a Q2_0 PLE file that is not on this machine), `expert_multi_test` (this Zen 3 CPU has no AVX-512), `platform_memory_test` (`mlock` at the shell's default `ulimit -l`).
- **Same-base baseline** (2026-10-05): the table above measured the baseline on `6f32ec0` + #835 (the binary this machine had; #835 touches only the prompt path). Rerun on `6f32ec0` unchanged, same method: 60.4 / 57.8 ms per window, GPU-reach wait 43.7 / 43.0, PCIe entries 4.07 / 3.91 per layer-window, 41.0 / 37.5 tok/s - the same as the table's baseline. The fix rerun the same day: 35.2 / 33.8 ms, PCIe 0.20 / 0.06, 63.0 / 64.4 tok/s.
- **Docs:** `docs/SECOND_GPU.md`'s helper section no longer says HIP is unmeasured: it states the two-card HIP measurement, the cost of the share before this change and the `--pcie-frac 0` workaround; details and the configuration in `bench/results/2026-10-04-rdna2-helper-pcie-share/README.md`.

Developed with an AI coding assistant; every number above was measured on 2x RX 6900 XT (gfx1030, PCIe 4.0 x8 each) / Ryzen 5 5600X, ROCm 10.0.

En el sitio

Enlaces a install, modelos, releases.