Pull requests / #854
#854 expert plan: a helper GPU's experts are not part of the PCIe share (--remote-expert-opt decode 40 -> 66 tok/s on 2x RX 6900 XT)
closed · @xjc10 · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
描述
## The bug The verify window's GPU plan (`expert_source.cpp`) counts the layer's misses and moves the last `pcie_num/256` of them to the primary over PCIe. The count and the pick leave out experts resident on the primary and those of the peer tier (`--peer-device`), but not those a helper GPU (`RemoteExperts`, `--expert-cache-deviceN`) holds: the plan runs before `RemoteExperts::begin()` claims its rows, and `begin()` only takes rows the plan left at kind -1. So with a PCIe share above 0, part of the helper's experts is read over the primary's link and computed there, while the helper sits idle for them. setup recommends `--remote-expert-opt` for 2+ GPUs (#578), and the PCIe share is probed per link (0.39 on PCIe 4.0 x8 here), so this hits the recommended two-card configuration. ## The fix `RemoteExperts::holds(layer, expert)` (the helper cache's `slot_of`, as `RemoteExpertOpt::owns` already reads it), and the plan treats such an expert as it treats the peer tier's: not a miss, never in the PCIe share. Its rows stay at kind -1 for `begin()`. The PCIe share now only takes experts no GPU holds, i.e. ones the CPU would otherwise compute. No default changes; without helpers nothing changes. ## Measurements 2x RX 6900 XT (PCIe 4.0 x8 each, probe 14.1 GB/s -> pcie_frac 0.39), ROCm 10.0, Qwen3.8-Flash-Next GSQ-RCO IQ3_S, `--expert-cache auto --expert-cache-device1 auto --remote-expert-opt`, `--spec 4`, two ~5K-token prompts with ~450-token answers, `STRATA_DECODE_TIMING=1`: | | ms / window | GPU-reach wait | PCIe entries / layer-window | CPU entries / layer-window | decode tok/s | | --- | ---: | ---: | ---: | ---: | ---: | | main (probed 0.39) | 59.6 / 58.5 | 43.4 | 4.00 | 0.7 | 40.6 / 39.3 | | **this change (probed 0.39)** | 37.2 / 33.8 | 19.5 / 18.6 | 0.26 / 0.06 | 1.17 / 0.52 | **65.6 / 67.7** | | this change, `--pcie-frac 0` | 36.5 / 35.1 | 18.8 | 0 | 1.39 / 0.76 | 66.9 / 67.1 | With the fix the probed share costs nothing against 0 and takes a little off the CPU. Until then, `--pcie-frac 0` is the workaround (main with it: 34.1-34.6 ms per window, 60.5 / 67.7 tok/s). ## Not tested CUDA (no CUDA here; the change is backend-independent host code). Three or four helper GPUs (`remote_count` > 1 is covered by the same loop). The prompt path is not affected (it does not use this plan). ## Validation and docs - **ctest** (`-DSTRATA_BUILD_TESTS=ON`, the same options as the engine build, gfx1030, ROCm 10.0): 60 of 65 pass, the same 60 as `6f32ec0` without this change. Skipped: `hip_prompt_attn_wmma` (no matrix cores on gfx1030), `hip_prefill_hipblaslt_gemm` (no hipBLASLt table). Failed for reasons outside the engine, identically with and without the change: `ple_parity` (a Q2_0 PLE file that is not on this machine), `expert_multi_test` (this Zen 3 CPU has no AVX-512), `platform_memory_test` (`mlock` at the shell's default `ulimit -l`). - **Same-base baseline** (2026-10-05): the table above measured the baseline on `6f32ec0` + #835 (the binary this machine had; #835 touches only the prompt path). Rerun on `6f32ec0` unchanged, same method: 60.4 / 57.8 ms per window, GPU-reach wait 43.7 / 43.0, PCIe entries 4.07 / 3.91 per layer-window, 41.0 / 37.5 tok/s - the same as the table's baseline. The fix rerun the same day: 35.2 / 33.8 ms, PCIe 0.20 / 0.06, 63.0 / 64.4 tok/s. - **Docs:** `docs/SECOND_GPU.md`'s helper section no longer says HIP is unmeasured: it states the two-card HIP measurement, the cost of the share before this change and the `--pcie-frac 0` workaround; details and the configuration in `bench/results/2026-10-04-rdna2-helper-pcie-share/README.md`. Developed with an AI coding assistant; every number above was measured on 2x RX 6900 XT (gfx1030, PCIe 4.0 x8 each) / Ryzen 5 5600X, ROCm 10.0.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。