Issues / #1272

#1272 0.1.40: with the shared-expert stream fork off by default, decode on 2x RX 6900 XT (gfx1030, Linux) is 5-13% slower; STRATA_SH_STREAM=1 restores it, and the #884 stall does not come back with it

open · @xjc10 · 1 评论 · 在 GitHub 查看

BenchmarksMulti-GPUAMD / HIPModels & quantsWindowsLinux

描述

On the 2x RX 6900 XT machine from #816 (gfx1030, Linux, Ubuntu 26.04 / Linux 7.0 / ROCm 10.0.0, Qwen3.8-Flash-Next GSQ-RCO IQ3_S, `--kv int8 --kv-resident 32768 --max-context 131072 --spec 4`), the 0.1.40 default - the shared-expert stream fork off on HIP - costs decode, the opposite of the RX 6800 / Windows report that led to it. Two measurements, both on a clean build of 0.1.40.1 (`82f46a8`), cold server per arm, temperature 0:

**The switch alone** (the same 50,517-token prompt, 256 tokens, three requests per arm, decode median of 3):

| | default (fork off) | `STRATA_SH_STREAM=1` |
| --- | ---: | ---: |
| one card | 37.8 tok/s | **40.4 (+7%)** |
| layer split over both cards | 58.6 tok/s | **64.5 (+10%)** |
| prompt read, one card / split | 458 / 745 tok/s | 458 / 753 (unchanged) |

**The full matrix** (the repository's `benchmark.py`, 4,096 / 32,768 / 128,000 tokens x 3, adaptive swaps on; `STRATA_HIP_PROMPT_F16=1` is on in the second arm as well, which does not touch decode): decode 47.5 / 46.9 / 48.3 -> 50.5 / 48.5 / 51.3 tok/s on one card, 62.1 / 67.0 / 59.0 -> 70.4 / 73.3 / 64.7 on the split, 71.3 / 66.4 / 63.9 -> 74.9 / 71.5 / 68.7 with the helper - 5-13%. The whole run is in the community report filed today (the folder `bench/results/2026-10-06-community-2x-rx-6900xt-0.1.40.1`).

**The fork and #884.** If the default was chosen partly because the fork is one of the stall's ingredients: on this machine, with GFXOFF at the kernel's default and the same request sequence that stalls 0.1.39 in 7 of 8 runs with the fork on (helper mode, prompts up to 32K, adaptive swaps on), 0.1.40.1 with `STRATA_SH_STREAM=1` stalled in **0 of 8**, measured the same afternoon, alternating. At 128,000-token prompts on the layer split, 0.1.40.1 stalls with the fork **off** (stock) and on alike - two runs, two stalls each, all while reading the prompt - and not at all with GFXOFF disabled (six two-card runs, 54 requests). So here the fork does not decide the stall; the power-gating state does. Details are in #884.

Since the sign of the switch differs between an RX 6800 on Windows and an RX 6900 XT on Linux, the default may need to be per platform, or measured: `calibrate.py` could add the fork as an arm and keep whichever is faster on the card at hand. Happy to write that PR if you want it, and to run anything else that helps decide.

Developed with an AI coding assistant; every number above was measured on 2x RX 6900 XT (gfx1030, PCIe 4.0 x8 each) / Ryzen 5 5600X, ROCm 10.0.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。