Pull requests / #1270

#1270 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K on 0.1.40.1 - one card, layer split and expert helper; stock, the two gfx103x switches, and #1150 #1151 #1167

closed · @xjc10 · 0 commentaires · Sur GitHub

BenchmarksMulti-GPUAMD / HIPModels & quantsDocumentationLinux

Description

Data only: one folder under `bench/results/` and its line in `docs/COMMUNITY_BENCHMARKS.md`. Re-filed from #927 (closed with the history rewrite), measured again on 0.1.40.1 instead of 0.1.39.

2x RX 6900 XT (gfx1030, PCIe 4.0 x8 each), Ryzen 5 5600X, 128 GB, Ubuntu 26.04 / Linux 7.0 / ROCm 10.0.0, Qwen3.8-Flash-Next GSQ-RCO IQ3_S at a 131,072-token context, the repository's `benchmark.py` (4,096 / 32,768 / 128,000 tokens x 3) and `needle_bench.py` (six needles), one cold server per configuration, adaptive swaps on. Nine configurations: one card, the layer split and the expert-helper mode, each as stock `82f46a8`, with the two switches 0.1.40 ships for gfx103x (`STRATA_HIP_PROMPT_F16=1 STRATA_SH_STREAM=1`), and with #1150 + #1151 + #1167 on top.

| | one card | layer split | helper |
| --- | ---: | ---: | ---: |
| prompt tok/s at 4K / 32K / 128K, stock | 458 / 473 / 458 | 446 / 723 / 821 | 459 / 474 / 459 |
| the two switches | 811 / 973 / 913 | 791 / 1,443 / 1,631 | 816 / 975 / 913 |
| + #1150 #1151 #1167 | 969 / 1,152 / 1,126 | 968 / 1,685 / 2,003 | 980 / 1,151 / 1,127 |
| decode tok/s (range over the three lengths), stock | 47-48 | 59-67 | 64-71 |
| the two switches | 49-51 | 65-73 | 69-75 |
| + the PRs | 50-52 | 70-75 | 66-72 |

All 54 needles found. Two things the README says in full: the two-card configurations were measured with GFXOFF disabled, because at the kernel's default the stock and the switches-only layer split both stalled on their first 128,000-token request with the adaptive swaps on (#884; the stalled runs' 4K and 32K rows are kept and match); and the decode gain over the 0.1.39 report from this machine (8-18%) is the swaps and the 0.1.40 changes together.

Developed with an AI coding assistant; every number above was measured on 2x RX 6900 XT (gfx1030, PCIe 4.0 x8 each) / Ryzen 5 5600X, ROCm 10.0.

Sur le site

Liens install, modèles, releases.