Pull requests / #1657
#1657 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K on 0.1.41 - layer split: stock, the two gfx103x switches, the open PRs, and the decode window
open · @xjc10 · 0 コメント · GitHub で見る
BenchmarksMulti-GPUAMD / HIPModels & quantsDocumentationWindows
本文
Results-only community report, no engine changes. Folder `bench/results/2026-10-09-community-2x-rx-6900xt-0.1.41/` plus one row in `bench/results/COMMUNITY.md`. Same machine as the 2026-10-05 and 2026-10-06 reports (2x RX 6900 XT, gfx1030, PCIe 4.0 x8 each; Ryzen 5 5600X; 128 GB; ROCm 10.0). Two-card layer split only, IQ3_S at 131K, five arms, two runs per length at 4K / 32K / 128K, no vision encoder, zero stalls: | arm | prompt tok/s 4K / 32K / 128K | decode tok/s | |---|---|---| | stock 0.1.41 (`fb58e0db`) | 464 / 767 / — | 62.6 / 65.0 / — | | stock + `STRATA_HIP_ADAPT_KERNEL_COPY=1` | 461 / 767 / 869 | 61.5 / 65.1 / 60.8 | | + `STRATA_HIP_PROMPT_F16=1 STRATA_SH_STREAM=1` | 832 / 1,619 / 1,823 | 69.2 / 74.6 / 67.5 | | + #1149 #1151 #1167, rocBLAS solution cache, RDNA2 MMQ patches in ggml | 968 / 1,714 / 2,029 | 64.7 / 68.8 / 68.3 | | the same with `--spec 3 --spec-min-p 0.7 --pipeline-windows 2` | 968 / 1,707 / 2,041 | 77.2 / 82.3 / 75.7 | What is new against the 0.1.40.1 report from this box: - Stock 0.1.41 reads prompts 4-6% faster than stock 0.1.40.1 on the split; with the two switches 5-12% faster. The pull requests land where they did (968 / 1,685 / 2,003 then). - `STRATA_HIP_ADAPT_KERNEL_COPY=1` costs nothing measurable and the stock split read its 128K prompts at the GFXOFF default without the #884 stall; the 0.1.40.1 report had to disable GFXOFF through debugfs for that. - Decode moves by configuration after all: the shorter draft with the higher threshold **together with** `--pipeline-windows 2` is +19 / +20 / +11%; each alone had measured neutral or worse here (a 16-point spec x min-p grid, and pipeline-windows with the default draft). The README notes that `--pipeline-windows 2` makes greedy output differ between server starts on this machine. - The split point moves 32K decode by about 8% (0-24/25-47 without the vision encoder vs 0-26/27-47 with it, same binary); written up as an observation, not swept. Per-arm folders have `config.json`, `summary.json`, `results.json` and `BUILD.txt`; the engine logs and one-second telemetry are in the companion repository (https://github.com/xjc10/dual-6900xt-flash-next-tuning, `bench/results/e388/`), which also holds the tuning history behind the "PRs" arm and the six llama.cpp patches. Developed with an AI coding assistant; every number above was measured on 2x RX 6900 XT (gfx1030, PCIe 4.0 x8 each) / Ryzen 5 5600X, ROCm 10.0.
関連リンク
インストール・モデル・リリースへの站内リンク。