Pull requests / #1293
#1293 Community benchmark: RX 7900 XTX 24 GB — ROCm 7.1.1 vs ROCm 10.2 nightly
open · @xyzzing · 0 comments · View on GitHub
BenchmarksAMD / HIPModels & quantsDocumentation
Description
Re-file of #745 — that PR was auto-closed on 2026-10-06 when this repository's main history was cleaned up (per the maintainer's note there, not a rejection). Same two commits re-based onto the new main; the content is a self-contained new folder under `bench/results/`, so nothing else in the diff moved. Original discussion: #745. --- ## Summary Results-only submission per docs/COMMUNITY_BENCHMARKS.md — no engine changes. **Hardware**: AMD Radeon RX 7900 XTX 24 GB (gfx1100), Ryzen 9 7900X, 96 GB DDR5, Fedora 44, PCIe Gen4 x16. **Model**: Qwen3.8-Flash-Next GSQ-RCO IQ3_S (ISTA, 2 shards), MTP loaded, expert cache `auto` pre-filled from the shipped routing profile, KV int8 with `--kv-resident 32768`, context 65536, prefill chunk 8192, greedy, 200-token cap, `--vram-reserve-mib 3072` (desktop co-resident). **Configurations tested**: the same serving config on two ROCm stacks — packaged **ROCm 7.1.1** vs a **ROCm 10.2 nightly SDK** (`libamdhip64.so.7.17.26392`, verified loaded). One identical request repeated back-to-back (~61 prompt tokens incl. template, 200 generated tokens, greedy), decode read from the API's per-request `timings.predicted_per_second`. The sequence deliberately measures the **warm-up curve**: request 1 cold, later requests warm. **Results** (medians): warm steady-state decode **88.4 tok/s on the nightly (n=5) vs 80.4 on 7.1.1 (n=6) = +10.0% toolchain delta**, config identical. Warm-up climbs ~+35% cold→warm on **both** stacks — a methodology finding: one-shot cells under-measure serving by 25–35%. **Secondary, 128K one-shot (one run each, flagged)**: decode flat in context (62.0 vs 61.0 tok/s, O(1) GDN); **128K prefill is confounded** — the nightly's hipBLASLt is 100500 and no gfx1100 table ships for it, so the engine falls back to plain hipBLAS (921 vs 1535 tok/s; the documented version-mismatch mechanism). **Follow-up (2026-10-04, second commit in this PR): the confound is resolved.** A matching `gfx1100-hipblaslt-100500` table was tuned on-device against the nightly SDK's own hipBLASLt (32 rows, tuner-gated) and the same 128K workload re-run on the same binary, arms differing only in `STRATA_HIPBLASLT_TUNING`: **prefill 1687.0 vs 926.1 tok/s median (3 clean cells each, ranges in the new folder) = +82% from the table alone**; the no-table arm reproduces this PR's 921 cell, confirming it measured the fallback, not the toolchain. Decode unchanged (~62, overlapping ranges). Table, shape set, tuning command and per-run data in `bench/results/2026-10-04-community-rx7900xtx-hipblaslt-100500/`. **Limitations**: single GPU; short chat-sized request (decode is flat in context for this model, measured 1K–128K elsewhere); per-request raw server logs were truncated by a restart between stacks — decode values were captured at request time from the API's `timings` field (runs.json marks one stall, kept and excluded); rocWMMA select arm compiled out of the nightly build; k8v4 untested (incompatible with `--kv-resident`). Per-run JSON (`runs.json` — every request, warm-up state and one marked stall), the candidate startup script, and a repeat script are in the folder. `bench_prefill.py` was not used for this comparison — the request/`timings` method above keeps both stacks identical. *Prepared with an AI engineering agent (GLM-5.3 / Z.ai) under human direction; every number is from our own recorded runs.*
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.