Pull requests / #1580

#1580 bench: add community benchmark report, AMD Radeon AI PRO R9700 (gfx12…

open · @KongChengZhi · 0 コメント · GitHub で見る

BenchmarksAMD / HIPModels & quantsDocumentationWindows

本文

…01), Windows

Decode and prompt throughput for Qwen3.8-Flash-Next IQ3_XXS on one R9700 at 17,154 / 99,696 / 132,886 / 199,316 actual prompt tokens, plus a single-variable A/B showing STRATA_SH_STREAM=0 worth +25% to +42% decode throughput, with the mechanism visible as a drop in the engine's per-window GPU-reach wait (27.4-28.8 ms -> 16.5-17.9 ms, ranges not overlapping) while draft acceptance, cache hit rate and tokens per window stay unchanged.

Results only: no engine, no docs and no source changes. 26 arm JSONs, the harness, the per-arm configs and the five drivers are included; model files, the generated pack and the engine logs are left out per docs/COMMUNITY_BENCHMARKS.md.

Each comparison uses a base anchor from the same window, and windows whose two base anchors differ by more than 5% are reported as invalid (two of five). n=3 per arm with no confidence intervals, TTFT not measured, prompt throughput n=1 per arm, synthetic prompts only with a 128-token output cap, and the GPU was shared with other sessions for part of the day.

No correctness claim is made: this build is not deterministic run to run (two cold base starts diverged at word 72), so the acceptance of STRATA_SH_STREAM=0 rests on a structural argument (it only moves streams and synchronisation points) and on unchanged output distributions, not on numerical equivalence.

## Title
<!-- If Applicable, reference the GitHub issue -->
Issue: Resolves # 

## Summary
<!-- Quick Summary of changes -->

## What changed
<!-- Specifics on files changed, and what changes were made there -->

## Extra Notes
<!-- Any extra notes, delete if there are none -->

関連リンク

インストール・モデル・リリースへの站内リンク。