Pull requests / #1292
#1292 Docs: AMD RX 7900 XTX (gfx1100) IQ3_S rates row (1K-128K)
open · @xyzzing · 0 Kommentare · Auf GitHub
BenchmarksAMD / HIPModels & quantsDocumentation
Beschreibung
Re-file of #319 — that PR was auto-closed on 2026-10-06 when this repository's main history was cleaned up (per the maintainer's note there, not a rejection). Same three commits, re-based onto the new main. One correction the re-file made: the original branch's first commit carried a stale copy of `docs/DETAILS.md` alongside the intended row, so its diff showed ~190 lines of unrelated deletions of upstream text. This refile carries exactly the rates-row payload — the whole branch diff is +12/−2, the two removed lines being the caption lead lines the AMD captions extend. Original discussion: #319. --- Adds one **IQ3_S (AMD RX 7900 XTX, gfx1100)** row to the README/DETAILS rates tables, per the maintainer's request to split the rates row out of #228 and re-measure. **What it adds** (docs only, current head): - README rates table: `| **IQ3_S** (AMD RX 7900 XTX) | 60 tokens/s | 1,641 tokens/s |` - DETAILS prompt table (1K/4K/32K/64K/128K/262K): `760 | 1,275 | 1,641 | 1,594 | 1,494 | -` - DETAILS output (decode) table: `65.1 | 60.4 | 64.4 | 61.8 | 59.2 | -` **Provenance (updated for the current head).** Engine 0.1.38 on a ROCm 10.2 nightly with the tuned gfx1100 hipBLASLt-100500 table (upstream as #755), median of 3 non-stalled cells per tier from a 1K-128K one-shot sweep (256 generated tokens, greedy; Ryzen 9 7900X, 96 GB). Default configuration — **no** WMMA opt-ins (see the thread). 262K stays `-`, same as the existing IQ3_S/IQ3_XXS rows. The flat decode is expected: GDN linear attention makes decode O(1) in context. On-device kernel-shape autotuning (#744) was run on this machine and kept every shipped shape — the row's kernels are the shipped defaults. Stalls were recorded and excluded from the medians. Docs-only change; no code. (The first revision of this description carried the 0.1.31 packaged-ROCm numbers; the commits above supersede them.) *Prepared with an AI engineering agent (GLM-5.3 / Z.ai) under human direction; every number is from our own recorded runs.*
Mehr auf der Site
Links zu Install, Modellen, Releases.