Pull requests / #1292

#1292 Docs: AMD RX 7900 XTX (gfx1100) IQ3_S rates row (1K-128K)

open · @xyzzing · 0 Kommentare · Auf GitHub

BenchmarksAMD / HIPModels & quantsDocumentation

Beschreibung

Re-file of #319 — that PR was auto-closed on 2026-10-06 when this repository's main history was cleaned up (per the maintainer's note there, not a rejection). Same three commits, re-based onto the new main.

One correction the re-file made: the original branch's first commit carried a stale copy of `docs/DETAILS.md` alongside the intended row, so its diff showed ~190 lines of unrelated deletions of upstream text. This refile carries exactly the rates-row payload — the whole branch diff is +12/−2, the two removed lines being the caption lead lines the AMD captions extend. Original discussion: #319.

---

Adds one **IQ3_S (AMD RX 7900 XTX, gfx1100)** row to the README/DETAILS rates tables, per the maintainer's request to split the rates row out of #228 and re-measure.

**What it adds** (docs only, current head):
- README rates table: `| **IQ3_S** (AMD RX 7900 XTX) | 60 tokens/s | 1,641 tokens/s |`
- DETAILS prompt table (1K/4K/32K/64K/128K/262K): `760 | 1,275 | 1,641 | 1,594 | 1,494 | -`
- DETAILS output (decode) table: `65.1 | 60.4 | 64.4 | 61.8 | 59.2 | -`

**Provenance (updated for the current head).** Engine 0.1.38 on a ROCm 10.2 nightly with the tuned gfx1100 hipBLASLt-100500 table (upstream as #755), median of 3 non-stalled cells per tier from a 1K-128K one-shot sweep (256 generated tokens, greedy; Ryzen 9 7900X, 96 GB). Default configuration — **no** WMMA opt-ins (see the thread). 262K stays `-`, same as the existing IQ3_S/IQ3_XXS rows. The flat decode is expected: GDN linear attention makes decode O(1) in context. On-device kernel-shape autotuning (#744) was run on this machine and kept every shipped shape — the row's kernels are the shipped defaults.

Stalls were recorded and excluded from the medians. Docs-only change; no code. (The first revision of this description carried the 0.1.31 packaged-ROCm numbers; the commits above supersede them.)

*Prepared with an AI engineering agent (GLM-5.3 / Z.ai) under human direction; every number is from our own recorded runs.*

Mehr auf der Site

Links zu Install, Modellen, Releases.