Pull requests / #319

#319 Docs: AMD RX 7900 XTX (gfx1100) IQ3_S rates row (1K-128K)

closed · @xyzzing · 0 comentários · No GitHub

BenchmarksAMD / HIPModels & quantsDocumentation

Descrição

Adds one **IQ3_S (AMD RX 7900 XTX, gfx1100)** row to the README/DETAILS rates tables, per the maintainer's request to split the rates row out of #228 and re-measure.

**What it adds** (docs only, current head):
- README rates table: `| **IQ3_S** (AMD RX 7900 XTX) | 60 tokens/s | 1,641 tokens/s |`
- DETAILS prompt table (1K/4K/32K/64K/128K/262K): `760 | 1,275 | 1,641 | 1,594 | 1,494 | -`
- DETAILS output (decode) table: `65.1 | 60.4 | 64.4 | 61.8 | 59.2 | -`

**Provenance (updated for the current head).** Engine 0.1.38 on a ROCm 10.2 nightly with the tuned gfx1100 hipBLASLt-100500 table (upstream as #755), median of 3 non-stalled cells per tier from a 1K-128K one-shot sweep (256 generated tokens, greedy; Ryzen 9 7900X, 96 GB). Default configuration — **no** WMMA opt-ins (see the thread). 262K stays `-`, same as the existing IQ3_S/IQ3_XXS rows. The flat decode is expected: GDN linear attention makes decode O(1) in context. On-device kernel-shape autotuning (#744) was run on this machine and kept every shipped shape — the row's kernels are the shipped defaults.

Stalls were recorded and excluded from the medians. Docs-only change; no code. (The first revision of this description carried the 0.1.31 packaged-ROCm numbers; the commits above supersede them.)

No site

Links install, modelos, releases.