Pull requests / #1321

#1321 bench: publish an IQ2_XS speed report for the Radeon AI PRO R9700

open · @alexhegit · 0 commentaires · Sur GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Description

## Summary
- Community IQ2_XS speed report for an AMD Radeon AI PRO R9700 (gfx1201), Ryzen AI 9 HX 370, 47 GB RAM, Ubuntu 24.04, ROCm 7.14 TheRock wheels (hipBLASLt 1.4.1, no gfx1201 table).
- Three runs at 4,096 and 32,768 prompt tokens with the default prompt kernel, then the same sweep with `STRATA_HIP_WMMA=1`.
- Does not change the Coder IQ1_M table in docs/AMD_HIP.md. 128,000-token prompts and the recall check were not run.

## Results
Default kernel, median of three runs, 256 tokens, temperature 0, reasoning off, 0 reused tokens:
- 4,096: prompt 1,270.3 tok/s, decode 94.1 tok/s, TTFT 3.256 s
- 32,768: prompt 1,574.1 tok/s, decode 86.1 tok/s, TTFT 20.875 s

`STRATA_HIP_WMMA=1`, same machine, engine restarted:
- 4,096: prompt 1,486.0 tok/s, decode 91.6 tok/s, TTFT 2.787 s
- 32,768: prompt 2,128.4 tok/s, decode 89.7 tok/s, TTFT 15.456 s

## Limitations
- The desktop is on this GPU, so 3,072 MiB of VRAM was reserved.
- Prompt GEMMs used plain hipBLAS.
- Matrix-core prompt attention is not bitwise-identical to the default kernel.
- These runs do not replace the RTX 5070 row, the RTX 5090 report, or the R9700 Coder table. Those used other engines, CPUs, and protocols.

## Test plan
- [ ] Confirm the report is linked from docs/COMMUNITY_BENCHMARKS.md
- [ ] Confirm the diff has no engine binary, model files, or packs
- [ ] Spot-check the medians against results.json and summary-wmma.json

Sur le site

Liens install, modèles, releases.