Pull requests / #1321
#1321 bench: publish an IQ2_XS speed report for the Radeon AI PRO R9700
open · @alexhegit · 0 comentarios · En GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Descripción
## Summary - Community IQ2_XS speed report for an AMD Radeon AI PRO R9700 (gfx1201), Ryzen AI 9 HX 370, 47 GB RAM, Ubuntu 24.04, ROCm 7.14 TheRock wheels (hipBLASLt 1.4.1, no gfx1201 table). - Three runs at 4,096 and 32,768 prompt tokens with the default prompt kernel, then the same sweep with `STRATA_HIP_WMMA=1`. - Does not change the Coder IQ1_M table in docs/AMD_HIP.md. 128,000-token prompts and the recall check were not run. ## Results Default kernel, median of three runs, 256 tokens, temperature 0, reasoning off, 0 reused tokens: - 4,096: prompt 1,270.3 tok/s, decode 94.1 tok/s, TTFT 3.256 s - 32,768: prompt 1,574.1 tok/s, decode 86.1 tok/s, TTFT 20.875 s `STRATA_HIP_WMMA=1`, same machine, engine restarted: - 4,096: prompt 1,486.0 tok/s, decode 91.6 tok/s, TTFT 2.787 s - 32,768: prompt 2,128.4 tok/s, decode 89.7 tok/s, TTFT 15.456 s ## Limitations - The desktop is on this GPU, so 3,072 MiB of VRAM was reserved. - Prompt GEMMs used plain hipBLAS. - Matrix-core prompt attention is not bitwise-identical to the default kernel. - These runs do not replace the RTX 5070 row, the RTX 5090 report, or the R9700 Coder table. Those used other engines, CPUs, and protocols. ## Test plan - [ ] Confirm the report is linked from docs/COMMUNITY_BENCHMARKS.md - [ ] Confirm the diff has no engine binary, model files, or packs - [ ] Spot-check the medians against results.json and summary-wmma.json
En el sitio
Enlaces a install, modelos, releases.