Pull requests / #1376
#1376 bench: add RX 7900 XT results for IQ3_S and IQ3_XXS
open · @matrix9neonebuchadnezzar2199-sketch · 0 comentários · No GitHub
BenchmarksAMD / HIPModels & quantsWindows
Descrição
Same Windows machine and engine, so the two sizes can be compared without an engine change. ## Hardware RX 7900 XT (20464 MiB, gfx1100), Ryzen 7 9800X3D, 95.6 GiB RAM, Windows 11 Pro build 26300. Ready-made HIP engine 0.1.40.2 at e8ca9af. Model files are on a separate NVMe. ## Model Qwen3.8-Flash-Next GSQ-RCO, IQ3_S and IQ3_XXS, one loaded at a time. Vision off. ## Configurations Three fresh runs each of a ~100-token prompt and a 38,184-token synthetic log, with max_tokens set to 256. Rates come from last_timings, not from dividing generated tokens by the whole request. cache_n was 0 on every run. - IQ3_S: context 131072, KV q4_0, expert cache auto (5512 slots, 10.48 GiB). Long prompt 1015.8 tok/s median (1015.0–1016.3), decode 24.3 tok/s median (21.4–25.4). - IQ3_XXS: context 262144, KV int8, expert cache auto (6376 slots, 10.31 GiB). Long prompt 1400.1 tok/s median (1399.7–1400.3), decode 72.3 tok/s median (59.4–73.1). Report: `bench/results/2026-10-07-community-rx-7900-xt/` ## Limitations One machine. No TTFT, no needle sweep, and no VRAM sample during generation. The short-prompt tok/s figure is overhead, not prefill speed. On IQ3_S the counted long-prompt decode (24.3 tok/s) was slower than its warm-up (51.7); the cause was not isolated. On IQ3_XXS one long run hit the 256-token cap and the visible answer was cut off at `amber-keel-2`. This does not replace the existing RX 7900 XTX report; that machine used different settings.
No site
Links install, modelos, releases.