Pull requests / #1376

#1376 bench: add RX 7900 XT results for IQ3_S and IQ3_XXS

open · @matrix9neonebuchadnezzar2199-sketch · 0 comentarios · En GitHub

BenchmarksAMD / HIPModels & quantsWindows

Descripción

Same Windows machine and engine, so the two sizes can be compared without an engine change.

## Hardware
RX 7900 XT (20464 MiB, gfx1100), Ryzen 7 9800X3D, 95.6 GiB RAM, Windows 11 Pro build 26300. Ready-made HIP engine 0.1.40.2 at e8ca9af. Model files are on a separate NVMe.

## Model
Qwen3.8-Flash-Next GSQ-RCO, IQ3_S and IQ3_XXS, one loaded at a time. Vision off.

## Configurations
Three fresh runs each of a ~100-token prompt and a 38,184-token synthetic log, with max_tokens set to 256. Rates come from last_timings, not from dividing generated tokens by the whole request. cache_n was 0 on every run.

- IQ3_S: context 131072, KV q4_0, expert cache auto (5512 slots, 10.48 GiB). Long prompt 1015.8 tok/s median (1015.0–1016.3), decode 24.3 tok/s median (21.4–25.4).
- IQ3_XXS: context 262144, KV int8, expert cache auto (6376 slots, 10.31 GiB). Long prompt 1400.1 tok/s median (1399.7–1400.3), decode 72.3 tok/s median (59.4–73.1).

Report: `bench/results/2026-10-07-community-rx-7900-xt/`

## Limitations
One machine. No TTFT, no needle sweep, and no VRAM sample during generation. The short-prompt tok/s figure is overhead, not prefill speed. On IQ3_S the counted long-prompt decode (24.3 tok/s) was slower than its warm-up (51.7); the cause was not isolated. On IQ3_XXS one long run hit the 256-token cap and the visible answer was cut off at `amber-keel-2`. This does not replace the existing RX 7900 XTX report; that machine used different settings.

En el sitio

Enlaces a install, modelos, releases.