Pull requests / #740

#740 Community benchmark: RX 7900 GRE (gfx1100, 16 GB), three Flash-Next quants

closed · draft · @jase100k · 0 Kommentare · Auf GitHub

BenchmarksAMD / HIPModels & quants

Beschreibung

Community benchmark report, no code change.

Measured on 2026-10-04 by @jase100k on an AMD Radeon RX 7900 GRE (gfx1100,
16 GB) with a Ryzen 7 5700X3D and 64 GB of RAM. Strata 0.1.38 (commit `99f3dbd`),
local source build, ROCm 7.2.3, hipBLASLt 1.2.2 with the repository gfx1100 tuning
table active.

Three arms whose settings differ only in the quantization: the Coder IQ1_M pack,
the original Flash-Next IQ2_XS and Flash-Next IQ3_XXS. Each arm ran
`tools/hip/bench_prefill.py` three times, with the engine restarted before every
run so that all 36 fresh prompts reused zero tokens, plus a
`tools/needle_bench.py` pass at 32k.

Median of six requests per cell, 8,830 fresh prompt tokens, 128-token greedy
output:

| Quant | Prompt tok/s | Decode tok/s | Expert cache slots | Peak VRAM |
| --- | ---: | ---: | ---: | ---: |
| Coder IQ1_M | 1,213.2 | 54.7 | 3,783 | 14.6 GiB |
| Flash-Next IQ2_XS | 1,154.7 | 65.3 | 5,959 | 14.6 GiB |
| Flash-Next IQ3_XXS | 1,140.1 | 59.1 | 4,608-4,621 | 14.6 GiB |

`needle_bench.py` found 3 of 3 needles at depths 10/50/90 for all three arms.
Peak host RAM used was 32.4 / 41.1 / 48.2 GiB in the same order.

Limitations: one machine and one card, with no cross-engine or cross-GPU
baseline; the Coder arm is a different fine-tune with a different expert profile
and cache sizing, so only IQ2_XS against IQ3_XXS isolates the quantization; the
context limit is 65,536 for all three arms; outputs are 128 tokens, greedy, with
reasoning off; time to first token was not measured because `bench_prefill.py`
uses the non-streaming endpoint; the desktop shares the measured GPU; the host
expert arena ran on 4 KB pages because `vm.nr_hugepages` is 0.

The report carries the full hardware and software list, the exact configs, the
method, per-request JSON for all nine runs, the engine logs, the 1 Hz memory
samples, computed aggregates and SHA-256 sums of the model and pack files.

Mehr auf der Site

Links zu Install, Modellen, Releases.