Pull requests / #740
#740 Community benchmark: RX 7900 GRE (gfx1100, 16 GB), three Flash-Next quants
closed · draft · @jase100k · 0 comments · View on GitHub
BenchmarksAMD / HIPModels & quants
Description
Community benchmark report, no code change. Measured on 2026-10-04 by @jase100k on an AMD Radeon RX 7900 GRE (gfx1100, 16 GB) with a Ryzen 7 5700X3D and 64 GB of RAM. Strata 0.1.38 (commit `99f3dbd`), local source build, ROCm 7.2.3, hipBLASLt 1.2.2 with the repository gfx1100 tuning table active. Three arms whose settings differ only in the quantization: the Coder IQ1_M pack, the original Flash-Next IQ2_XS and Flash-Next IQ3_XXS. Each arm ran `tools/hip/bench_prefill.py` three times, with the engine restarted before every run so that all 36 fresh prompts reused zero tokens, plus a `tools/needle_bench.py` pass at 32k. Median of six requests per cell, 8,830 fresh prompt tokens, 128-token greedy output: | Quant | Prompt tok/s | Decode tok/s | Expert cache slots | Peak VRAM | | --- | ---: | ---: | ---: | ---: | | Coder IQ1_M | 1,213.2 | 54.7 | 3,783 | 14.6 GiB | | Flash-Next IQ2_XS | 1,154.7 | 65.3 | 5,959 | 14.6 GiB | | Flash-Next IQ3_XXS | 1,140.1 | 59.1 | 4,608-4,621 | 14.6 GiB | `needle_bench.py` found 3 of 3 needles at depths 10/50/90 for all three arms. Peak host RAM used was 32.4 / 41.1 / 48.2 GiB in the same order. Limitations: one machine and one card, with no cross-engine or cross-GPU baseline; the Coder arm is a different fine-tune with a different expert profile and cache sizing, so only IQ2_XS against IQ3_XXS isolates the quantization; the context limit is 65,536 for all three arms; outputs are 128 tokens, greedy, with reasoning off; time to first token was not measured because `bench_prefill.py` uses the non-streaming endpoint; the desktop shares the measured GPU; the host expert arena ran on 4 KB pages because `vm.nr_hugepages` is 0. The report carries the full hardware and software list, the exact configs, the method, per-request JSON for all nine runs, the engine logs, the 1 Hz memory samples, computed aggregates and SHA-256 sums of the model and pack files.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.