Pull requests / #404

#404 bench(gfx1100): 0.1.31 vs 0.1.30 prefill/decode, curated entry + raw trials

closed · @rwkeyes · 0 comentários · No GitHub

BenchmarksAMD / HIPModels & quantsDocumentation

Descrição

Adds a community benchmark entry in the docs/benchmarks/ shape, plus the raw bench_prefill.py trials under bench/results/2026-10-01-0.1.31-release/.

Platform: Ryzen 7 7700X, one RX 7900 XTX 24 GB (gfx1100), ROCm 7.1, model on a rotational drive.

Method: tools/hip/bench_prefill.py against an otherwise idle local server, one run per release binary, identical pack (coder-iq1_m), template and settings. Sequential same-config runs roughly an hour apart - not interleaved, so treat this as indicative of the throughput difference rather than a tightly controlled A/B. Candidate is the stock v0.1.31 release binary, sha256 17f60a01ff553593af813b2ded8b288179ee9285d8a270a10b90f6bb63e5de1d, source commit 9259cad (tag v0.1.31). Token cap 128; throughput only, not a coding-quality benchmark.

Fresh-prefill trials (over 512 fresh tokens): 0.1.30 prefill 225/545/700 tps min/med/max, decode 16.0-34.2 tps; 0.1.31 prefill 260/900/994 tps, decode 16.7-64.7 tps.

Matches the release notes' direction: both warm prefill and warm decode improved. The spread between trials is dominated by first-touch reads from the rotational drive rather than by the engine.

No site

Links install, modelos, releases.