Pull requests / #404
#404 bench(gfx1100): 0.1.31 vs 0.1.30 prefill/decode, curated entry + raw trials
closed · @rwkeyes · 0 Kommentare · Auf GitHub
BenchmarksAMD / HIPModels & quantsDocumentation
Beschreibung
Adds a community benchmark entry in the docs/benchmarks/ shape, plus the raw bench_prefill.py trials under bench/results/2026-10-01-0.1.31-release/. Platform: Ryzen 7 7700X, one RX 7900 XTX 24 GB (gfx1100), ROCm 7.1, model on a rotational drive. Method: tools/hip/bench_prefill.py against an otherwise idle local server, one run per release binary, identical pack (coder-iq1_m), template and settings. Sequential same-config runs roughly an hour apart - not interleaved, so treat this as indicative of the throughput difference rather than a tightly controlled A/B. Candidate is the stock v0.1.31 release binary, sha256 17f60a01ff553593af813b2ded8b288179ee9285d8a270a10b90f6bb63e5de1d, source commit 9259cad (tag v0.1.31). Token cap 128; throughput only, not a coding-quality benchmark. Fresh-prefill trials (over 512 fresh tokens): 0.1.30 prefill 225/545/700 tps min/med/max, decode 16.0-34.2 tps; 0.1.31 prefill 260/900/994 tps, decode 16.7-64.7 tps. Matches the release notes' direction: both warm prefill and warm decode improved. The spread between trials is dominated by first-touch reads from the rotational drive rather than by the engine.
Mehr auf der Site
Links zu Install, Modellen, Releases.