Pull requests / #1406
#1406 bench: community report, RX 7900 XTX (gfx1100) on WSL2, Coder IQ1_M: 0.1.33 vs 0.1.40, prefill streaming, 16k-256k sweep
open · @theGiallo · 0 comentários · No GitHub
BenchmarksAMD / HIPModels & quantsDocumentationWindowsLinux
Descrição
Issue: n/a ## Summary Results only, no engine changes. Adds one community benchmark report for an RX 7900 XTX 24 GB (gfx1100, clock capped at 2850 MHz), Ryzen 9 5900X, 64 GB DDR4 (54 GB visible to WSL2), Windows 11 + Ubuntu 26.04 in WSL2, ROCm 7.14. Model: Qwen3.8-Flash-Next GSQ-RCO IQ1_M Coder pack + MTP, `--mmap-experts`, `--kv int8`. ## What changed Only `bench/results/2026-10-06-community-rx7900xtx-wsl2/` is added: README following docs/COMMUNITY_BENCHMARKS.md, per-run JSONL/CSV, engine logs, RSS samples, needle results, and the scripts that produced them. Measured: - 0.1.33 vs 0.1.40 with the community script (short / ~10k / follow-up, 3-9 runs each) and needle at 32k and 128k. - Cold ~10k prompt prefill is 66.7 tok/s on 0.1.40 defaults vs 442.9 on 0.1.33; STRATA_PREFILL_STREAM_MIN=65536 gives 455.9. Env A/B in runs/envab/. - Prompt-length sweep 16k-256k (-c 262144): prefill flat from ~60k at ~96-104 tok/s (0.1.33), 85-93 (0.1.40 defaults), 499-530 (0.1.40 with STREAM_MIN=65536). Decode ~40-60 tok/s to 261k. - Community measurements repeated at DDR4-3200 (XMP): little change. ## Extra Notes Limitations (also in the README): one machine, model and pack; sweep prompts are source code, one text per length; page-cache state and disk read rate not recorded in the sweep (compute vs file cache is open); the step in 0.1.33 prefill between 17k and 32k tokens was not investigated; long-context quality checked only with needle (to 128k). Branch is two commits on current main; nothing outside the results folder is touched.
No site
Links install, modelos, releases.