Pull requests / #1406

#1406 bench: community report, RX 7900 XTX (gfx1100) on WSL2, Coder IQ1_M: 0.1.33 vs 0.1.40, prefill streaming, 16k-256k sweep

open · @theGiallo · 0 评论 · 在 GitHub 查看

BenchmarksAMD / HIPModels & quantsDocumentationWindowsLinux

描述

Issue: n/a

## Summary
Results only, no engine changes. Adds one community benchmark report for an RX 7900 XTX 24 GB (gfx1100, clock capped at 2850 MHz), Ryzen 9 5900X, 64 GB DDR4 (54 GB visible to WSL2), Windows 11 + Ubuntu 26.04 in WSL2, ROCm 7.14. Model: Qwen3.8-Flash-Next GSQ-RCO IQ1_M Coder pack + MTP, `--mmap-experts`, `--kv int8`.

## What changed
Only `bench/results/2026-10-06-community-rx7900xtx-wsl2/` is added: README following docs/COMMUNITY_BENCHMARKS.md, per-run JSONL/CSV, engine logs, RSS samples, needle results, and the scripts that produced them.

Measured:
- 0.1.33 vs 0.1.40 with the community script (short / ~10k / follow-up, 3-9 runs each) and needle at 32k and 128k.
- Cold ~10k prompt prefill is 66.7 tok/s on 0.1.40 defaults vs 442.9 on 0.1.33; STRATA_PREFILL_STREAM_MIN=65536 gives 455.9. Env A/B in runs/envab/.
- Prompt-length sweep 16k-256k (-c 262144): prefill flat from ~60k at ~96-104 tok/s (0.1.33), 85-93 (0.1.40 defaults), 499-530 (0.1.40 with STREAM_MIN=65536). Decode ~40-60 tok/s to 261k.
- Community measurements repeated at DDR4-3200 (XMP): little change.

## Extra Notes
Limitations (also in the README): one machine, model and pack; sweep prompts are source code, one text per length; page-cache state and disk read rate not recorded in the sweep (compute vs file cache is open); the step in 0.1.33 prefill between 17k and 32k tokens was not investigated; long-context quality checked only with needle (to 128k).
Branch is two commits on current main; nothing outside the results folder is touched.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。