Pull requests / #1263
#1263 Community benchmark: RTX 5090 Laptop GPU (24 GB), Windows 11, Strata 0.1.40.1
closed · @wolffahrer · 0 评论 · 在 GitHub 查看
BenchmarksNVIDIA / CUDAModels & quantsWindows
描述
- **Hardware:** Schenker XMG NEO 16 (A25) laptop, RTX 5090 Laptop GPU 24 GB (175 W limit, PCIe Gen 4 x8), Ryzen 9 9955HX3D, 64 GB DDR5-5600, Samsung 990 PRO NVMe, Windows 11. - **Model:** Flash-Next Q2_0 (ISTA-DASLab GSQ-RCO), Strata 0.1.40.1 built from source (engine 0.1.40), 131,072 context, INT8 KV. - **Tested:** the workload of the 2026-09-30 desktop RTX 5090 report: 4,096 / 32,768 / 128,000 fresh prompt tokens × 3, 256-token cap, greedy, reasoning off. Decode medians 136.9 / 137.7 / 134.6 tok/s, prompt 1,952 / 2,878 / 2,764 tok/s; needle recall 6/6; small coding check 10/10. - **Limitations:** one laptop; Q2_0 vs the desktop's IQ2_XS plus different engine, OS, power limit and PCIe link, so the side-by-side is not a controlled comparison; no thermal soak, vision, tools or concurrency. - Results only, no engine changes. I'll walk through this run in a German video on the KI SOUVERÄN YouTube channel: https://www.youtube.com/@KISouver%C3%A4n (Telegram: https://t.me/lokale_ki). 🤖 Generated with [Claude Code](https://claude.com/claude-code)
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。