Pull requests / #1243

#1243 Bench/2026 10 05 community rtx 5080

closed · @KalAbaddon · 0 コメント · GitHub で見る

BenchmarksNVIDIA / CUDAModels & quantsWindows

本文

Community benchmark: RTX 5080 (PCIe 5.0 x8) + Ryzen 9 9900X3D, 96 GiB, Windows 11, engine 0.1.39

Results-only submission in bench/results/2026-10-05-community-rtx5080-9900x3d/.

Configurations: IQ3_S, Unsloth UD-IQ4_XS (before and after --calibrate), Unsloth UD-Q4_K_XL, and Coder IQ1_M; 262,144 context.
Speed: 3 runs each at 1k, 4k, 32k, 64k, 128k and 262k prompt tokens, 256-token cap, greedy, reasoning off. Per-run JSON and engine log lines included.
Needles: 9 cases each on four configurations; all found 9 of 9.
Headline: IQ3_S was the fastest full model (2,344–3,603 prompt tok/s from 4k up, 80–98 decode). Calibration raised UD-IQ4_XS decode by 13–44%.
Limitations: PCIe x8 rather than x16; tight RAM with the Unsloth models; many answers stopped before the 256-token cap; one calibrated 128k run failed with an engine error and is excluded.

関連リンク

インストール・モデル・リリースへの站内リンク。