贡献 / #1726

#1726 bench: community report, RTX 5080 16 GB + Xeon w7-2475X + 128 GB, Unsloth UD-IQ4_XS at a 786K context (YaRN x3), engine 0.1.41

open · @enkynakamura · 0 评论 · 去 GitHub 看

BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows

说明

Results-only submission per docs/COMMUNITY_BENCHMARKS.md, no engine changes.

Hardware: one RTX 5080 16 GB, Xeon w7-2475X (20C, AVX-512), 128 GB DDR5-4800, Windows 11, source build of v0.1.41 (CUDA 13.3, sm_120).
Model: Unsloth UD-IQ4_XS, all experts resident in RAM (--resident-budget-gib 55), --max-context 786432 with --rope-scaling yarn --rope-scale 3, KV int8, MTP --spec 5. Greedy throughout.

Measured:
- standard sweep 4K / 32K / 128K, 3 runs each, 256-token cap: prompt 1,814 / 3,299 / 3,398 tok/s (medians), decode 68.6 / 70.2 / 64.1 tok/s
- needle (tools/needle_bench.py) 32K / 128K / 262K / 512K at depths 10/50/90: 12 of 12, the 512K rows beyond the native window through YaRN
- a 352K synthetic prompt (1 run): 3,111 tok/s prompt, 60.7 decode
- a 352K real agent conversation (private, not reproducible, reported for scale): 2,253 tok/s prompt as the first request after start, 2,635 with a warm expert cache, 56.3 tok/s decode (3 runs) on the reused prefix

One observation the standard sweep does not show: at the same length and cache state, the real conversation reads 15% slower than the repeated-paragraph synthetic prompt, and the decode expert cache hits 57-61% instead of 68% (engine log lines in the folder). Repeated-text prompts understate the prompt cost of real text.

Limitations: single machine and quantization; the real-conversation rows are one private prompt; memory peaks not measured; model revision and PCIe generation under load not recorded.

本站相关内容

相关页面的快捷入口。