Pull requests / #1726
#1726 bench: community report, RTX 5080 16 GB + Xeon w7-2475X + 128 GB, Unsloth UD-IQ4_XS at a 786K context (YaRN x3), engine 0.1.41
open · @enkynakamura · 0 comentários · No GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows
Descrição
Results-only submission per docs/COMMUNITY_BENCHMARKS.md, no engine changes. Hardware: one RTX 5080 16 GB, Xeon w7-2475X (20C, AVX-512), 128 GB DDR5-4800, Windows 11, source build of v0.1.41 (CUDA 13.3, sm_120). Model: Unsloth UD-IQ4_XS, all experts resident in RAM (--resident-budget-gib 55), --max-context 786432 with --rope-scaling yarn --rope-scale 3, KV int8, MTP --spec 5. Greedy throughout. Measured: - standard sweep 4K / 32K / 128K, 3 runs each, 256-token cap: prompt 1,814 / 3,299 / 3,398 tok/s (medians), decode 68.6 / 70.2 / 64.1 tok/s - needle (tools/needle_bench.py) 32K / 128K / 262K / 512K at depths 10/50/90: 12 of 12, the 512K rows beyond the native window through YaRN - a 352K synthetic prompt (1 run): 3,111 tok/s prompt, 60.7 decode - a 352K real agent conversation (private, not reproducible, reported for scale): 2,253 tok/s prompt as the first request after start, 2,635 with a warm expert cache, 56.3 tok/s decode (3 runs) on the reused prefix One observation the standard sweep does not show: at the same length and cache state, the real conversation reads 15% slower than the repeated-paragraph synthetic prompt, and the decode expert cache hits 57-61% instead of 68% (engine log lines in the folder). Repeated-text prompts understate the prompt cost of real text. Limitations: single machine and quantization; the real-conversation rows are one private prompt; memory peaks not measured; model revision and PCIe generation under load not recorded.
No site
Links install, modelos, releases.