Pull requests / #418

#418 bench: community entry, 2x RTX PRO 4500, Swift IQ3_XXS, engine 0.1.30 and 0.1.36

closed · @qni-live · 0 评论 · 在 GitHub 查看

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentation

描述

Community benchmark for the Swift 1.5 IQ3_XXS pack on two RTX PRO 4500 Blackwell (32 GB each, SM 12.0),
Threadripper 7960X, 64 GB DDR5, engine 0.1.30, --spec 5, --layer-split 25, --kv k8v4.

- 1K / 4K / 32K / 128K prompts, 3 runs each, 256 output tokens, greedy, no prompt reuse
- Output: 124 / 115 / 120 / 93 tok/s; prompt: 903 / 1428 / 2454 / 2785 tok/s
- Needle recall at 32K and 128K, depths 10/50/90: 6 of 6 found
- A --spec 3/4/5/6 sweep: differences are small (3-5% run spread); 5 was slightly ahead

Not the Flash-Next IQ2_XS pack of the first entry, and two GPUs, so not directly comparable to single-GPU entries.
Single stream only. Cold start not measured. Details, scripts, raw runs and settings.json are in the folder README.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。