Pull requests / #1617
#1617 bench: RTX 5090 community report, unsloth UD-Q4_K_XL and UD-Q5_K_XL with images on 0.1.41
open · @CYoung83 · 0 comentários · No GitHub
BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentation
Descrição
## Title Issue: none (results only). Related: #967, #1611 / #1612 ## Summary A community benchmark report from an RTX 5090 with a Ryzen 9 9950X3D and 89 GiB of RAM, on engine 0.1.41 with #1612. It covers unsloth UD-Q4_K_XL and UD-Q5_K_XL with and without images at a 48 GiB RAM budget: - speed at 4,096 and 32,768 prompt tokens, 3 runs each; - KL against unsloth Q8_0 on the same engine, over 30 public-text requests; - a three-question image check on each quant. This is the sample #967 asked for. Decode medians, 4K / 32K, in tok/s: | Quant | Without images | With images | | --- | --- | --- | | UD-Q4_K_XL | 109.7 / 111.8 | 101.2 / 104.1 | | UD-Q5_K_XL | 60.6 / 64.4 | 51.8 / 60.9 | Mean KL against Q8_0: UD-Q4_K_XL 0.0558, UD-Q5_K_XL 0.0346. The paired difference is +0.0212 [+0.0122, +0.0307]. For scale, the same Q8_0 weights on 0.1.40.3 against 0.1.41 differ by 0.0224. Image check: 3/3 on both quants. ## What changed - bench/results/2026-10-08-community-rtx-5090-ud-q4-q5-0.1.41/ (new): - README.md, laid out like the template in docs/COMMUNITY_BENCHMARKS.md; - per-run speed data, configs, engine logs and GPU telemetry; - the quality report, its summary and the 30 requests (gzipped, 0.4 MB); - the image check's picture and answers; - the scripts that produced them, as run; - build options, engine commits and model hashes. - docs/COMMUNITY_BENCHMARKS.md: one entry in the community reports list. Nothing outside those two paths changes. No engine or setup change; the loader fix is #1612. ## Extra Notes - UD-Q5_K_XL needs #1612, or a copy of shard 1 padded to 10,946,624 bytes. The build used -DSTRATA_Q6K_EXPERTS=ON because the Q5 pack lists Q6_K gate/up tensors for some layers. - Perplexity doesn't order the two quants the way KL does (overall 52.81 for Q4, 54.34 for Q5, 51.87 for Q8_0). The README gives both as measured. - Not measured: 128K prompts on 0.1.41, RAM budgets other than 48 GiB, task benchmarks, and harder image questions. The README lists the limitations. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01NWf9g8Qf9dTawR22quzAX3
No site
Links install, modelos, releases.