贡献 / #1617

#1617 bench: RTX 5090 community report, unsloth UD-Q4_K_XL and UD-Q5_K_XL with images on 0.1.41

open · @CYoung83 · 0 评论 · 去 GitHub 看

BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentation

说明

## Title
Issue: none (results only). Related: #967, #1611 / #1612

## Summary
A community benchmark report from an RTX 5090 with a Ryzen 9 9950X3D and 89 GiB of RAM, on engine 0.1.41 with #1612. It covers unsloth UD-Q4_K_XL and UD-Q5_K_XL with and without images at a 48 GiB RAM budget:

- speed at 4,096 and 32,768 prompt tokens, 3 runs each;
- KL against unsloth Q8_0 on the same engine, over 30 public-text requests;
- a three-question image check on each quant. This is the sample #967 asked for.

Decode medians, 4K / 32K, in tok/s:

| Quant | Without images | With images |
| --- | --- | --- |
| UD-Q4_K_XL | 109.7 / 111.8 | 101.2 / 104.1 |
| UD-Q5_K_XL | 60.6 / 64.4 | 51.8 / 60.9 |

Mean KL against Q8_0: UD-Q4_K_XL 0.0558, UD-Q5_K_XL 0.0346. The paired difference is +0.0212 [+0.0122, +0.0307]. For scale, the same Q8_0 weights on 0.1.40.3 against 0.1.41 differ by 0.0224. Image check: 3/3 on both quants.

## What changed
- bench/results/2026-10-08-community-rtx-5090-ud-q4-q5-0.1.41/ (new):
  - README.md, laid out like the template in docs/COMMUNITY_BENCHMARKS.md;
  - per-run speed data, configs, engine logs and GPU telemetry;
  - the quality report, its summary and the 30 requests (gzipped, 0.4 MB);
  - the image check's picture and answers;
  - the scripts that produced them, as run;
  - build options, engine commits and model hashes.
- docs/COMMUNITY_BENCHMARKS.md: one entry in the community reports list.

Nothing outside those two paths changes. No engine or setup change; the loader fix is #1612.

## Extra Notes
- UD-Q5_K_XL needs #1612, or a copy of shard 1 padded to 10,946,624 bytes. The build used -DSTRATA_Q6K_EXPERTS=ON because the Q5 pack lists Q6_K gate/up tensors for some layers.
- Perplexity doesn't order the two quants the way KL does (overall 52.81 for Q4, 54.34 for Q5, 51.87 for Q8_0). The README gives both as measured.
- Not measured: 128K prompts on 0.1.41, RAM budgets other than 48 GiB, task benchmarks, and harder image questions. The README lists the limitations.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01NWf9g8Qf9dTawR22quzAX3

本站相关内容

相关页面的快捷入口。