Pull requests / #1509

#1509 bench: community report, 2x RTX 2080 Ti 22 GB at full 250 W, IQ3_XXS (vs #1225 at 100 W)

open · @caolonghao · 0 commentaires · Sur GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentation

Description

Results-only PR: one report folder, `bench/results/2026-10-08-community-2x-2080ti-22gb/`. Companion to #1508 (the same PC's 2x V100-PCIE pair, idle during these runs).

The **second** 2x 2080 Ti data point after #1225 and the **first at full power limits** - #1225's 22 GB mods ran a 100 W cap; these ran 250 W (max 280 W). Same quant (GSQ-RCO IQ3_XXS at setup's pinned revision), same harness unmodified, same toolchain (engine 0.1.40.2 source-built for sm_75 with CUDA 12.8; #1225 was 0.1.40 at `82f46a8`), auto split K=26 (caches 20,209/24,576 pairs, ~99.5%):

- 4K / 32K fresh prompts: **981.3 / 1,684.4 tok/s prompt, 84.1 / 75.1 tok/s decode** (3 runs each, ranges in the README).
- vs #1225 (changed settings listed in the README: power cap, host, engine version, KV streaming): 4K decode **+59%**, 32K prompt **+253%** - the power cap is the obvious first suspect.
- On this PC, also directly comparable with the V100 pair's UD-IQ4_XS run: reads 4K prompts 23% / 32K prompts 28% slower than the V100 pair on the larger 4-bit model, decodes 9% faster at 4K and 6% slower at 32K on the smaller 3-bit one.

Limitations in the README (no needle/correctness runs, one machine, shared storage, no memory sampling). Raw dumps and logs pre-trimmed per the repository's practice; COMMUNITY.md index rows left to the maintainers.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Sur le site

Liens install, modèles, releases.