Pull requests / #1509
#1509 bench: community report, 2x RTX 2080 Ti 22 GB at full 250 W, IQ3_XXS (vs #1225 at 100 W)
open · @caolonghao · 0 Kommentare · Auf GitHub
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentation
Beschreibung
Results-only PR: one report folder, `bench/results/2026-10-08-community-2x-2080ti-22gb/`. Companion to #1508 (the same PC's 2x V100-PCIE pair, idle during these runs). The **second** 2x 2080 Ti data point after #1225 and the **first at full power limits** - #1225's 22 GB mods ran a 100 W cap; these ran 250 W (max 280 W). Same quant (GSQ-RCO IQ3_XXS at setup's pinned revision), same harness unmodified, same toolchain (engine 0.1.40.2 source-built for sm_75 with CUDA 12.8; #1225 was 0.1.40 at `82f46a8`), auto split K=26 (caches 20,209/24,576 pairs, ~99.5%): - 4K / 32K fresh prompts: **981.3 / 1,684.4 tok/s prompt, 84.1 / 75.1 tok/s decode** (3 runs each, ranges in the README). - vs #1225 (changed settings listed in the README: power cap, host, engine version, KV streaming): 4K decode **+59%**, 32K prompt **+253%** - the power cap is the obvious first suspect. - On this PC, also directly comparable with the V100 pair's UD-IQ4_XS run: reads 4K prompts 23% / 32K prompts 28% slower than the V100 pair on the larger 4-bit model, decodes 9% faster at 4K and 6% slower at 32K on the smaller 3-bit one. Limitations in the README (no needle/correctness runs, one machine, shared storage, no memory sampling). Raw dumps and logs pre-trimmed per the repository's practice; COMMUNITY.md index rows left to the maintainers. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Mehr auf der Site
Links zu Install, Modellen, Releases.