Pull requests / #1433

#1433 community bench: RTX 5090 + RTX PRO 4000 Blackwell, UD-Q4_K_XL helper cache, engine 0.1.40.3

open · @sophieandreikina · 0 comments · View on GitHub

BenchmarksNVIDIA / CUDA

Description

Helper-cache config (all 48 layers on RTX 5090, RTX PRO 4000 as 5000-expert cache), UD-Q4_K_XL native pack, engine 0.1.40.3. Standard ladder 4k/32k/128k, 1 warm-up + 3 measured, temp 0, 256 out. Medians: prompt 1513.7 / 2048.0 / 2200.6 tok/s, decode 103.3 / 92.2 / 93.9 tok/s. Background vLLM embedding server holds ~12 GiB of GPU 1 (caps cache at 5000 slots); results-only, no engine changes.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.