Pull requests / #1245

#1245 bench: UD-IQ4_XS first NVIDIA measurement (RTX 5090 Laptop, 24 GB)

closed · @ManfredCh · 0 comments · View on GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentation

Description

Unsloth UD-IQ4_XS was marked 'not measured on NVIDIA yet' in docs/UNSLOTH_Q4.md — here is a first data point.

**Setup:** RTX 5090 Laptop 24 GB (PCIe Gen5 x16, driver 595.91.07), Core Ultra 9 275HX (8P+16E), 188 GiB RAM, Ubuntu 24.04, engine 0.1.40, locally compiled sm_120. Context 262K, INT8 KV, resident-budget 55 GiB (all non-VRAM experts pinned), CJK draft vocab.

**Headline:** 22–33 tok/s decode (greedy, cold→warm) vs 53–59 tok/s for IQ3_S on the same box — the ~4-bit pack is ~2x slower here, and STRATA_DECODE_TIMING shows the GPU itself is the bottleneck (Q8_0 dense side + bigger expert blobs), not the CPU pool / PCIe / RAM residency.

**Actionable findings in the report:**
1. --calibrate aborts mid-sweep on this config: the MTP draft head (212.9 MiB, CJK vocab) no longer fits during the 23-worker restart (213–274 MiB free), so the whole tuning run keeps defaults. Sweep winners are printed before the failure (PCIe 0.55, floor 0.70); applied by hand. Maybe treat a non-fitting restart as 'setting loses' instead of failing the run.
2. STRATA_EXCHANGE_ROTATE=1 unavailable for this pack (unequal-size expert blocks: IQ3_S/IQ4_NL mix).
3. 213 MiB VRAM free with everything loaded at 262K + CJK draft head (engine warns of stalls).
4. Adaptive swap counts grow during use; decode warms 21.6 → 33.0 tok/s over consecutive requests.

All three shards SHA-256-verified. Details, raw bench lines and the full config in bench/results/2026-10-07-community-rtx-5090-laptop-ud-iq4xs/.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.