Pull requests / #1608

#1608 bench: RTX 5090 Laptop GPU follow-up for 0.1.41

open · @wolffahrer · 0 comments · View on GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows

Description

Follow-up to #1263 and #1462: same laptop (RTX 5090 Laptop GPU 24 GB, Windows 11), model files, configuration and scripts. Only the Strata version changed. Results only, no engine changes.

New folder: `bench/results/2026-10-08-community-rtx-5090-laptop-engine-0.1.41/`

**Median decode / prompt tok/s (Q2_0, 131K context, greedy, reasoning off, 256-token cap):**

| Prompt tokens | 0.1.40.3 | 0.1.41 |
| ---: | ---: | ---: |
| 4,096 | 138.1 / 1,949.7 | 142.5 / 1,949.8 |
| 32,768 | 144.4 / 2,893.3 | 141.5 / 2,888.3 |
| 128,000 | 131.0 / 2,780.1 | 127.8 / 2,776.8 |

- No measurable change on this machine: single GPU and 64 GB RAM, so the multi-GPU and low-RAM gains don't apply here. The installer suggested the same arguments as for 0.1.40.3.
- Needles 6/6, coding check 10/10. No regression on our private agent task suite (124/132 runs, best of the four versions).
- Peak GPU memory 23,565 MiB vs 23,221 MiB on 0.1.40.3, and the expert cache sized 3 slots smaller (11,739 vs 11,742). Both are small, mentioned in case it is unexpected.
- The README links to the 0.1.40.3 folder from #1462. If you merge this one first, that link stays dead until #1462 lands.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.