Pull requests / #1462

#1462 bench: RTX 5090 Laptop GPU follow-up for 0.1.40.2 and 0.1.40.3

open · @wolffahrer · 0 comentários · No GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentation

Descrição

Follow-up to #1263 (now in `bench/results/2026-10-06-community-rtx-5090-laptop`): same laptop, model files, configuration and scripts, only the Strata version changed. Results only, no engine changes.

New folder: `bench/results/2026-10-08-community-rtx-5090-laptop-engine-0.1.40.3/`

**Median decode / prompt tok/s (Q2_0, 131K context, greedy, reasoning off, 256-token cap):**

| Prompt tokens | 0.1.40.1 | 0.1.40.2 | 0.1.40.3 |
| ---: | ---: | ---: | ---: |
| 4,096 | 136.9 / 1,952.5 | 139.3 / 1,960.0 | 138.1 / 1,949.7 |
| 32,768 | 137.7 / 2,877.8 | 141.6 / 2,890.0 | 144.4 / 2,893.3 |
| 128,000 | 134.6 / 2,764.1 | 137.2 / 2,764.1 | 131.0 / 2,780.1 |

- No measurable speed change on this hardware and no regression; differences are inside single-run min–max ranges.
- Needles 6/6 on both versions. Coding check 10/10 on 0.1.40.3.
- The 0.1.40.2 coding check scored 0/10. Re-runs with raw responses show the same greedy reasoning loop in 0.1.40.1 (first request after a fresh start, identical text in both versions, ended by `reasoning_loop_recovery: "stop"`), so we don't think it is a 0.1.40.2 regression. Details are in the README.
- Raw data is included for 0.1.40.3 only. `telemetry.jsonl` is left out, `memory-summary.json` is kept.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

No site

Links install, modelos, releases.