Pull requests / #1462
#1462 bench: RTX 5090 Laptop GPU follow-up for 0.1.40.2 and 0.1.40.3
open · @wolffahrer · 0 commentaires · Sur GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentation
Description
Follow-up to #1263 (now in `bench/results/2026-10-06-community-rtx-5090-laptop`): same laptop, model files, configuration and scripts, only the Strata version changed. Results only, no engine changes. New folder: `bench/results/2026-10-08-community-rtx-5090-laptop-engine-0.1.40.3/` **Median decode / prompt tok/s (Q2_0, 131K context, greedy, reasoning off, 256-token cap):** | Prompt tokens | 0.1.40.1 | 0.1.40.2 | 0.1.40.3 | | ---: | ---: | ---: | ---: | | 4,096 | 136.9 / 1,952.5 | 139.3 / 1,960.0 | 138.1 / 1,949.7 | | 32,768 | 137.7 / 2,877.8 | 141.6 / 2,890.0 | 144.4 / 2,893.3 | | 128,000 | 134.6 / 2,764.1 | 137.2 / 2,764.1 | 131.0 / 2,780.1 | - No measurable speed change on this hardware and no regression; differences are inside single-run min–max ranges. - Needles 6/6 on both versions. Coding check 10/10 on 0.1.40.3. - The 0.1.40.2 coding check scored 0/10. Re-runs with raw responses show the same greedy reasoning loop in 0.1.40.1 (first request after a fresh start, identical text in both versions, ended by `reasoning_loop_recovery: "stop"`), so we don't think it is a 0.1.40.2 regression. Details are in the README. - Raw data is included for 0.1.40.3 only. `telemetry.jsonl` is left out, `memory-summary.json` is kept. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sur le site
Liens install, modèles, releases.