Pull requests / #1226

#1226 Community benchmark: 2x RTX PRO 4500 Blackwell, Swift IQ3_XXS, engine 0.1.40.1 (solo to 256k, --batch after #776)

closed · @qni-live · 0 comments · View on GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationLinux

Description

Adds one results folder, `bench/results/2026-10-06-community-2x-rtx-pro-4500-engine-0.1.40.1/`, and nothing else. No engine changes.

- Hardware: 2x RTX PRO 4500 Blackwell 32 GB (sm_120), Threadripper 7960X, 61 GiB RAM, Linux, driver 595.91.07, CUDA 13.4. Source build of v0.1.40.1 (82f46a8).
- Model: Swift 1.5 IQ3_XXS, layer split over both cards, 262144 context.
- Single stream at 1k, 4k, 32k, 128k and 256k prompt tokens (3 runs each), engine 0.1.40.1 against 0.1.39 measured right after it. Prompt reading 1,889 to 6,127 tok/s; decode 134 to 162 tok/s. The decode difference to 0.1.39 is not established.
- Concurrent requests with and without `--batch` (the fix from 0.1.40 for #776), 1 to 4 requests: no engine exit, no failed request; 2 slots +8% at 2 requests, 4 slots +29% at 4 requests (total tok/s including prompt reading).
- Needle recall 6 of 6, tool calls 9 of 9 valid.
- README follows `docs/COMMUNITY_BENCHMARKS.md`; per-run JSON, scripts, configs and the build commit are in the folder. Untested items are listed at the end of the README.

This replaces my earlier PR #418, which GitHub closed when `main` was force-pushed. Its engine 0.1.30 and 0.1.36 folders are not part of this PR; I can re-add them if you want them.

The measurements were run by an AI agent (Claude Code) on my machine. I decided the plan and reviewed the results and the README before opening this PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.