Pull requests / #1531

#1531 bench: community report, TITAN RTX 24 GB single card, IQ3_S at 262K: speed sweep and recall to 250K

open · @victorgabr · 0 comments · View on GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentationLinux

Description

First single-TITAN-RTX community row (#389 is the 2-card NVLink box). Engine 0.1.40.2, Linux, Xeon E5-2660 v3, 126 GiB RAM, NVMe.

**Speed sweep** (bench.py, streamed /v1/chat/completions, max_tokens 256, 3 runs per size, every prefill fully fresh - engine lines show 0 reused):

| Config | Prompt tok/s (median) | Decode tok/s (median) | TTFT s |
| --- | ---: | ---: | ---: |
| ~3.6K | 1,006 | 61.4 | 3.7 |
| ~28.6K | 1,213 | 56.1 | 23.8 |
| ~115.4K | 1,061 | 53.4 | 109.5 |

**Recall**: 9/9 needles at 32K/128K/250K x depths 10/50/90; fresh-prefill holds 862-865 tok/s at 234K fresh tokens inside the 262K context. First 250K-column data from this card.

Limitations: single session, expert cache warm during the sweep (88-91% hit; cold-box rows will sit lower); engine 0.1.40.2, not 0.1.40.4; PCIe link not measured; the first 32K needle attempt was cancelled (recorded in log-excerpt.txt) and its retry reused the prefix, so the fresh 32K rate is taken from the fully-fresh 32K/50 run.

Files: README.md (template), runs.json + log-excerpt.txt (raw), needles.json, bench.py, env.json, and the run config strata-iq3_s.json with absolute paths replaced by <STRATA>/<STRATA_DATA> placeholders (it binds host 0.0.0.0 with no API key as actually run - noted in the README).

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.