Pull requests / #1561
#1561 Community benchmark: dual V100-SXM2 16GB on Windows, IQ2_XS with 256K Vision
open · @intproduct · 0 comments · View on GitHub
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAWindows
Description
This adds a results-only report for Strata 0.1.40 on two Tesla V100-SXM2-16GB GPUs, an EPYC 7A23 and 64 GB RAM on Windows 10. IQ2_XS, INT8 KV, MTP and Vision stayed at the existing settings; the RTX PRO 4000 desktop GPU was excluded by UUID. Three fresh-prompt runs at each of 4096 / 32768 / 128000 / 258000 input tokens, with 256 output tokens and zero prefix reuse, measured decode medians of 93.05 / 91.66 / 87.02 / 78.99 tok/s. The report includes separate prefill/decode timing, TTFT, latency, sampled memory, per-run data and reproduction scripts. Official needle recall passed 9/9; small tool, image and coding checks are included. Limitations include one prompt family, three repetitions, an older installed release and unresolved long-session compatibility observations disclosed separately. Large raw streams, historical datasets and private translation logs are omitted; the full audit remains local. No engine or installation changes are proposed.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.