Pull requests / #1333
#1333 bench: dual RTX 5090 production and synthetic Strata 0.1.40 results
open · @b1naryagent · 0 commentaires · Sur GitHub
BenchmarksMulti-GPUNVIDIA / CUDA
Description
Adds the reviewed dual-RTX-5090 Strata 0.1.40 performance report and a community benchmark index entry. The final production dataset contains 340 completed requests: 230.7 tok/s median native decode and 238.1 tok/s time-weighted decode. It includes 63 requests in the 192K–224K prompt bucket at 230.5 tok/s median. These aggregate production observations are separate from controlled synthetic layer-split versus peer-expert tests through 250,053 prompt tokens and a limited Q4 screen. The report documents hardware, sampling for the synthetic tests, overclocks, timing definitions, retention limits, and missing provenance. Private prompts and task content are excluded. No engine changes. Previously posted as #1285, now closed because this is a benchmark contribution. Validation: checked the final production values against the supplied report, verified bucket counts total 340 and uncached tokens equal prompt minus reused tokens, and preserved the synthetic sections unchanged. Only two Markdown files are changed by this PR. No new inference tests were run for this documentation submission.
Sur le site
Liens install, modèles, releases.