Pull requests / #1330

#1330 Add community benchmark: 2x Tesla P100 16GB

open · @touchtop · 0 comentarios · En GitHub

BenchmarksMulti-GPUModels & quantsDocumentation

Descripción

First reported P100 result for Strata (docs/OLDER_GPUS.md lists P100 as "not measured").

- 2x Tesla P100-PCIE-16GB, layer-split across GPUs 2+3 (same NUMA node)
- Qwen3.8-Flash-Next IQ3_S, 131,072-token context
- Decode: 31.1 tok/s (short), 28.0 tok/s (4K), 33.3 tok/s (32K)
- Prefill: 42 tok/s (short), 282 tok/s (4K), 454 tok/s (32K)
- Exceeds P40 community prompt ceiling (374 tok/s) by ~21% at 32K

See bench/results/2026-10-07-community-2x-p100/README.md for full hardware, software, and method details.

En el sitio

Enlaces a install, modelos, releases.