Pull requests / #1192

#1192 Community benchmark: 4x Tesla P100 16 GB (Pascal), IQ3_XXS, CUDA 12 engine

closed · @apollo-mg · 0 comentarios · En GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentation

Descripción

Results-only report for the experimental Pascal path, which OLDER_GPUS.md lists as not measured for the P100. 4x P100-PCIE-16GB, 2x Xeon E5-2650 v3, 128 GB; Strata v0.1.39, CUDA 12 engine built by setup with CUDA 12.4 and gcc-13 as the host compiler. IQ3_XXS across all four cards: decode 33.5 tok/s on short prompts and 28-29 tok/s after 4.4K and 36.4K prompts; prompt reading 94, 239 and 594 tok/s; decode expert cache hit rate 98.7-99.7%. Clocks pinned at 1063 MHz, single machine, no accuracy checks. Per-run data and engine timing lines are in the README folder.

(Test designed and run with my agent, Claude; I reviewed it before posting.)

En el sitio

Enlaces a install, modelos, releases.