Pull requests / #823
#823 bench: community report, Tesla V100 32 GB with 16 GB of RAM (Coder IQ1_M)
closed · @christopherrobertbrooks-tech · 0 comments · View on GitHub
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quants
Description
Community benchmark report: **Tesla V100-PCIE-32GB with 16 GB of system RAM**, Coder IQ1_M, engine 0.1.39 (experimental CUDA 12 build for sm_70), in `bench/results/2026-10-04-community-v100-16gb-ram/`. Results only, no code changes. - **Hardware:** V100 32 GB on PCIe 3.0 x4, i7-13700KF, 16 GB DDR5 (below the Coder's usual 32 GB), SATA SSD. - **Configuration:** setup's low-RAM choice (`--resident-experts`), KV int8 in VRAM, vision on, MTP `--spec 4`. 11,650 of 12,288 experts in the GPU cache with vision on (all 12,288 without it); 0.0 MB read from the model files in every run. - **Method:** the `benchmark.py` from the 2x MI50 report, unchanged: 3 runs each at 4,096 / 32,768 / 128,000 fresh prompt tokens, 256-token output cap, reasoning none. - **Results (medians):** prompt 1,233 / 1,580 / 1,394 tok/s; decode 69.3 / 68.6 / 67.1 tok/s; TTFT 3.35 / 20.8 / 92.0 s. - **Correctness:** needle checks 6/6 (32K and 128K, depths 10/50/90). - **Limitations:** one machine, one size, three runs per configuration; requests went through llama-swap's proxy on the same machine; not tested with a second card. The report was drafted with an AI assistant (Claude) from our measurements, and checked by me.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.