Pull requests / #823
#823 bench: community report, Tesla V100 32 GB with 16 GB of RAM (Coder IQ1_M)
closed · @christopherrobertbrooks-tech · 0 comentários · No GitHub
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quants
Descrição
Community benchmark report: **Tesla V100-PCIE-32GB with 16 GB of system RAM**, Coder IQ1_M, engine 0.1.39 (experimental CUDA 12 build for sm_70), in `bench/results/2026-10-04-community-v100-16gb-ram/`. Results only, no code changes. - **Hardware:** V100 32 GB on PCIe 3.0 x4, i7-13700KF, 16 GB DDR5 (below the Coder's usual 32 GB), SATA SSD. - **Configuration:** setup's low-RAM choice (`--resident-experts`), KV int8 in VRAM, vision on, MTP `--spec 4`. 11,650 of 12,288 experts in the GPU cache with vision on (all 12,288 without it); 0.0 MB read from the model files in every run. - **Method:** the `benchmark.py` from the 2x MI50 report, unchanged: 3 runs each at 4,096 / 32,768 / 128,000 fresh prompt tokens, 256-token output cap, reasoning none. - **Results (medians):** prompt 1,233 / 1,580 / 1,394 tok/s; decode 69.3 / 68.6 / 67.1 tok/s; TTFT 3.35 / 20.8 / 92.0 s. - **Correctness:** needle checks 6/6 (32K and 128K, depths 10/50/90). - **Limitations:** one machine, one size, three runs per configuration; requests went through llama-swap's proxy on the same machine; not tested with a second card. The report was drafted with an AI assistant (Claude) from our measurements, and checked by me.
No site
Links install, modelos, releases.