Pull requests / #823

#823 bench: community report, Tesla V100 32 GB with 16 GB of RAM (Coder IQ1_M)

closed · @christopherrobertbrooks-tech · 0 Kommentare · Auf GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quants

Beschreibung

Community benchmark report: **Tesla V100-PCIE-32GB with 16 GB of system RAM**, Coder IQ1_M, engine 0.1.39 (experimental
CUDA 12 build for sm_70), in `bench/results/2026-10-04-community-v100-16gb-ram/`. Results only, no code changes.

- **Hardware:** V100 32 GB on PCIe 3.0 x4, i7-13700KF, 16 GB DDR5 (below the Coder's usual 32 GB), SATA SSD.
- **Configuration:** setup's low-RAM choice (`--resident-experts`), KV int8 in VRAM, vision on, MTP `--spec 4`.
  11,650 of 12,288 experts in the GPU cache with vision on (all 12,288 without it); 0.0 MB read from the model files in every run.
- **Method:** the `benchmark.py` from the 2x MI50 report, unchanged: 3 runs each at 4,096 / 32,768 / 128,000 fresh prompt
  tokens, 256-token output cap, reasoning none.
- **Results (medians):** prompt 1,233 / 1,580 / 1,394 tok/s; decode 69.3 / 68.6 / 67.1 tok/s; TTFT 3.35 / 20.8 / 92.0 s.
- **Correctness:** needle checks 6/6 (32K and 128K, depths 10/50/90).
- **Limitations:** one machine, one size, three runs per configuration; requests went through llama-swap's proxy on the
  same machine; not tested with a second card.

The report was drafted with an AI assistant (Claude) from our measurements, and checked by me.

Mehr auf der Site

Links zu Install, Modellen, Releases.