Pull requests / #707

#707 docs: community benchmark row - Tesla V100-PCIE-32GB (sm_70 source build, before/after calibration)

closed · @noahark · 0 コメント · GitHub で見る

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows

本文

Adds the V100 row invited in #617: a field-measurement report at
`bench/results/2026-10-03-community-v100/README.md` plus the index entry in
`docs/COMMUNITY_BENCHMARKS.md`.

What it covers:

- Tesla V100-PCIE-32GB (sm_70, PCIe Gen3) + dual Xeon E5-2696 v3 + 128 GB DDR4, Windows;
  v0.1.38 source build (CUDA 12.4, MSVC; the link fix from #585 applied)
- Flash-Next IQ3_XXS, 131,072 context, KV int8 streamed, MTP on
- Decode speeds across real workloads before `--calibrate` (7.5 warm-up to 16.8 long-form)
  and the doubling to 33.5 tok/s after, with the full pcie-frac / spec-min-p / pool-workers
  sweep - including the dual-socket NUMA finding that 18 CPU workers beat 35 by 20%+
- Prompt-processing scaling (57 tok/s at ~90-token prompts to 863 tok/s on a 4.3k-token
  document), MTP acceptance 65-98%, expert cache hit 95-97%

These are field measurements from daily use and a capability battery rather than a
fixed-cap harness; the README labels them as such per the guide. Happy to add a
controlled short/long-prompt run with a fixed 256-token cap if that fits the format better.

関連リンク

インストール・モデル・リリースへの站内リンク。