Pull requests / #707
#707 docs: community benchmark row - Tesla V100-PCIE-32GB (sm_70 source build, before/after calibration)
closed · @noahark · 0 comentários · No GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
Descrição
Adds the V100 row invited in #617: a field-measurement report at `bench/results/2026-10-03-community-v100/README.md` plus the index entry in `docs/COMMUNITY_BENCHMARKS.md`. What it covers: - Tesla V100-PCIE-32GB (sm_70, PCIe Gen3) + dual Xeon E5-2696 v3 + 128 GB DDR4, Windows; v0.1.38 source build (CUDA 12.4, MSVC; the link fix from #585 applied) - Flash-Next IQ3_XXS, 131,072 context, KV int8 streamed, MTP on - Decode speeds across real workloads before `--calibrate` (7.5 warm-up to 16.8 long-form) and the doubling to 33.5 tok/s after, with the full pcie-frac / spec-min-p / pool-workers sweep - including the dual-socket NUMA finding that 18 CPU workers beat 35 by 20%+ - Prompt-processing scaling (57 tok/s at ~90-token prompts to 863 tok/s on a 4.3k-token document), MTP acceptance 65-98%, expert cache hit 95-97% These are field measurements from daily use and a capability battery rather than a fixed-cap harness; the README labels them as such per the guide. Happy to add a controlled short/long-prompt run with a fixed 256-token cap if that fits the format better.
No site
Links install, modelos, releases.