Pull requests / #1647

#1647 bench: community report - a Tesla V100 (sm_70) confirms the opt-in Volta decode kernels

open · @Mitsuasa513 · 0 コメント · GitHub で見る

BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows

本文

A V100-SXM2-32GB owner's confirmation of the three opt-in Volta decode kernels
(`STRATA_SM70_TABLE=1`), which `docs/NVIDIA_V100.md` asks a V100 owner to check before they
become the sm_70 default.

- Hardware: V100-SXM2-32GB (sm_70, PCIe Gen3 x16), Threadripper 2990WX without AVX-512,
  96 GB RAM, Windows 11, driver 560.94, CUDA 12.6; engine `fb58e0d` (v0.1.41) built for sm_70
  (`-DSTRATA_EXPERIMENTAL_SM60=ON -DCMAKE_CUDA_ARCHITECTURES=70`).
- Model: Unsloth UD-IQ4_XS with a native-experts pack, 262,144 context, `--kv int8`,
  `--prefill auto`, `--spec 3`.
- Decode: +9.8% and +15.0% in two order-swapped pairs (28.16/28.64 -> 30.92/32.94 tok/s).
  Prompt throughput unchanged (1619-1633 tok/s). The two A runs agree within 1.7%.
- Correctness: all four runs produced the same token sequence (sha256 258d9017223cd4d36c7b47bf,
  256 tokens) and the same draft acceptance, so the kernels change the speed and not the arithmetic.
- Also reported honestly: `tools/ab_engine.py`'s three short chats are below this machine's noise
  floor; the long-prompt arm agrees in direction (+5%). Raw rows in `ab-engine.jsonl`.
- Limits: one machine, one model, greedy; loading excluded; no needle/image test here.

This is a results-only report; it changes no engine code.

関連リンク

インストール・モデル・リリースへの站内リンク。