Pull requests / #1647
#1647 bench: community report - a Tesla V100 (sm_70) confirms the opt-in Volta decode kernels
open · @Mitsuasa513 · 0 comments · View on GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows
Description
A V100-SXM2-32GB owner's confirmation of the three opt-in Volta decode kernels (`STRATA_SM70_TABLE=1`), which `docs/NVIDIA_V100.md` asks a V100 owner to check before they become the sm_70 default. - Hardware: V100-SXM2-32GB (sm_70, PCIe Gen3 x16), Threadripper 2990WX without AVX-512, 96 GB RAM, Windows 11, driver 560.94, CUDA 12.6; engine `fb58e0d` (v0.1.41) built for sm_70 (`-DSTRATA_EXPERIMENTAL_SM60=ON -DCMAKE_CUDA_ARCHITECTURES=70`). - Model: Unsloth UD-IQ4_XS with a native-experts pack, 262,144 context, `--kv int8`, `--prefill auto`, `--spec 3`. - Decode: +9.8% and +15.0% in two order-swapped pairs (28.16/28.64 -> 30.92/32.94 tok/s). Prompt throughput unchanged (1619-1633 tok/s). The two A runs agree within 1.7%. - Correctness: all four runs produced the same token sequence (sha256 258d9017223cd4d36c7b47bf, 256 tokens) and the same draft acceptance, so the kernels change the speed and not the arithmetic. - Also reported honestly: `tools/ab_engine.py`'s three short chats are below this machine's noise floor; the long-prompt arm agrees in direction (+5%). Raw rows in `ab-engine.jsonl`. - Limits: one machine, one model, greedy; loading excluded; no needle/image test here. This is a results-only report; it changes no engine code.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.