Pull requests / #902
#902 Community benchmark: Tesla V100-SXM2-32GB on Windows (CUDA 12.6 / MSVC)
closed · @Mitsuasa513 · 0 comments · View on GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Description
Adds a community benchmark report for a **Tesla V100-SXM2-32GB** running the experimental Volta build on **Windows 11 with CUDA 12.6 and MSVC 19.44**, and lists it in `docs/COMMUNITY_BENCHMARKS.md`. `docs/NVIDIA_V100.md` already documents the kernel paths and a Linux measurement. This report adds what that one does not cover: * **The Windows toolchain.** `-DSTRATA_EXPERIMENTAL_SM60=ON` with CUDA 12.6 / MSVC 19.44, `-allow-unsupported-compiler`, and the note that `CMAKE_CUDA_RUNTIME_LIBRARY` must be left alone because the flag already selects the shared runtime (#585). Configure and build: 0 errors. * **A quantization comparison on one machine** - the official GSQ-RCO `Q2_0` and `IQ2_XS` against Unsloth `UD-IQ4_XS`, same prompt, same settings: | model | decode | prefill | VRAM expert cache | | --- | ---: | ---: | --- | | Q2_0 | 76.4 tok/s | 1268 tok/s | ~18 900 slots / 24.4 GiB | | IQ2_XS | 71.2 (75.7 at 256K) tok/s | 1183 tok/s| ~18 200 slots | | UD-IQ4_XS | 26.3 tok/s | 665 tok/s | 10 078 slots | The 2-bit files are ~3x faster, and the engine's own numbers show why: half the bytes per expert means nearly twice as many fit in the same VRAM, and experts per layer left to the CPU drop from 3.9 to 0.7. * **A five-question correctness battery** (wolf-goat-cabbage, reverse arithmetic, 3L/5L jugs, chickens-and-rabbits, a `two_sum` implementation) through the resident server. `Q2_0` answered all five; `IQ2_XS` answered everything it finished, but spent its whole 1024-token budget on thinking for the wolf problem. One 2-bit run also fell into a verbatim repetition loop, which **enabling sampling escaped** - recorded because it is a failure mode worth knowing about. * **The PCIe link as the dominant variable**: the same 30 000-token prompt read at 179 tok/s on a Gen3 **x4** slot and 682 tok/s on **x16** (the engine's `pcie_frac` moves from 0.00 to 0.36). A factor of 3.8 from the slot alone. * **Two Windows-specific traps**: batch files containing non-ASCII bytes get mis-parsed under a GBK code page, and `taskkill` answers `Access denied` for the engine (leaving a 34 GB orphan) where `TerminateProcess` works. **Limits, stated in the report:** one machine and one operator; 1-2 runs per configuration and 5 prompts for the correctness table (enough to separate 2-bit from 4-bit, not enough to rank `Q2_0` against `IQ2_XS` precisely); engine-side numbers rather than client timings, with prompt sizes of 8 000 and 30 000 tokens rather than the usual 4K/32K/128K sweep; and an AVX2-only CPU, which is the slowest case for the expert pool and therefore the widest gap between quantizations.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.