Pull requests / #902
#902 Community benchmark: Tesla V100-SXM2-32GB on Windows (CUDA 12.6 / MSVC)
closed · @Mitsuasa513 · 0 Kommentare · Auf GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Beschreibung
Adds a community benchmark report for a **Tesla V100-SXM2-32GB** running the experimental Volta build on **Windows 11 with CUDA 12.6 and MSVC 19.44**, and lists it in `docs/COMMUNITY_BENCHMARKS.md`. `docs/NVIDIA_V100.md` already documents the kernel paths and a Linux measurement. This report adds what that one does not cover: * **The Windows toolchain.** `-DSTRATA_EXPERIMENTAL_SM60=ON` with CUDA 12.6 / MSVC 19.44, `-allow-unsupported-compiler`, and the note that `CMAKE_CUDA_RUNTIME_LIBRARY` must be left alone because the flag already selects the shared runtime (#585). Configure and build: 0 errors. * **A quantization comparison on one machine** - the official GSQ-RCO `Q2_0` and `IQ2_XS` against Unsloth `UD-IQ4_XS`, same prompt, same settings: | model | decode | prefill | VRAM expert cache | | --- | ---: | ---: | --- | | Q2_0 | 76.4 tok/s | 1268 tok/s | ~18 900 slots / 24.4 GiB | | IQ2_XS | 71.2 (75.7 at 256K) tok/s | 1183 tok/s| ~18 200 slots | | UD-IQ4_XS | 26.3 tok/s | 665 tok/s | 10 078 slots | The 2-bit files are ~3x faster, and the engine's own numbers show why: half the bytes per expert means nearly twice as many fit in the same VRAM, and experts per layer left to the CPU drop from 3.9 to 0.7. * **A five-question correctness battery** (wolf-goat-cabbage, reverse arithmetic, 3L/5L jugs, chickens-and-rabbits, a `two_sum` implementation) through the resident server. `Q2_0` answered all five; `IQ2_XS` answered everything it finished, but spent its whole 1024-token budget on thinking for the wolf problem. One 2-bit run also fell into a verbatim repetition loop, which **enabling sampling escaped** - recorded because it is a failure mode worth knowing about. * **The PCIe link as the dominant variable**: the same 30 000-token prompt read at 179 tok/s on a Gen3 **x4** slot and 682 tok/s on **x16** (the engine's `pcie_frac` moves from 0.00 to 0.36). A factor of 3.8 from the slot alone. * **Two Windows-specific traps**: batch files containing non-ASCII bytes get mis-parsed under a GBK code page, and `taskkill` answers `Access denied` for the engine (leaving a 34 GB orphan) where `TerminateProcess` works. **Limits, stated in the report:** one machine and one operator; 1-2 runs per configuration and 5 prompts for the correctness table (enough to separate 2-bit from 4-bit, not enough to rank `Q2_0` against `IQ2_XS` precisely); engine-side numbers rather than client timings, with prompt sizes of 8 000 and 30 000 tokens rather than the usual 4K/32K/128K sweep; and an AVX2-only CPU, which is the slowest case for the expert pool and therefore the widest gap between quantizations.
Mehr auf der Site
Links zu Install, Modellen, Releases.