Pull requests / #154
#154 Correctness fixes from #149 (s_gemv barrier, bf16 NaN, QSA page mask, verify-window bounds, MSVC, test build)
closed · @gputier · 0 comentarios · En GitHub
BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
Descripción
The correctness fixes offered in #149, on top of 0.1.26 (ac8b251). One commit per fix; each is small and changes no result outside the case it fixes. - `s_gemv`: in `s_gemv_q8_split_kernel`, a warp with no row returned before the `__syncthreads()` that follows the codebook load. The codebook is now loaded first, then the rowless warp returns. `s_gemv_q8k_parity` gains output widths that are not a multiple of the eight rows per block. - `bf16`: the rounding add in `bf16_from_f32` carried a NaN out of the mantissa (0x7FFFFFFF came back as -0, 0x7F800001 as +inf). A NaN now stays a quiet NaN, as in ggml; `bf16_bits_test` checks all 2^32 inputs against ggml's rule, and `elementwise_parity` covers the device side, including the second copy in `dequant_bf16.cu`. The HIP `__builtin_memcpy` branch is kept as is; the NaN check comes right after it, so it applies to every backend. - `qsa`: a cell whose KV-streaming page did not resolve (-1) was read anyway in `qsa_decode_attn`; it is now masked with weight 0. `kv_stream` stops with a message when asked to sweep fewer slots than the resolve block, where two threads could take the same slot (the engine never goes below 5,120 slots, so this is a guard). - `expert_source`: the verify window's fixed tables (`kind[128]`, `distinct[128]`) are bounded by a `static_assert` and a runtime check on `n_tok * k`, and `s2_expert_grouped` asserts its group limit covers the window. - `gguf_reader.hpp`: `NOMINMAX` before `windows.h`, so every file that reaches this header can use `std::min`/`std::max` under MSVC (`gguf_reader_test.cpp`, `native_dense.cpp` and `native_head.cpp` include it without defining it). - build: 0.1.26 already guards `graph_registry_test` and `pinned_capture` on their file. The two `bench/micro/` targets still unguarded, `native_mmvq_multi` and `hit_cpu_order_parity`, get the same `EXISTS` check, with the backend variables. Before the rebase: tested on an RTX 5090 (Windows 11, MSVC 2022, CUDA 13.4, sm_120), ctest 30 of 33, the three failures being missing model files on my side (`ple_parity` wants a Q2_0 PLE table, `expert_parity` and `pool_test` want `pack/full/experts.bin`). After the rebase on 0.1.26: the four code fixes are byte-identical to the tested ones (`git range-diff`); only the `bf16_bits.hpp` conflict and the CMake commit changed. Checked on x86-64 Linux, CPU only: configure and build clean, `bf16_bits_test` passes (0 differences from ggml over 2^32 inputs), 10 of 11 CPU tests pass, and the failing one, `expert_multi_test`, fails the same way on plain 0.1.26. The CUDA and HIP paths were not rebuilt after the rebase, so your byte-identity gate is the check there. `qsa_decode_attn` and `s2_expert_grouped` have no ctest target, so the build is their only check.
En el sitio
Enlaces a install, modelos, releases.