Pull requests / #154

#154 Correctness fixes from #149 (s_gemv barrier, bf16 NaN, QSA page mask, verify-window bounds, MSVC, test build)

closed · @gputier · 0 comentários · No GitHub

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Descrição

The correctness fixes offered in #149, on top of 0.1.26 (ac8b251). One commit per fix; each is small and changes no result outside the case it fixes.

- `s_gemv`: in `s_gemv_q8_split_kernel`, a warp with no row returned before the `__syncthreads()` that follows the codebook load. The codebook is now loaded first, then the rowless warp returns. `s_gemv_q8k_parity` gains output widths that are not a multiple of the eight rows per block.
- `bf16`: the rounding add in `bf16_from_f32` carried a NaN out of the mantissa (0x7FFFFFFF came back as -0, 0x7F800001 as +inf). A NaN now stays a quiet NaN, as in ggml; `bf16_bits_test` checks all 2^32 inputs against ggml's rule, and `elementwise_parity` covers the device side, including the second copy in `dequant_bf16.cu`. The HIP `__builtin_memcpy` branch is kept as is; the NaN check comes right after it, so it applies to every backend.
- `qsa`: a cell whose KV-streaming page did not resolve (-1) was read anyway in `qsa_decode_attn`; it is now masked with weight 0. `kv_stream` stops with a message when asked to sweep fewer slots than the resolve block, where two threads could take the same slot (the engine never goes below 5,120 slots, so this is a guard).
- `expert_source`: the verify window's fixed tables (`kind[128]`, `distinct[128]`) are bounded by a `static_assert` and a runtime check on `n_tok * k`, and `s2_expert_grouped` asserts its group limit covers the window.
- `gguf_reader.hpp`: `NOMINMAX` before `windows.h`, so every file that reaches this header can use `std::min`/`std::max` under MSVC (`gguf_reader_test.cpp`, `native_dense.cpp` and `native_head.cpp` include it without defining it).
- build: 0.1.26 already guards `graph_registry_test` and `pinned_capture` on their file. The two `bench/micro/` targets still unguarded, `native_mmvq_multi` and `hit_cpu_order_parity`, get the same `EXISTS` check, with the backend variables.

Before the rebase: tested on an RTX 5090 (Windows 11, MSVC 2022, CUDA 13.4, sm_120), ctest 30 of 33, the three failures being missing model files on my side (`ple_parity` wants a Q2_0 PLE table, `expert_parity` and `pool_test` want `pack/full/experts.bin`).

After the rebase on 0.1.26: the four code fixes are byte-identical to the tested ones (`git range-diff`); only the `bf16_bits.hpp` conflict and the CMake commit changed. Checked on x86-64 Linux, CPU only: configure and build clean, `bf16_bits_test` passes (0 differences from ggml over 2^32 inputs), 10 of 11 CPU tests pass, and the failing one, `expert_multi_test`, fails the same way on plain 0.1.26. The CUDA and HIP paths were not rebuilt after the rebase, so your byte-identity gate is the check there.

`qsa_decode_attn` and `s2_expert_grouped` have no ctest target, so the build is their only check.

No site

Links install, modelos, releases.