Pull requests / #532
#532 tests: prefill_mmq_kquant_test waits for its uploads before the kernels
closed · draft · @eadra · 0 comentários · No GitHub
Server & APINVIDIA / CUDAModels & quantsWindows
Descrição
## What `prefill_mmq_kquant_test` uploads its inputs and sets the NaN sentinel with `cudaMemcpy` / `cudaMemset` on the legacy stream, then runs the MMQ kernels on a `cudaStreamNonBlocking` stream, which does not wait for the legacy stream. A pageable copy (or the sentinel memset) can still be in flight when the kernels start. This adds a `cudaDeviceSynchronize()` before the kernels, the same fix the fused tests got in 0.1.36 (`prefill_fused_moe_test.cpp`, `prefill_fused_iq_test.cpp`). Test-only: the engine is unchanged. ## Measured RTX 4090 (sm_89, 128 SMs), Windows 11, CUDA 13.3, a `-DSTRATA_MMQ_KQUANTS=ON -DSTRATA_BUILD_TESTS=ON` build: | | runs | result | | --- | ---: | --- | | without the sync | 3 | 3 failed: "unwritten or non-finite output", in a different product each run (Q4_K, Q5_K, Q5_1 in the multi-expert batches) | | with the sync | 5 | 5 passed, every product inside the screen | The failing outputs were whole 128-row tiles left at the sentinel (NaN), or zero when the sentinel was skipped. The kernels themselves are fine: the same products pass once the inputs are guaranteed to be on the device. ## Note [#478](https://github.com/Niko1221/Strata/pull/478) (Q6_K) edits the same file, in the header comment and `main()`. This change is inside `product()`, so the two apply independently. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
No site
Links install, modelos, releases.