Pull requests / #532

#532 tests: prefill_mmq_kquant_test waits for its uploads before the kernels

closed · draft · @eadra · 0 Kommentare · Auf GitHub

Server & APINVIDIA / CUDAModels & quantsWindows

Beschreibung

## What

`prefill_mmq_kquant_test` uploads its inputs and sets the NaN sentinel with `cudaMemcpy` / `cudaMemset` on the legacy
stream, then runs the MMQ kernels on a `cudaStreamNonBlocking` stream, which does not wait for the legacy stream. A
pageable copy (or the sentinel memset) can still be in flight when the kernels start. This adds a
`cudaDeviceSynchronize()` before the kernels, the same fix the fused tests got in 0.1.36
(`prefill_fused_moe_test.cpp`, `prefill_fused_iq_test.cpp`). Test-only: the engine is unchanged.

## Measured

RTX 4090 (sm_89, 128 SMs), Windows 11, CUDA 13.3, a `-DSTRATA_MMQ_KQUANTS=ON -DSTRATA_BUILD_TESTS=ON` build:

| | runs | result |
| --- | ---: | --- |
| without the sync | 3 | 3 failed: "unwritten or non-finite output", in a different product each run (Q4_K, Q5_K, Q5_1 in the multi-expert batches) |
| with the sync | 5 | 5 passed, every product inside the screen |

The failing outputs were whole 128-row tiles left at the sentinel (NaN), or zero when the sentinel was skipped. The
kernels themselves are fine: the same products pass once the inputs are guaranteed to be on the device.

## Note

[#478](https://github.com/Niko1221/Strata/pull/478) (Q6_K) edits the same file, in the header comment and `main()`.
This change is inside `product()`, so the two apply independently.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Mehr auf der Site

Links zu Install, Modellen, Releases.