Pull requests / #1260
#1260 tests: order MMQ parity transfers on the compute stream
closed · @mcuygjsliy429 · 0 comments · View on GitHub
AMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindowsLinux
Description
## Summary
Test-only fix for the same synchronization issue reported in #532. That PR was closed during the history rewrite; this branch starts from the current `main` (`82f46a8c`). Instead of adding a device-wide barrier, keep each test product's transfers and compute on its existing non-blocking stream.
`prefill_mmq_kquant_test` currently uploads pageable buffers and initializes the NaN sentinel on the default stream, then launches kernels on `cudaStreamNonBlocking`. The two streams have no ordering guarantee. A pageable H2D copy may return before its final DMA completes, and the device memset is host-asynchronous. See NVIDIA's [API synchronization](https://docs.nvidia.com/cuda/cuda-runtime-api/api-sync-behavior.html) and [stream synchronization](https://docs.nvidia.com/cuda/cuda-runtime-api/stream-sync-behavior.html) documentation.
## Changes
- Queue all input copies, the NaN sentinel, kernels and result copy on stream `s`; synchronize it before inspecting host results.
- Add `--stress-default-stream` and a second CTest entry. A short host callback adds unrelated default-stream work while the product uses `s`; this is scheduling stress, not a claim that every unsynchronized run must fail.
- Cover empty experts and uneven batches around tile boundaries: `{0,1,0,3,0}`, `{1,7,8,9}`, `{1,127,1}`.
- Link the test with `Threads::Threads` for the callback's sleep. Use the default-stream null handle, which the existing gfx906 HIP compatibility path also accepts.
No inference code, quantizers, reference math or numerical thresholds change.
## Validation
Linux, GCC 13.3, CUDA 12.8; RTX 3090 (sm_86) and RTX 4080 SUPER (sm_89):
- Configured the patched current-main source with `STRATA_ENABLE_CUDA=ON`, `STRATA_BUILD_TESTS=ON`, `STRATA_MMQ_KQUANTS=ON`, `STRATA_Q6K_EXPERTS=ON`, and architectures `86;89`.
- Compiled the test object through the generated Ninja rules and checked registration of both CTest entries (120-second timeouts).
- Separately linked the exact patched test source against the existing **unmodified** MMQ and ggml-base archives, using pinned ggml `3cf03257f219afbe7334045ff7c6a06ac68c627d`. The adapter source/header match current main byte-for-byte. This was not a full engine rebuild.
- Both normal and stress invocations passed on both GPUs: 4 runs, 35 matrix products each, 140 checks total. No tolerance was relaxed. Windows and HIP execution have not been tested.
- `git diff --check` passed. The serving engine was not restarted or replaced.
To run the two registered tests in a built CUDA/Q6 test tree:
```sh
ctest --test-dir build -R '^prefill_mmq_(kquant|stream_order)_test$' --output-on-failure
```
## Separate limitation
During earlier investigation, Compute Sanitizer also reported tail row-ID reads past the allocation in the pinned ggml MMQ kernel, independently of this test-stream race. This PR does **not** fix that dependency issue or claim a sanitizer-clean engine. Clean sanitizer and full-model results obtained with an additional private kernel patch are deliberately not attributed to this test-only PR.
AI-assisted implementation and validation, submitted with the operator's authorization.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.