Pull requests / #568

#568 parity tests: ple_parity runs without the missing fixtures, gr_parity and ple_parity wait for their setup

closed · @gputier · 0 commentaires · Sur GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quants

Description

Two test fixes on top of v0.1.37. Engine code is untouched.

## ple_parity fails on every release

On v0.1.37, `ple_parity` fails at every launch (200 out of 200, RTX 5090) with:

    ple_parity: required block fixtures are missing, truncated or incompatible: bench/micro/ple_in.bin / bench/micro/ple_out.bin

Neither file is in any release or in any commit of the repository. Their generator, `ple_layer_xcheck.cpp`, shows up only in comments, since the initial commit f2a08d4. It was never shipped.

## Commit 1: gr_parity, ple_parity wait for the default-stream setup

Several test sections prepare their data with `cudaMemcpy` or `cudaMemset` on the legacy default stream, then launch kernels on a `cudaStreamNonBlocking` stream without waiting. A non-blocking stream does not synchronize with the legacy stream, and a `cudaMemcpy` from pageable memory can return before the DMA has landed, so the kernels may read data that is not there yet. The same pattern was fixed in the tests merged in 0.1.36.

The commit adds a `cudaDeviceSynchronize()` before the first launch at each affected site: 7 calls, 12 lines with their comments, in `gr_parity.cpp` (scalar section, multi section) and `ple_parity.cpp` (history regression, native key stream). Nothing else changes.

I could not reproduce the race. `gr_parity` gave 0 failures out of 210 launches before the commit and 0 out of 210 after, on an RTX 5090. The fix rests on reading the code and on the CUDA stream semantics, not on a failing run. `ple_parity` cannot show it either, since it stops on the missing fixtures before reaching these sites.

The same pattern may exist in three HIP tests. I read them but did not run or touch them:
- `tests/hip/ple_iq4.cpp:154-157`
- `tests/hip/prefill_mmq_parity.cpp:135-140` and `:252`
- `tests/hip/prefill_native_batch.cpp:56-59` and `:274`

## Commit 2: ple_parity checks against a double-precision reference

The test now computes the PLE block in double precision, inside the test, from the real weights (Q2_0 pack `pack/full/dense.bin`, offsets from `index.txt`, and the table rows from the GGUF). It compares each stage of the block with that reference, with a tolerance of 1e-5. `--in` and `--out` are gone, along with `STRATA_PLE_FIXTURE_DIR`.

This changes what the test proves. The original oracle was a capture of ggml's CPU graph. The new reference is arithmetic written in the test, so it checks that the kernel computes the arithmetic it announces, not that it matches ggml bit for bit or within ggml's own rounding. If you still have the generator or the two files, publishing them restores the ggml oracle, and the double reference can stay next to it.

The engine is unchanged: `ple.cu` and `ple.hpp` change in comments only, to stop pointing at the missing files. `CMakeLists.txt` changes only the `ple_parity` test declaration and the comment above it.

Measured on an RTX 5090, v0.1.37 with both commits:
- `ple_parity`: 10 launches out of 10 pass, worst stage 1.148e-07 against the 1e-5 tolerance.
- Counter-test: a 1% error put in `silu_f` in a copy fails 3 launches out of 3, with 6 failing stages (`conv_out` at 1.000e-02 and `result` between 1.5e-04 and 1.8e-04, for each of the 3 tokens). The error was removed from the copy afterwards.

Not run: the full ctest suite. The counter-test only breaks a kernel stage, not a weight of the pack.

Sur le site

Liens install, modèles, releases.