Pull requests / #148
#148 tests: bitwise check of native_mmvq's multi_exact contract, with a negative control
closed · @enkynakamura · 0 コメント · GitHub で見る
本文
native_mmvq.hpp states that with multi_exact on, "every column bitwise equal to a single-column call", and verify.hpp counts it among the reasons the verify window is bit-exact. In the published tree nothing checks it: iq_parity compares two columns against a float64 reference at 2e-2. CMakeLists.txt registers a native_mmvq_multi --selftest from bench/micro/, which the published source leaves out, so its coverage could not be verified. This adds mmvq_multi_parity, built with STRATA_ENABLE_CUDA alone and registered with add_test: the same Q8_1 bytes through one T-column call and T single-column calls, compared bit for bit, for 7 cases (Q4_0, Q8_0, Q4_K, Q5_K, Q6_K, IQ4_XS, a wide Q5_K) and T = 1, 2, 3, 4, 5, 6, 8; random weight bytes with each block's fp16 scales rewritten as normal values, so every output is finite: on the GPU every NaN has the same bits, so a NaN output would compare equal whatever produced it. A non-finite output fails the test; a negative control: with multi_exact off, every case must show differences at some T > 4, or the test fails for lack of power. At T ≤ 4 the control's silence is reported, not skipped: the generic layout uses NW = NCOLS <= 4 ? 4 : 2 warps, and with NW = 4 each row runs the same operations in the same order as the exact layout. IQ4_XS uses n_in = 4096: at 2048 (8 blocks, 16 per iteration with 4 warps, 8 with 2) the two layouts coincide at every T. Result on an RTX 5080 (sm_120), built from a tree whose native_mmvq sources are identical to this base: with multi_exact on, 0 of 148,480 outputs differ, 0 non-finite; with it off, 63,901 differ, all at T > 4, in every case.
関連リンク
インストール・モデル・リリースへの站内リンク。