Pull requests / #1399
#1399 qsa_select_bench: accuracy floor per sample (fails on an RTX 5090 / MSVC build)
open · @sergqwer · 0 Kommentare · Auf GitHub
Beschreibung
`qsa_select_bench` (a CUDA test since 65ce3299) fails on an RTX 5090 with a Windows (MSVC) build of v0.1.40.2: ``` FAIL accuracy vs FP64 (score scale 229): warp kernel max err 2.17e-05, fast scorer max err 0.000236 (1e-06 of scale) ctx 32768, 256 queries x 8193 blocks: ... selections identical 256/256, cells differing 0.0000% ``` The tensor-core scorer is fine: 1.03e-6 of the score scale against a floor of 1.00e-6, and every selection is the warp kernel's. The floor sits at the method's own error. 3xTF32 accumulates 48 MMAs (16 k-steps x 3) in the tensor cores' FP32, so its error is about 1e-6 of the score scale by construction. Which side of 1.00e-6 a run lands on depends on the data, and the data depend on the standard library. `std::normal_distribution<float>` is implementation-defined: MSVC's STL draws other numbers than libstdc++ from the same `mt19937(7)`. That is likely why it passes on Linux. **Change (test only):** the floor is now per sample and scales with that sample's own sum |q * k|. It is set at the worst case of a K = 128 FP32 dot product, 128 x 2^-24 = 2^-17 (7.6e-6) of that sum. The "no worse than 4x the warp kernel" part is unchanged. 3xTF32's bound is within the floor: 48 accumulations at <= 2^-23 each (tensor cores may truncate) plus the dropped lo * lo terms, about 7e-6. The line also prints the worst ratio and the floor. RTX 5090, CUDA 13.3, MSVC 17.14: | run | before | after: worst error / sum \|q*k\| | result | |---|---|---|---| | ctest `qsa_select_bench` (32K, tensor cores) | FAIL, 1.03e-6 of scale | 1.3e-6 (floor 7.6e-6) | PASS | | ctest `qsa_select_bench_fp32_tiled` (131K, `STRATA_SELECT_SIMT=1`) | PASS | 3.2e-7 | PASS | | default args (131K, tensor cores) | FAIL, 1.01e-6 of scale | 1.2e-6 | PASS | | negative control: plain TF32 (the two lo MMAs commented out) | - | 1.5e-4 (20x the floor; 17 of 256 selections differ) | FAIL | No engine code changes. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
Mehr auf der Site
Links zu Install, Modellen, Releases.