Pull requests / #1399

#1399 qsa_select_bench: accuracy floor per sample (fails on an RTX 5090 / MSVC build)

open · @sergqwer · 0 Kommentare · Auf GitHub

NVIDIA / CUDAWindowsLinux

Beschreibung

`qsa_select_bench` (a CUDA test since 65ce3299) fails on an RTX 5090 with a Windows (MSVC) build of v0.1.40.2:

```
FAIL accuracy vs FP64 (score scale 229): warp kernel max err 2.17e-05, fast scorer max err 0.000236 (1e-06 of scale)
ctx 32768, 256 queries x 8193 blocks: ... selections identical 256/256, cells differing 0.0000%
```

The tensor-core scorer is fine: 1.03e-6 of the score scale against a floor of 1.00e-6, and every selection is the warp kernel's. The floor sits at the method's own error. 3xTF32 accumulates 48 MMAs (16 k-steps x 3) in the tensor cores' FP32, so its error is about 1e-6 of the score scale by construction. Which side of 1.00e-6 a run lands on depends on the data, and the data depend on the standard library. `std::normal_distribution<float>` is implementation-defined: MSVC's STL draws other numbers than libstdc++ from the same `mt19937(7)`. That is likely why it passes on Linux.

**Change (test only):** the floor is now per sample and scales with that sample's own sum |q * k|. It is set at the worst case of a K = 128 FP32 dot product, 128 x 2^-24 = 2^-17 (7.6e-6) of that sum. The "no worse than 4x the warp kernel" part is unchanged. 3xTF32's bound is within the floor: 48 accumulations at <= 2^-23 each (tensor cores may truncate) plus the dropped lo * lo terms, about 7e-6. The line also prints the worst ratio and the floor.

RTX 5090, CUDA 13.3, MSVC 17.14:

| run | before | after: worst error / sum \|q*k\| | result |
|---|---|---|---|
| ctest `qsa_select_bench` (32K, tensor cores) | FAIL, 1.03e-6 of scale | 1.3e-6 (floor 7.6e-6) | PASS |
| ctest `qsa_select_bench_fp32_tiled` (131K, `STRATA_SELECT_SIMT=1`) | PASS | 3.2e-7 | PASS |
| default args (131K, tensor cores) | FAIL, 1.01e-6 of scale | 1.2e-6 | PASS |
| negative control: plain TF32 (the two lo MMAs commented out) | - | 1.5e-4 (20x the floor; 17 of 256 selections differ) | FAIL |

No engine code changes.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Mehr auf der Site

Links zu Install, Modellen, Releases.