Issues / #1522
#1522 STRATA_GDN_CONVL2=1 is not bitwise on CUDA 13.3 / sm_120: gdn_l2_kernel compiles without the fma the fused kernel spells out
open · @sergqwer · 0 comments · View on GitHub
Description
`STRATA_GDN_CONVL2=1` (9668e324, "bitwise"; on by default on gfx1151 through arch_defaults) changes the output on CUDA 13.3 / sm_120. **Measured**: v0.1.40.3 (d5ea7133), RTX 5090, Windows / MSVC 17.14, CUDA 13.3. IQ2_XS, 2K prompt, `--expert-cache 12000 --pcie-frac 0.25`, 32 greedy tokens: | | first-token logits (md5) | tokens (md5) | |---|---|---| | default | 1e96c370f5 | c0e83a232a | | `STRATA_GDN_CONVL2=1` | d94aaabd3e | 737517ad5f | First-token KL 0.138; the top token is the same. **Why**: `gdn_conv_l2_kernel` spells out the reduction "as the compiler built" `gdn_l2_kernel`. The first xor step is an fma of the own unrounded square with the partner's rounded one, and the rest are plain adds. On this toolchain `gdn_l2_kernel` has no fma at all. `cuobjdump -sass -arch sm_120` of it: ``` FMUL R5, R0, R0 SHFL.BFLY PT, R6, R5, 0x10, 0x1f FADD R6, R5, R6 SHFL.BFLY PT, R7, R6, 0x8, 0x1f FADD R7, R6, R7 ... ``` So `__fmaf_rn(v, v, partner)` rounds the first step differently, `ss` differs in its last bits, and the recurrence carries that through the layers. Whether the two kernels match depends on how the compiler contracts `warp_sum(v * v)`. Possible fixes: 1. Spell out the same explicit intrinsics in `gdn_l2_kernel` as well, so both kernels are toolchain-independent. The default's bits would then move once, on toolchains that had contracted. 2. Use the same `warp_sum(v * v)` expression in the fused kernel. This is likely the same code on most toolchains, but not guaranteed. 3. Mark the switch as rounding-level on CUDA and keep it off there, as it already is. I can send a PR for whichever you prefer. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.