Issues / #1522

#1522 STRATA_GDN_CONVL2=1 is not bitwise on CUDA 13.3 / sm_120: gdn_l2_kernel compiles without the fma the fused kernel spells out

open · @sergqwer · 0 commentaires · Sur GitHub

AMD / HIPNVIDIA / CUDAWindows

Description

`STRATA_GDN_CONVL2=1` (9668e324, "bitwise"; on by default on gfx1151 through arch_defaults) changes the output on CUDA 13.3 / sm_120.

**Measured**: v0.1.40.3 (d5ea7133), RTX 5090, Windows / MSVC 17.14, CUDA 13.3. IQ2_XS, 2K prompt, `--expert-cache 12000 --pcie-frac 0.25`, 32 greedy tokens:

| | first-token logits (md5) | tokens (md5) |
|---|---|---|
| default | 1e96c370f5 | c0e83a232a |
| `STRATA_GDN_CONVL2=1` | d94aaabd3e | 737517ad5f |

First-token KL 0.138; the top token is the same.

**Why**: `gdn_conv_l2_kernel` spells out the reduction "as the compiler built" `gdn_l2_kernel`. The first xor step is an fma of the own unrounded square with the partner's rounded one, and the rest are plain adds. On this toolchain `gdn_l2_kernel` has no fma at all. `cuobjdump -sass -arch sm_120` of it:

```
FMUL R5, R0, R0
SHFL.BFLY PT, R6, R5, 0x10, 0x1f
FADD R6, R5, R6
SHFL.BFLY PT, R7, R6, 0x8, 0x1f
FADD R7, R6, R7
...
```

So `__fmaf_rn(v, v, partner)` rounds the first step differently, `ss` differs in its last bits, and the recurrence carries that through the layers. Whether the two kernels match depends on how the compiler contracts `warp_sum(v * v)`.

Possible fixes:
1. Spell out the same explicit intrinsics in `gdn_l2_kernel` as well, so both kernels are toolchain-independent. The default's bits would then move once, on toolchains that had contracted.
2. Use the same `warp_sum(v * v)` expression in the fused kernel. This is likely the same code on most toolchains, but not guaranteed.
3. Mark the switch as rounding-level on CUDA and keep it off there, as it already is.

I can send a PR for whichever you prefer.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Sur le site

Liens install, modèles, releases.