Pull requests / #379

#379 gr_parity: compare the split read (STRATA_GR_V3=1) within float rounding (the test fails with it on)

closed · @sergqwer · 0 コメント · GitHub で見る

NVIDIA / CUDA

本文

With `STRATA_GR_V3=1`, `gr_parity` fails on 0.1.31 main:

```
fused GR multi max-T=8 LDS launch and changing graph replay FAIL
```

The test compares the multi-token hyper-connection read with the single-token kernel bit for bit. That breaks under v3 for two reasons:
- **Summation order.** v3 splits `down` over (row group, stream) blocks and sums in another order, so it is equal to the reference only within float rounding.
- **`lo`.** The test also compares `lo`, the default kernels' workspace between down and up. v3 keeps that in shared memory and never writes it.

Under v3 the test now checks `R_out`, `rs`, `inject` and `mixed` within 2e-6 of the largest reference value, and leaves `lo` out. It prints the worst difference when that fails. Without v3 it compares bit for bit, as before.

Passes with and without `STRATA_GR_V3=1` (RTX 5090). Test-only: no engine code changes.

We run v3 on by default in our NVFP4 fork (decode rounds 3-5% shorter there). This PR only makes the test usable for anyone who turns it on.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

関連リンク

インストール・モデル・リリースへの站内リンク。