Pull requests / #1170

#1170 test(gr): check the fused multi-token read at T=4 and T=2

open · @InB4DevOps · 0 评论 · 在 GitHub 查看

NVIDIA / CUDAWindows

描述

Follow-up to #853, which was closed when `main` was force-pushed.

**The kernel half of #853 is no longer needed.** `104a848` (#783 PR-g) already templates `gr_up_multi_kernel<MAX_T, EXACT_T>` and dispatches exact `T` for `ct 1..6` (`src/kernels/cuda/fused_gr.cu:907-919`, `:891-896`). Nothing here changes a kernel.

**What is still missing is the test.** `gr_parity` checked the fused multi-token read only at `kFusedGrMaxT = 8` - the one launch that requests the whole dynamic LDS. The exact-`T` instantiations the engine actually launches were never compared with the single-token reference.

`fused_multi_lds_parity` now takes `T` and the selftest runs it at 8, 4 and 2: the full 40 KiB LDS request, and two short windows through the exact-`T` dispatch. The `q8_1` (`STRATA_QFUSE`) check stays at the full window, where its group counters are exercised at the full size.

Measured on an RTX 3060 12 GB (driver 580.95.05), built from `82f46a8`:

before, `gr_parity --selftest`:
```
  fused GR multi max-T=8 LDS launch and changing graph replay pass
```
after:
```
  fused GR multi T=8 LDS launch and changing graph replay pass
  fused GR multi T=4 LDS launch and changing graph replay pass
  fused GR multi T=2 LDS launch and changing graph replay pass
  ...
gr_read/gr_write: 0 failures
gr_parity OK
```

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。