Pull requests / #1170
#1170 test(gr): check the fused multi-token read at T=4 and T=2
open · @InB4DevOps · 0 comentarios · En GitHub
Descripción
Follow-up to #853, which was closed when `main` was force-pushed. **The kernel half of #853 is no longer needed.** `104a848` (#783 PR-g) already templates `gr_up_multi_kernel<MAX_T, EXACT_T>` and dispatches exact `T` for `ct 1..6` (`src/kernels/cuda/fused_gr.cu:907-919`, `:891-896`). Nothing here changes a kernel. **What is still missing is the test.** `gr_parity` checked the fused multi-token read only at `kFusedGrMaxT = 8` - the one launch that requests the whole dynamic LDS. The exact-`T` instantiations the engine actually launches were never compared with the single-token reference. `fused_multi_lds_parity` now takes `T` and the selftest runs it at 8, 4 and 2: the full 40 KiB LDS request, and two short windows through the exact-`T` dispatch. The `q8_1` (`STRATA_QFUSE`) check stays at the full window, where its group counters are exercised at the full size. Measured on an RTX 3060 12 GB (driver 580.95.05), built from `82f46a8`: before, `gr_parity --selftest`: ``` fused GR multi max-T=8 LDS launch and changing graph replay pass ``` after: ``` fused GR multi T=8 LDS launch and changing graph replay pass fused GR multi T=4 LDS launch and changing graph replay pass fused GR multi T=2 LDS launch and changing graph replay pass ... gr_read/gr_write: 0 failures gr_parity OK ```
En el sitio
Enlaces a install, modelos, releases.