Pull requests / #186
#186 decode: hyper-connection read in 2 kernels, stream-split (-1.5 ms per verify window)
closed · @q8atnight · 0 comentários · No GitHub
Multi-GPUNVIDIA / CUDAModels & quants
Descrição
## Summary In `fused_gr_read_multi` the norm kernel ran on only T blocks (~16 us of pure latency per call, 96 calls per verify window) and `down` on 41 blocks (~280 GB/s). v3 does the hyper-connection read in 2 kernels instead of 3: - `down` is split over (row group, stream) = 164 blocks; each block stages its stream's R' * w_norm for the T tokens, reduces that stream's sum of squares itself, and writes UNSCALED per-stream partial dots - since the rms scale is per stream, w_down . xn = sum_c rs[c] * (w_down[:, c] . (R'[c] * w_norm[c])), and `up` applies it in its prologue (lo, inject, rs) One commit on top of 0.1.27 (a790805); all four of my current branches were compile-tested together on 0.1.27. ## Measured Measured on 2x RTX 3090, IQ3_S, with the engine this was developed in (0.1.24 + the dual-GPU work): decode verify window 23.25 -> 21.69 ms (about +7 % decode). ## Correctness Same maths, another summation order, so NOT bitwise identical to the old kernels. Our dual-vs-single exactness gate passes with it (both paths use it). ## Switches `STRATA_GR_V3=0` = the old 3 kernels. A/B knobs measured neutral or slower, left off by default: `STRATA_GR_KSPLIT=2`, `STRATA_GR_DOWN_R2=2`, `STRATA_GR_UP_PF=1`.
No site
Links install, modelos, releases.