Pull requests / #186

#186 decode: hyper-connection read in 2 kernels, stream-split (-1.5 ms per verify window)

closed · @q8atnight · 0 コメント · GitHub で見る

Multi-GPUNVIDIA / CUDAModels & quants

本文

## Summary

In `fused_gr_read_multi` the norm kernel ran on only T blocks (~16 us of pure latency per call, 96 calls per verify
window) and `down` on 41 blocks (~280 GB/s). v3 does the hyper-connection read in 2 kernels instead of 3:

- `down` is split over (row group, stream) = 164 blocks; each block stages its stream's R' * w_norm for the T tokens,
  reduces that stream's sum of squares itself, and writes UNSCALED per-stream partial dots
- since the rms scale is per stream, w_down . xn = sum_c rs[c] * (w_down[:, c] . (R'[c] * w_norm[c])), and `up`
  applies it in its prologue (lo, inject, rs)

One commit on top of 0.1.27 (a790805); all four of my current branches were compile-tested together on 0.1.27.

## Measured

Measured on 2x RTX 3090, IQ3_S, with the engine this was developed in (0.1.24 + the dual-GPU work): decode verify
window 23.25 -> 21.69 ms (about +7 % decode).

## Correctness

Same maths, another summation order, so NOT bitwise identical to the old kernels. Our dual-vs-single exactness gate
passes with it (both paths use it).

## Switches

`STRATA_GR_V3=0` = the old 3 kernels. A/B knobs measured neutral or slower, left off by default:
`STRATA_GR_KSPLIT=2`, `STRATA_GR_DOWN_R2=2`, `STRATA_GR_UP_PF=1`.

関連リンク

インストール・モデル・リリースへの站内リンク。