Pull requests / #188

#188 prefill: software-pipelined loads in the column-split GDN recurrence (bit-identical)

closed · @q8atnight · 0 commentaires · Sur GitHub

Multi-GPUNVIDIA / CUDAModels & quants

Description

## Summary

The port Niko mentioned when closing #108 ("your software-pipelined loads in the recurrence go further than ours
(-29 % vs our -17 %)"): `gdn_rec_cols_kernel` with the next token's q/k rows, v, gate and beta loaded into registers
while the current token computes. Templated on columns per block (32 = upstream's split; 16 available via
`STRATA_GDN_CB=16`, measured neutral).

One commit on top of 0.1.27 (a790805); all four of my current branches were compile-tested together on 0.1.27.

## Measured

Measured on 2x RTX 3090, IQ3_S, with the engine this was developed in (0.1.24 + the dual-GPU work): recurrence time
-29 % vs -17 % for the column split alone (as quoted in #108); prompt speed at 32K within a percent or two overall.

## Correctness

The same arithmetic in the same order: bit-identical.

## Switch

`STRATA_GDN_PIPELINE=0` = upstream's kernel.

Sur le site

Liens install, modèles, releases.