Pull requests / #187

#187 decode: QSA block scores read each key block once per window (bit-identical)

closed · @q8atnight · 0 commentaires · Sur GitHub

Multi-GPUNVIDIA / CUDAModels & quants

Description

## Summary

In a captured decode window `qsa_block_scores` launched a (max_blocks / 8) x nq grid (~24,600 mostly idle blocks per
layer) and re-read every pooled key block once per query. New kernel: a fixed grid strides over the key blocks, reads
each once and scores it for all of the window's queries (up to 8); per (block, query) the same arithmetic in the same
order as `block_scores_kernel`. Used only when no active-block count is given (the captured decode window).

One commit on top of 0.1.27 (a790805); all four of my current branches were compile-tested together on 0.1.27.

## Measured

Measured on 2x RTX 3090, IQ3_S, with the engine this was developed in (0.1.24 + the dual-GPU work): at 32K context the
"scores+topk" stage 1.02 -> 0.90 ms per window; verify window 20.28 -> 20.04 ms.

## Correctness

Bit-identical scores.

## Switch

`STRATA_SCORES_MULTI=0` = the old grid.

Sur le site

Liens install, modèles, releases.