Pull requests / #187
#187 decode: QSA block scores read each key block once per window (bit-identical)
closed · @q8atnight · 0 comentarios · En GitHub
Multi-GPUNVIDIA / CUDAModels & quants
Descripción
## Summary In a captured decode window `qsa_block_scores` launched a (max_blocks / 8) x nq grid (~24,600 mostly idle blocks per layer) and re-read every pooled key block once per query. New kernel: a fixed grid strides over the key blocks, reads each once and scores it for all of the window's queries (up to 8); per (block, query) the same arithmetic in the same order as `block_scores_kernel`. Used only when no active-block count is given (the captured decode window). One commit on top of 0.1.27 (a790805); all four of my current branches were compile-tested together on 0.1.27. ## Measured Measured on 2x RTX 3090, IQ3_S, with the engine this was developed in (0.1.24 + the dual-GPU work): at 32K context the "scores+topk" stage 1.02 -> 0.90 ms per window; verify window 20.28 -> 20.04 ms. ## Correctness Bit-identical scores. ## Switch `STRATA_SCORES_MULTI=0` = the old grid.
En el sitio
Enlaces a install, modelos, releases.