Pull requests / #185
#185 sampler: the sampled path over 64 vocabulary chunks per token (identical picks, ~12% faster sampled decode)
closed · @q8atnight · 0 comentários · No GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quants
Descrição
## Summary The sampled path in `sampler.cu` put ONE block per token over the 248,320-logit vocabulary and ran top_k (20) argmax rounds, each a full scan with an inner "already taken?" loop: ~3 ms per sampled verify window on 3 of 82 SMs. This splits the same selection into two kernels: - kernel 1: the vocabulary in 64 chunks per token, each chunk block keeps its own top-k (k rounds over its ~3,900 values in shared memory, a taken value set to -inf) - kernel 2: one block per token merges the chunks' candidates with k more rounds, then runs the unchanged top_p / min_p / temperature / Philox chain, line for line The global top-k in (value desc, index asc) order is the union of the chunks' top-k's, so the kept list - set AND order - is exactly the serial one. One commit on top of 0.1.27 (a790805); all four of my current branches were compile-tested together on 0.1.27. ## Measured Measured on 2x RTX 3090, IQ3_S, with the engine this was developed in (0.1.24 + the dual-GPU work): sampled decode (temperature 0.7, top_p 0.95, top_k 20) at 1K context 102 -> 114 tok/s; about 3 ms less per sampled verify window. Greedy is unaffected (it does not go through this path). ## Correctness `STRATA_SAMPLER_CHECK=1` runs the old one-block kernel beside the new one and compares every pick: 1,001 of 1,001 identical. ## Switch `STRATA_SAMPLER_FAST=0` = the old one-block kernel.
No site
Links install, modelos, releases.