Pull requests / #185

#185 sampler: the sampled path over 64 vocabulary chunks per token (identical picks, ~12% faster sampled decode)

closed · @q8atnight · 0 comentários · No GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quants

Descrição

## Summary

The sampled path in `sampler.cu` put ONE block per token over the 248,320-logit vocabulary and ran top_k (20) argmax
rounds, each a full scan with an inner "already taken?" loop: ~3 ms per sampled verify window on 3 of 82 SMs.
This splits the same selection into two kernels:

- kernel 1: the vocabulary in 64 chunks per token, each chunk block keeps its own top-k (k rounds over its ~3,900
  values in shared memory, a taken value set to -inf)
- kernel 2: one block per token merges the chunks' candidates with k more rounds, then runs the unchanged
  top_p / min_p / temperature / Philox chain, line for line

The global top-k in (value desc, index asc) order is the union of the chunks' top-k's, so the kept list - set AND
order - is exactly the serial one.

One commit on top of 0.1.27 (a790805); all four of my current branches were compile-tested together on 0.1.27.

## Measured

Measured on 2x RTX 3090, IQ3_S, with the engine this was developed in (0.1.24 + the dual-GPU work): sampled decode
(temperature 0.7, top_p 0.95, top_k 20) at 1K context 102 -> 114 tok/s; about 3 ms less per sampled verify window.
Greedy is unaffected (it does not go through this path).

## Correctness

`STRATA_SAMPLER_CHECK=1` runs the old one-block kernel beside the new one and compares every pick:
1,001 of 1,001 identical.

## Switch

`STRATA_SAMPLER_FAST=0` = the old one-block kernel.

No site

Links install, modelos, releases.