Pull requests / #1019
#1019 kv_q8: the gather decodes 8 int8 per thread (V100 experimental build, +11%)
closed · @ATIVX928 · 0 comments · View on GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentation
Description
## What The int8 KV gather (`kv_gather_q8_kernel`, used by the 32K prompt path) decodes four int8 codes per thread. This decodes eight, with `uint2`/`uint4` loads and stores, and falls back to the four-wide path when the alignment does not allow the wider access. It is the int8-gather piece of the V100 work in #627, split out on its own. ## Gating The wide kernel is compiled only under `-DSTRATA_EXPERIMENTAL_SM60=ON` and only launched on a device below cc 7.5 (`pre75_q8_gather()`), with `STRATA_Q8_GATHER_WIDE=0` as the A/B arm. The ready-made engine compiles the original kernel; as a check, the non-experimental object's SASS sha256 is identical to main's (`a4af40d8…`, 54376 bytes). ## Measured V100-SXM2-16GB, CUDA 12.8, `kv_q8_parity --bench`, 2,051 selected cells at 32K: | path | time | bandwidth | | --- | ---: | ---: | | 4-wide (main) | 15.9 us | 399.8 GB/s | | 8-wide (this PR) | 14.2 us | 449.4 GB/s | +11.3% on the gather kernel. ## 32K end-to-end 2x V100-SXM2-16GB, 32,504-token prompt, int8 KV, MTP spec 4, expert-cache auto, prompt cache off; median of 5 rounds after a warm-up. Single-card prefill 1550.4 tok/s and dual-card 2203.0 vs main's 1557.1 / 2207.0 - within the +-1% session noise. The gather is a few tens of microseconds of a 21-second prefill, so the +11.3% stays a kernel-level gain. ## Tests - `kv_q8_parity` PASS (worst INT8-vs-FP16 error 0.59 quantization steps), both the wide default and `STRATA_Q8_GATHER_WIDE=0`. ## Notes `bench/results/2026-10-03-v100-kv-int8-read/README.md` records the measurement and the footprint/bandwidth comparison of int8 vs q4_0 vs k8v4.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.