Pull requests / #1095

#1095 kv_q8: the gather decodes 8 int8 per thread (V100 experimental build, +11%)

open · @ATIVX928 · 0 Kommentare · Auf GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentation

Beschreibung

> Resubmitted from #1019: the original was auto-closed by the 2026-10-06 history cleanup. Rebased onto the new `main` (82f46a8); build and parity re-verified on 2x V100.

## What

The int8 KV gather (`kv_gather_q8_kernel`, used by the 32K prompt path) decodes four int8 codes per thread. This decodes eight, with `uint2`/`uint4` loads and stores, and falls back to the four-wide path when the alignment does not allow the wider access. It is the int8-gather piece of the V100 work in #627, split out on its own.

## Gating

The wide kernel is compiled only under `-DSTRATA_EXPERIMENTAL_SM60=ON` and only launched on a device below cc 7.5 (`pre75_q8_gather()`), with `STRATA_Q8_GATHER_WIDE=0` as the A/B arm. The ready-made engine compiles the original kernel; as a check, the non-experimental object's SASS sha256 is identical to main's (`a4af40d8…`, 54376 bytes).

## Measured

V100-SXM2-16GB, CUDA 12.8, `kv_q8_parity --bench`, 2,051 selected cells at 32K:

| path | time | bandwidth |
| --- | ---: | ---: |
| 4-wide (main) | 15.9 us | 399.8 GB/s |
| 8-wide (this PR) | 14.2 us | 449.4 GB/s |

+11.3% on the gather kernel.

## 32K end-to-end

2x V100-SXM2-16GB, 32,504-token prompt, int8 KV, MTP spec 4, expert-cache auto, prompt cache off; median of 5 rounds after a warm-up. Single-card prefill 1550.4 tok/s and dual-card 2203.0 vs main's 1557.1 / 2207.0 - within the +-1% session noise. The gather is a few tens of microseconds of a 21-second prefill, so the +11.3% stays a kernel-level gain.

## Tests

- `kv_q8_parity` PASS (worst INT8-vs-FP16 error 0.59 quantization steps), both the wide default and `STRATA_Q8_GATHER_WIDE=0`.

## Notes

`bench/results/2026-10-03-v100-kv-int8-read/README.md` records the measurement and the footprint/bandwidth comparison of int8 vs q4_0 vs k8v4.

Mehr auf der Site

Links zu Install, Modellen, Releases.