Pull requests / #1096

#1096 qsa_prompt_attn: the Volta kernel's latency pipeline (32K int8 12.1 -> 10.2 ms, +1.2% e2e)

open · @ATIVX928 · 0 comments · View on GitHub

BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDA

Description

> Resubmitted from #1020: the original was auto-closed by the 2026-10-06 history cleanup. Rebased onto the new `main` (82f46a8); build and parity re-verified on 2x V100.

## What

Main already runs the Volta prompt attention on `prompt_attn_v70_kernel` (#600). This adds the kernel's latency pipeline: the pool-row lookups are double-buffered (the next chunk's rows are prefetched under the previous chunk's p.v), and a thread's staging loads are batched before its stores. No arithmetic, layout or output changes - the sums are formed in the same order.

This is the remaining piece of the Volta prompt attention in #627 (the base kernel was already taken from #600; only the pipeline commit is left).

## Gating

All changes are inside the existing `#if !defined(__HIPCC__) && defined(STRATA_EXPERIMENTAL_SM60)` block in `src/kernels/cuda/qsa_prompt_attn.cu`; the ready-made engine is untouched.

## Measured

V100-SXM2-16GB, CUDA 12.8, `qsa_prompt_attn_sm70_bench`, 2,048 queries of a 32K selection, 5 reps:

| KV | before | after | vs FP32 before -> after |
| --- | ---: | ---: | ---: |
| int8 | 12.111 ms | 10.247 ms | 1.73x -> **2.05x** |
| fp16 | 25.700 ms | 17.390 ms | 0.80x -> 1.18x |
| K8V4 | 25.713 ms | 24.119 ms | 0.85x -> 0.91x |

## Tests

- `qsa_prompt_attn_parity 32768 2048 5`: PASS, 0 failures. Every FP64 and new-vs-old figure is identical before/after the patch (int8 32768: new vs old 1.25e-05, 3.3e-06 of the output scale): the port is numerically bit-identical.
- The bench's K8V4 `max |old-new|` is unchanged and covers mode 3, which the parity test skips on sm_70.

## 32K end-to-end

Same setup. Interleaved A/B (baseline / this branch, 5 rounds each): on the rebased branch against the new main, dual-card prefill 2219.9 -> 2245.5 tok/s (+1.2%); the same A/B on the pre-cleanup main read 2191.6 -> 2225.2 (+1.5%), with single-card 1557.1 -> 1575.1 (+1.2%). The branch repeats within 0.2% across sessions, and so does the baseline. Decode is unchanged (within its +-15% session noise).

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.