Pull requests / #1020
#1020 qsa_prompt_attn: the Volta kernel's latency pipeline (32K int8 12.1 -> 10.2 ms, +1.5% e2e)
closed · @ATIVX928 · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installAMD / HIPNVIDIA / CUDA
描述
## What Main already runs the Volta prompt attention on `prompt_attn_v70_kernel` (#600). This adds the kernel's latency pipeline: the pool-row lookups are double-buffered (the next chunk's rows are prefetched under the previous chunk's p.v), and a thread's staging loads are batched before its stores. No arithmetic, layout or output changes - the sums are formed in the same order. This is the remaining piece of the Volta prompt attention in #627 (the base kernel was already taken from #600; only the pipeline commit is left). ## Gating All changes are inside the existing `#if !defined(__HIPCC__) && defined(STRATA_EXPERIMENTAL_SM60)` block in `src/kernels/cuda/qsa_prompt_attn.cu`; the ready-made engine is untouched. ## Measured V100-SXM2-16GB, CUDA 12.8, `qsa_prompt_attn_sm70_bench`, 2,048 queries of a 32K selection, 5 reps: | KV | before | after | vs FP32 before -> after | | --- | ---: | ---: | ---: | | int8 | 12.111 ms | 10.247 ms | 1.73x -> **2.05x** | | fp16 | 25.700 ms | 17.390 ms | 0.80x -> 1.18x | | K8V4 | 25.713 ms | 24.119 ms | 0.85x -> 0.91x | ## Tests - `qsa_prompt_attn_parity 32768 2048 5`: PASS, 0 failures. Every FP64 and new-vs-old figure is identical before/after the patch (int8 32768: new vs old 1.25e-05, 3.3e-06 of the output scale): the port is numerically bit-identical. - The bench's K8V4 `max |old-new|` is unchanged and covers mode 3, which the parity test skips on sm_70. ## 32K end-to-end Same setup. Interleaved A/B (baseline / this branch, 5 rounds each): dual-card prefill 2191.6 -> 2225.2 tok/s (+1.5%), single-card 1557.1 -> 1575.1 (+1.2%). The branch repeats within 0.2% across sessions, and so does the baseline. Decode is unchanged (within its +-15% session noise).
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。