Pull requests / #1091

#1091 qsa_select: FP32 tiled block scores below sm_80 (+22% 128K / +44% 248K prefill on a 2080 Ti)

closed · @konijiwa110 · 0 评论 · 在 GitHub 查看

BenchmarksNVIDIA / CUDAModels & quants

描述

Reopening #742 here. It was auto-closed by the force-push of `main`. Same commit, cherry-picked onto the new `main` (82f46a8) without conflicts. I re-measured on that alone, without any of my other PRs. #741 isn't coming back since `STRATA_BF16_TC=1` covers it now.

## What

Below sm_80 there is no TF32 tensor-core block scorer, so the prompt path's QSA selection scores blocks on the warp
kernel (4 products per lane, 5 shuffles): ~1 TFLOPS on an RTX 2080 Ti, 23-37% of a 128K-248K prompt there.

`qsa_block_scores_tc` now runs an FP32 tiled GEMM on those cards (`block_scores_simt_kernel`): rows = query x indexer
head, K = 128 through shared memory, each thread 2 queries x 4 heads x 4 blocks, relu per head summed in registers.
~6.9 TFLOPS. The tail block keeps the tail kernel's arithmetic. Like the tensor-core scorer it is FP32 in another
summation order, so not bitwise equal to the warp kernel.

`STRATA_SELECT_SIMT=0` (or the existing `STRATA_SELECT_OLD=1`) keeps the warp kernel.

`qsa_select_bench` is now also registered as a CUDA test: at 32K, and at 131K with `STRATA_QSA_WARP=1` so the FP32
tiled scorer is exercised on any card. Gate: error vs an FP64 reference no worse than 4x the warp kernel's.

## Measured

RTX 2080 Ti, this branch, `qsa_select_bench` (256 queries, capacity 262144):

| context | warp kernel | FP32 tiled |
|---|---|---|
| 32K | 2.110 ms | 0.329 ms (6.4x) |
| 131K | 8.516 ms | 1.253 ms (6.8x) |

Selections identical to the warp kernel's 256/256. FP64 error 5.6e-5 on a score scale of ~268.

End to end: Swift 1.5 IQ3_XXS, 256K context, `--kv k8v4 --kv-resident 32768`, the same build with
`STRATA_SELECT_SIMT=0` / unset. Prefill tok/s:

| prompt | warp | FP32 tiled | `qsa select` |
|---|---:|---:|---|
| 4K | 972 | 984 | 38 -> 21 ms |
| 32K | 1173 | 1221 (+4%) | 1.7 -> 0.5 s |
| 128K | 1003 | 1228 (+22%) | 29.4 -> 5.7 s |
| 248K | 802 | 1151 (+44%) | 113.4 -> 20.9 s |

Decode is unchanged. 248K gains more than in #742 (+33% there) because the 0.1.40 top-k change made the scores the
larger part of `qsa select`.

`qsa_topk_active_parity` and both `qsa_select_bench` runs pass on this branch.

Only tested on sm_75.

Developed with an AI coding assistant; all numbers measured on an RTX 2080 Ti 22 GB / i5-12490F / 64 GB.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。