Pull requests / #1091
#1091 qsa_select: FP32 tiled block scores below sm_80 (+22% 128K / +44% 248K prefill on a 2080 Ti)
closed · @konijiwa110 · 0 commentaires · Sur GitHub
BenchmarksNVIDIA / CUDAModels & quants
Description
Reopening #742 here. It was auto-closed by the force-push of `main`. Same commit, cherry-picked onto the new `main` (82f46a8) without conflicts. I re-measured on that alone, without any of my other PRs. #741 isn't coming back since `STRATA_BF16_TC=1` covers it now. ## What Below sm_80 there is no TF32 tensor-core block scorer, so the prompt path's QSA selection scores blocks on the warp kernel (4 products per lane, 5 shuffles): ~1 TFLOPS on an RTX 2080 Ti, 23-37% of a 128K-248K prompt there. `qsa_block_scores_tc` now runs an FP32 tiled GEMM on those cards (`block_scores_simt_kernel`): rows = query x indexer head, K = 128 through shared memory, each thread 2 queries x 4 heads x 4 blocks, relu per head summed in registers. ~6.9 TFLOPS. The tail block keeps the tail kernel's arithmetic. Like the tensor-core scorer it is FP32 in another summation order, so not bitwise equal to the warp kernel. `STRATA_SELECT_SIMT=0` (or the existing `STRATA_SELECT_OLD=1`) keeps the warp kernel. `qsa_select_bench` is now also registered as a CUDA test: at 32K, and at 131K with `STRATA_QSA_WARP=1` so the FP32 tiled scorer is exercised on any card. Gate: error vs an FP64 reference no worse than 4x the warp kernel's. ## Measured RTX 2080 Ti, this branch, `qsa_select_bench` (256 queries, capacity 262144): | context | warp kernel | FP32 tiled | |---|---|---| | 32K | 2.110 ms | 0.329 ms (6.4x) | | 131K | 8.516 ms | 1.253 ms (6.8x) | Selections identical to the warp kernel's 256/256. FP64 error 5.6e-5 on a score scale of ~268. End to end: Swift 1.5 IQ3_XXS, 256K context, `--kv k8v4 --kv-resident 32768`, the same build with `STRATA_SELECT_SIMT=0` / unset. Prefill tok/s: | prompt | warp | FP32 tiled | `qsa select` | |---|---:|---:|---| | 4K | 972 | 984 | 38 -> 21 ms | | 32K | 1173 | 1221 (+4%) | 1.7 -> 0.5 s | | 128K | 1003 | 1228 (+22%) | 29.4 -> 5.7 s | | 248K | 802 | 1151 (+44%) | 113.4 -> 20.9 s | Decode is unchanged. 248K gains more than in #742 (+33% there) because the 0.1.40 top-k change made the scores the larger part of `qsa select`. `qsa_topk_active_parity` and both `qsa_select_bench` runs pass on this branch. Only tested on sm_75. Developed with an AI coding assistant; all numbers measured on an RTX 2080 Ti 22 GB / i5-12490F / 64 GB. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sur le site
Liens install, modèles, releases.