Pull requests / #270

#270 qsa_prompt_attn: run the f16 tensor-core path on Turing (sm_75)

closed · @kenh0u · 0 comments · View on GitHub

BenchmarksNVIDIA / CUDA

Description

## What this does

`mma.m16n8k16` (f16) needs sm_80, so main compiles the tensor-core prompt attention to a trap on sm_75 and those cards run the fp32-FMA kernel — the gap #87 left in ("the sm_80 guard on the tensor-core prompt attention"). Turing (sm_75) has `mma.m16n8k8` with the **same A/B/C register mapping**, so each k=16 step becomes two k=8 steps on the fragments exactly as they are already laid out (a[0]/a[1] with b[0], a[2]/a[3] with b[1]). The sum rounds twice, so the two paths do not agree bit for bit. Turing therefore takes the v1 kernel, which has no cp.async dependency.

## Measured (Quadro RTX 8000, CUDA 12.9, arch 75)

- `qsa_prompt_attn_parity` (`-DCMAKE_CUDA_ARCHITECTURES=75`): **PASS, FAILURES: 0**; vs FP64 new 6.99e-06, old 2.37e-06 (output scale 3.62); int8 kernel 62.3 → 19.6 ms per chunk (**3.19x**), fp16 65.0 → 25.1 ms (2.59x)
- Engine, three Turing cards (RTX 8000 + 2× RTX 6000), on 0.1.29 main: 9.7K-token prompt **1,050 → 1,177 tok/s**, 41.7K **1,124 → 1,251 tok/s**; an 80K needle test answers correctly
- Single-card relevance: #139's thread measured main 0.1.27 at PP 731-738 tok/s on one RTX 2080 Ti — this kernel is that card's prompt path too

## AI tooling
Developed with Claude Code (GLM-5.3-Flash): implementation, sm_75 parity runs and the measurements above.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.