Pull requests / #883
#883 sm_60: PTX vmad for the __dp4a fallback (bit-exact, ~2.2x in isolation)
closed · @willowmerestudios · 0 commentaires · Sur GitHub
BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
Description
## sm_60: PTX `vmad` for the `__dp4a` fallback (bit-exact, ~2.2-2.3x in isolation) ### What On compute capability 6.0 (Tesla P100 / GP100) there is no `__dp4a`, so `STRATA_DP4A` falls back to `strata_dp4a()` in `include/strata/kernels/dp4a.hpp`. Today that is the byte-wise C form: four sign-extended byte extractions and four multiplies. This PR replaces the body with four PTX `vmad.s32.s32.s32` instructions using byte selectors (`%1.b0` … `%1.b3`), which do the signed-byte multiply-accumulate directly: `cuobjdump -sass` shows four `VMAD.S8.S8` per call and no byte extraction. ### Scope Only the `__CUDA_ARCH__ < 610` branch changes. sm_61+ still uses hardware `__dp4a`; HIP builds never take this branch (`__CUDA_ARCH__` undefined). No call sites change. ### Proof of bit-exactness New `tools/sm60_dp4a_check.cu` includes the committed header and compares `STRATA_DP4A` against the byte-wise reference over 1,073,741,824 `(a, b, c)` triples (exit 1 on any mismatch or CUDA error): - every pairing of the edge bytes 0x7f, 0x80, 0xff and 0x00 in `a` and `b`; - exhaustive per lane: every (byte lane, `a` byte, `b` byte) combination, 4 × 256 × 256, with random other lanes; - the rest pseudo-random (splitmix32 per thread). No CMake target; build by hand: ``` nvcc -O3 -arch=sm_60 -Iinclude tools/sm60_dp4a_check.cu -o sm60_dp4a_check && ./sm60_dp4a_check ``` Output on a Tesla P100-PCIE-16GB (CUDA 12.4), 3 runs: | run | mismatches | byte loop | `STRATA_DP4A` | speedup | |---|---|---|---|---| | 1 | 0 / 1,073,741,824 | 92.3 ms | 39.3 ms | 2.35x | | 2 | 0 | 84.4 ms | 37.3 ms | 2.26x | | 3 | 0 | 84.4 ms | 37.1 ms | 2.27x | (Sanity check of the checker: with the reference deliberately off by one it reports 1,073,741,824 mismatches, exit 1.) ### End-to-end Measured on v0.1.38 (`99f3dbd`), not re-run on this branch's base (v0.1.39); the commit applies to both unchanged. Same tree built twice, with and without this change. 3 interleaved arms each on one P100, Qwen3.8-Flash-Next Coder IQ1_M, temperature 0, 512-token decode ×3 + 8K-token prefill ×3 per arm. | | decode tok/s (mean of runs 2-3, per arm) | prefill tok/s (run 3, per arm) | greedy output sha256 | |---|---|---|---| | byte-wise (stock) | 21.05 / 20.85 / 21.25 (mean 21.05) | 320.3 / 321.5 / 321.3 | same in all 6 arms | | vmad | 22.10 / 21.75 / 20.55 (mean 21.47) | 320.8 / 321.0 / 322.9 | same in all 6 arms | Honest read: on this model the end-to-end difference (+2% decode, prefill flat) is inside run-to-run noise, so the gain here is kernel-level, not a headline tok/s number. The useful end-to-end result is correctness: for each of the 6 temperature-0 requests per arm (two prompts × 3 runs), the sha256 of the output text (compared on the first 12 hex digits) is the same in all six arms, stock and vmad. Workloads that spend a larger share of time in `STRATA_DP4A` would see more of it (not measured here). ### Build check Full engine build at this branch with `-DCMAKE_CUDA_ARCHITECTURES=60 -DSTRATA_EXPERIMENTAL_SM60=ON -DSTRATA_ENABLE_CUDA=ON`, nvcc 12.4: exit 0, no errors.
Sur le site
Liens install, modèles, releases.