Pull requests / #883

#883 sm_60: PTX vmad for the __dp4a fallback (bit-exact, ~2.2x in isolation)

closed · @willowmerestudios · 0 Kommentare · Auf GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants

Beschreibung

## sm_60: PTX `vmad` for the `__dp4a` fallback (bit-exact, ~2.2-2.3x in isolation)

### What
On compute capability 6.0 (Tesla P100 / GP100) there is no `__dp4a`, so `STRATA_DP4A` falls back to
`strata_dp4a()` in `include/strata/kernels/dp4a.hpp`. Today that is the byte-wise C form: four sign-extended
byte extractions and four multiplies. This PR replaces the body with four PTX `vmad.s32.s32.s32` instructions
using byte selectors (`%1.b0` … `%1.b3`), which do the signed-byte multiply-accumulate directly: `cuobjdump -sass`
shows four `VMAD.S8.S8` per call and no byte extraction.

### Scope
Only the `__CUDA_ARCH__ < 610` branch changes. sm_61+ still uses hardware `__dp4a`; HIP builds never take
this branch (`__CUDA_ARCH__` undefined). No call sites change.

### Proof of bit-exactness
New `tools/sm60_dp4a_check.cu` includes the committed header and compares `STRATA_DP4A` against the byte-wise
reference over 1,073,741,824 `(a, b, c)` triples (exit 1 on any mismatch or CUDA error):
- every pairing of the edge bytes 0x7f, 0x80, 0xff and 0x00 in `a` and `b`;
- exhaustive per lane: every (byte lane, `a` byte, `b` byte) combination, 4 × 256 × 256, with random other lanes;
- the rest pseudo-random (splitmix32 per thread).

No CMake target; build by hand:

```
nvcc -O3 -arch=sm_60 -Iinclude tools/sm60_dp4a_check.cu -o sm60_dp4a_check && ./sm60_dp4a_check
```

Output on a Tesla P100-PCIE-16GB (CUDA 12.4), 3 runs:

| run | mismatches | byte loop | `STRATA_DP4A` | speedup |
|---|---|---|---|---|
| 1 | 0 / 1,073,741,824 | 92.3 ms | 39.3 ms | 2.35x |
| 2 | 0 | 84.4 ms | 37.3 ms | 2.26x |
| 3 | 0 | 84.4 ms | 37.1 ms | 2.27x |

(Sanity check of the checker: with the reference deliberately off by one it reports 1,073,741,824 mismatches, exit 1.)

### End-to-end
Measured on v0.1.38 (`99f3dbd`), not re-run on this branch's base (v0.1.39); the commit applies to both unchanged.
Same tree built twice, with and without this change. 3 interleaved arms each on one P100,
Qwen3.8-Flash-Next Coder IQ1_M, temperature 0, 512-token decode ×3 + 8K-token prefill ×3 per arm.

| | decode tok/s (mean of runs 2-3, per arm) | prefill tok/s (run 3, per arm) | greedy output sha256 |
|---|---|---|---|
| byte-wise (stock) | 21.05 / 20.85 / 21.25 (mean 21.05) | 320.3 / 321.5 / 321.3 | same in all 6 arms |
| vmad | 22.10 / 21.75 / 20.55 (mean 21.47) | 320.8 / 321.0 / 322.9 | same in all 6 arms |

Honest read: on this model the end-to-end difference (+2% decode, prefill flat) is inside run-to-run noise, so
the gain here is kernel-level, not a headline tok/s number. The useful end-to-end result is correctness: for each
of the 6 temperature-0 requests per arm (two prompts × 3 runs), the sha256 of the output text (compared on the first
12 hex digits) is the same in all six arms, stock and vmad. Workloads that
spend a larger share of time in `STRATA_DP4A` would see more of it (not measured here).

### Build check
Full engine build at this branch with `-DCMAKE_CUDA_ARCHITECTURES=60 -DSTRATA_EXPERIMENTAL_SM60=ON
-DSTRATA_ENABLE_CUDA=ON`, nvcc 12.4: exit 0, no errors.

Mehr auf der Site

Links zu Install, Modellen, Releases.