Pull requests / #1279
#1279 perf(cuda): use exact GP100 VMAD for native DP4A emulation
closed · @bschrib · 0 comments · View on GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Description
Tesla P100's sm_60 native DP4A fallback extracts and multiplies four signed bytes in scalar code. This uses four exact integer VMAD operations on GP100. Other CUDA architectures and HIP retain their existing source paths. The pinned llama.cpp MMQ helper is unchanged. The original optimization idea and P100 work belong to **[shinbunbun](https://github.com/shinbunbun)**, specifically [llama-cpp-p100-patches patch 01](https://github.com/shinbunbun/llama-cpp-p100-patches/blob/939dd95f9f926c87769bd0442fbcf8a40aff2597/patches/01-vmad-dp4a-sm60.patch). This submission adapts that idea to Strata, adds a scoped accumulator to avoid output/input register aliasing, and supplies focused tests and Strata measurements. The original MIT notice is retained. Validation: - Both P100s passed 1,813,290 exact arithmetic cases; Compute Sanitizer reported zero errors. Native IQ/MMVQ fixtures and all 48 real expert layers passed their existing parity tests. - Offline widened/modulo oracle and defined scalar UBSan checks passed. Non-sm_60 preprocessed paths match; CUDA 12.9 sm_61/sm_75 probe SASS matches control exactly. No modern GPU device execution was performed. - At 119,628 actual input tokens, decode median changed **36.55 to 40.00 tok/s, +9.44%**. A separate 19,790-token corpus showed +10.17%. Prompt throughput was effectively unchanged; all text/reasoning matched across twelve fresh-engine ABBA trials per corpus. Measurements used pinned Strata v0.1.39 on two PCIe P100 16 GB cards at Gen3 x8/x8, Ryzen 7900X, Ubuntu 24.04.5, R580 and CUDA 12.9.86. Both arms used 128K context, INT8 KV with 32K resident, split25/23, MTP spec4 and zero prompt-cache reuse. This PR targets newer main with the same pristine DP4A header; the full current-main engine was not built or benchmarked. Its header passed separate CUDA compile-only probes. Eight API prompt pairs passed full 248,320-value first-position bitwise logit parity and matching text/reasoning/usage. These checks used separately identified diagnostic twins with identical API instrumentation, relinked with unchanged original kernel libraries. The existing standalone dump hook does not cover the API; the rejected missing-dump attempt is retained separately. This does not establish every generated-position logit or general quantization quality. Instantaneous power excursions above configured/enforced200W are retained in the evidence. Mechanism, commands, limits, source/build/model identities and raw timing rows are in [docs/P100_VMAD.md](https://github.com/bschrib/Strata/blob/perf/p100-vmad-dp4a/docs/P100_VMAD.md) and its linked benchmark JSON files. The original pristine node baseline remains archived and unchanged.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.