Pull requests / #330
#330 V100 (sm_70): the compute-capability floor drops from 7.5 to 7.0
closed · @taweili · 0 コメント · GitHub で見る
Setup & installAMD / HIPNVIDIA / CUDADocumentation
本文
The three places the 7.5 floor was written down all go to 7.0, and the messages now name the card the check begins with instead of the one it used to exclude: * CMakeLists.txt - configure time. `_base` is stripped of its -real/-virtual suffix first, so 70 and 70-virtual both pass; STRATA_EXPERIMENTAL_SM60 still admits a Pascal build below it. * src/core/device.cu - run time. This is the gate that matters even for a locally built binary: a binary carrying sm_70 code can be carried to a Pascal machine and would otherwise silently take whatever path the driver chose. * setup.py - gpu_problem(), plus the hint in check_gpus() and the "none of your GPUs" message. Without this the installer reports the V100 as unsupported and never reaches the compile step, so the other two gates would be invisible. A fourth place the floor is written down is engine/BUILD.json, which is not edited: build_engine() and the HIP builder both write it, and the prebuilt carries its own copy. Until the release tooling (tools/make_release.py, absent from this tree) emits an sm_70 asset, a V100 gets the compile-locally path - get_prebuilt() sees no code for arch 70 and falls through to build_engine(), which compiles for gpu["arch"]. The same shape as the unsupported-GPU model in docs/AMD_HIP.md. No kernel changes. The sm_80 guards were written for the experimental sm_75 build and key off cc_major < 8 or __CUDA_ARCH__ < 800, which sm_70 satisfies: qsa_block_scores_tc and qsa_prompt_attn_batch both refuse a V100 and prefill runs the old warp and decode kernels, and fused_gr sizes its chunk from cudaDevAttrMaxSharedMemoryPerBlockOptin, which a V100 answers with 96 KB - more than the 64 KB the Turing comment already anticipated. Only Volta's missing cp.async and TF32 mma cost anything, so the risk is performance rather than correctness. docs/NVIDIA_V100.md: the gate-by-gate walk, which translation units actually use sm_80 instructions (and that native_qsa_score has no caller outside its HIP test), the Volta/Turing/Ampere comparison, and the open points - the cc_major < 8 guards are coarse enough that an eventual Turing tensor-core path needs a finer split, and the V100's doubled resident warps per SM mean the fallback kernels' register and shared-memory pressure should be measured rather than assumed.
関連リンク
インストール・モデル・リリースへの站内リンク。