Issues / #1071

#1071 vmm.cpp: the CUDA 12 engine does not build with a toolkit older than 12.5 (missing CUDART_VERSION guard)

closed · @zhouxihong1 · 1 comments · View on GitHub

Setup & installMulti-GPUAMD / HIPNVIDIA / CUDADocumentationLinux

Description

## Problem

Building the experimental CUDA 12 engine (`--cuda 12`) with a toolkit older than **CUDA 12.5** fails on the new `src/core/vmm.cpp`:

```
[51/136] Building CXX object CMakeFiles/strata_engine.dir/src/core/vmm.cpp.o
FAILED: CMakeFiles/strata_engine.dir/src/core/vmm.cpp.o
/usr/bin/g++-10 ... -isystem /usr/local/cuda-12.2/targets/x86_64-linux/include ... -c src/core/vmm.cpp
src/core/vmm.cpp: In instantiation of 'bool strata::core::{anonymous}::resolve(const char*, F&)':
src/core/vmm.cpp:43:50:   required from here
src/core/vmm.cpp:28:41: error: 'cudaGetDriverEntryPointByVersion' was not declared in this scope;
                                     did you mean 'cudaGetDriverEntryPointFlags'?
```

`cudaGetDriverEntryPointByVersion` arrived in CUDA 12.5. The same call is already version-gated one file over:

```cpp
// src/core/expert_cache.cpp:225-231
#if CUDART_VERSION >= 12050
    const cudaError_t e = cudaGetDriverEntryPointByVersion(name, &p, 12000, cudaEnableDefault, &q);
#else
    const cudaError_t e = cudaGetDriverEntryPoint(name, &p, cudaEnableDefault, &q);
#endif
```

`src/core/vmm.cpp:28` misses that guard. Both files resolve the **same** driver entry points (`cuDeviceGetAttribute`, `cuMemGetAllocationGranularity`, `cuMemAddressReserve`, `cuMemAddressFree`, `cuMemCreate`, `cuMemRelease`, `cuMemMap`, `cuMemUnmap`, `cuMemSetAccess`), so the `#else` pattern is already the one that ships next door.

**Why it is a regression:** `src/core/vmm.cpp` is new in 0.1.40 (added in `1735d64`, the release tip). 0.1.39 built with CUDA 12.2; 0.1.40 does not. (This also matches #1066, which hits the same new file from the gfx906 side.)

**Also inconsistent with setup.py:** `setup.py:2505` accepts CUDA 12.0 for arch < 120:

```python
need12 = (12, 8) if max(archs) >= 120 else (12, 0)     # sm_120 needs CUDA 12.8 or newer
```

while the code needs 12.5. `docs/OLDER_GPUS.md:34` and the `--build` failure message (`setup.py:2514`, "install CUDA 12.9") both say the CUDA 12 engine is built with CUDA 12.9 — the check is simply too permissive.

## Fix

The same guard as `expert_cache.cpp`:

```diff
 template <class F> bool resolve(const char* name, F& f) {
     cudaDriverEntryPointQueryResult q{};
     void* p = nullptr;
+#if CUDART_VERSION >= 12050   // cudaGetDriverEntryPointByVersion arrived in CUDA 12.5
+    const cudaError_t e = cudaGetDriverEntryPointByVersion(name, &p, 12000, cudaEnableDefault, &q);
+#else
+    const cudaError_t e = cudaGetDriverEntryPoint(name, &p, cudaEnableDefault, &q);
+#endif
-    if (cudaGetDriverEntryPointByVersion(name, &p, 12000, cudaEnableDefault, &q) != cudaSuccess ||
-        q != cudaDriverEntryPointSuccess || p == nullptr)
+    if (e != cudaSuccess || q != cudaDriverEntryPointSuccess || p == nullptr)
         return false;
     f = (F) p;
     return true;
 }
```

`cudaGetDriverEntryPoint(name, &p, cudaEnableDefault, &q)` returns the entry point for the runtime's own version; the pinned `12000` asks for the CUDA 12.0 ABI of the same symbol, and these `cu*` VMM entry points have stable signatures across 12.x. On CUDA >= 12.5 the guard takes exactly today's call, so it only affects older toolkits.

## Scope / safety

- Only the CUDA branch of `vmm.cpp` changes. The HIP path is inside `#if !defined(STRATA_USE_HIP)` at the top of the file and is untouched.
- On CUDA >= 12.5 (`CUDART_VERSION >= 12050`) the compiled code is identical to today's, so CUDA 13 and the ready-made engines are unaffected.
- #1066 is the HIP/gfx906 counterpart on the same new file (a missing `STRATA_HIP_GFX906` gate rather than a `CUDART_VERSION` one).

## Verified

- 4x RTX 4090 (sm_89), Ubuntu, GCC 10 (`-std=c++2a`), CUDA Toolkit 12.2, `--cuda 12`, `--layer-split 12,24,36`, engine v0.1.40 (`1735d647`).
- With the guard above, `src/core/vmm.cpp` compiles and the build proceeds past `51/136`.
- The call site is reached in every build, not only in an unused template: `vmm.cpp:43` instantiates `resolve("cuDeviceGetAttribute", attr)` from `api()`.

Workaround for anyone on CUDA 12.0-12.4 today: build with CUDA 12.9 as `docs/OLDER_GPUS.md` intends, or carry the guard above locally.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.