Issues / #585

#585 Windows: source build for sm_70 (Volta) fails to link - LNK1169 (cudart_static.lib vs cudart.lib)

closed · @noahark · 1 commentaires · Sur GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Description

# Windows: source build for sm_70 (Volta) fails to link - LNK1169 (cudart_static.lib x cudart.lib)

Thanks for the great project - a 125B model running on a 2017-era workstation feels like magic. I hit a build bug on the **Windows + sm_70 experimental path** (PR #295), with a one-line fix that I verified locally.

## Environment

- Strata v0.1.38 (source, git clone)
- Windows 10 Enterprise LTSC 2021 (19044)
- Tesla V100-PCIE-32GB (sm_70, TCC), driver 581.80
- MSVC 14.44 (VS Build Tools), CUDA Toolkit 12.4 (system-wide, pre-existing)
- `STRATA_EXPERIMENTAL_SM60=1` set, `SETUP.bat --setup --yes --build ...`

## What happens

All 129 targets compile; the final link of `strata.exe` fails:

```
cudart_static.lib(cudart_generated_cuda_runtime_api.obj) : error LNK2005:
  cudaDeviceCanAccessPeer already defined in cudart.lib(cudart64_12.dll)
  [~40 more symbols: cudaDeviceSynchronize, cudaEventCreate, ...]
strata.exe : fatal error LNK1169: one or more multiply defined symbols found
```

## Root cause

CMake's CUDA language links the runtime per `CMAKE_CUDA_RUNTIME_LIBRARY`, whose **default is `Static`** (`cudart_static.lib`), while `CMakeLists.txt` additionally links the imported `CUDA::cudart` target (the dynamic import lib) - e.g. `_strata_gpu_runtime_target` and several `target_link_libraries(... CUDA::cudart)`. Both end up on the link line:

```
link.exe ... cudart.lib ... cudadevrt.lib cudart_static.lib ...
```

MSVC rejects the duplicated definitions outright. GNU ld on Linux tolerates them, which is presumably why the Pascal/Volta path (tested in #295 on Linux) never hit this on Windows.

## Fix (verified)

Configure with a shared CUDA runtime:

```
cmake -S . -B build -DCMAKE_CUDA_RUNTIME_LIBRARY=Shared
```

then re-run setup (or the build). Upstream one-liner if you prefer it in-tree:

```cmake
set(CMAKE_CUDA_RUNTIME_LIBRARY Shared)   # before enable_language(CUDA)
```

## After the fix: V100-PCIE-32GB benchmarks

Worth adding to `docs/COMMUNITY_BENCHMARKS.md` - I couldn't find any V100 numbers there, and #236 (same hardware class) didn't post any.

**Setup:** V100-PCIE-32GB (solo, PCIe Gen3) - dual Xeon E5-2696 v3 - 128 GB RAM - IQ3_XXS - 131,072 context (KV streamed to RAM) - MTP draft (q2_0) - 15,553/24,576 experts (63%) cached in VRAM (25.2 GiB).

| Scenario | Output | Speed | Expert cache hit |
|---|---|---|---|
| First request ever (JIT warm-up) | 175 tok | 7.5 tok/s | 81.0% |
| Short answer (repetitive list) | 39 tok | 10.1 tok/s | 94.8% |
| Medium answer | 289 tok | 13.7 tok/s | 95.5% |
| ~400-word essay (thinking + prose) | 322 tok | 14.2 tok/s | 95.6% |
| Long-form writing | 1,492 tok | 15.8 tok/s | 96.6% |
| Long multi-method math reasoning | 1,761 tok | **16.8 tok/s** | 97.3% |

- Prompt processing: 61.7 tok/s (warm)
- MTP acceptance: 97.8% repetitive / 64.9% free prose
- Decode is memory-bound on this card: GPU util ~19%, 42-48 W of 250 W, 49 deg C; the CPU computes the RAM-resident experts (~55% load). Speed also ramps within a request as the expert cache warms (13.6 to 15.8 tok/s observed).

## Minor note: setup in restricted-network regions

`get_llama_cpp()` retries the GitHub zip download forever (`WinError 10060`) in regions where github.com is unreachable. Pre-populating `third_party/llama.cpp` at the pinned commit (e.g. via a proxy) makes setup skip it - maybe worth a line in INSTALL.md.

Happy to turn the fix into a PR if it helps.

Sur le site

Liens install, modèles, releases.