Issues / #1485
#1485 [Bug]: update.sh cannot rebuild a locally compiled engine on a PC with a Volta card beside an RTX 40 (nvcc: Unsupported gpu architecture 'compute_70')
open · @bh611 · 0 评论 · 在 GitHub 查看
Setup & installNVIDIA / CUDAModels & quants
描述
## What happened
`./update.sh` on a checkout whose engine is compiled locally (`engine/BUILD.json` → `"source": "local"`) cannot rebuild the engine. The rebuild picks the wrong architecture and at the end `--update` reports success while keeping the old engine:
```
Getting the newest Strata (git pull) ...
已经是最新的。
[ok] build tools present (CUDA 13.0)
Compiling the engine for sm_70, sm_89 (a card it had no code for; 10-20 minutes, once) ...
> .../.venv/bin/cmake -G Ninja ... -S .../Strata -B .../Strata/build -DCMAKE_BUILD_TYPE=Release \
-DSTRATA_ENABLE_CUDA=ON -DSTRATA_BUILD_TESTS=OFF -DCMAKE_CUDA_ARCHITECTURES=70;89 \
-DCMAKE_CUDA_COMPILER=/usr/local/cuda-13.0/bin/nvcc -DSTRATA_EXPERIMENTAL_SM60=ON
FAILED: [code=1] CMakeFiles/strata_core.dir/src/core/device.cu.o
/usr/local/cuda-13.0/bin/nvcc -forward-unknown-to-host-compiler -DSTRATA_EXPERIMENTAL_SM60=1 \
-DSTRATA_VERSION=\"0.1.40.3\" ... "--generate-code=arch=compute_70,code=[compute_70,sm_70]" \
"--generate-code=arch=compute_89,code=[compute_89,sm_89]" ... -c src/core/device.cu
nvcc fatal : Unsupported gpu architecture 'compute_70'
... (15 more .cu objects, same error) ...
ninja: build stopped: subcommand failed.
(the build stopped - trying it once more)
... same 15 failures ...
[X] command failed (exit 1): cmake
Setup stopped. Fix the item above and run it again - everything already done is kept and skipped.
[!] could not compile the updated engine: starting the installed one
[ok] qwen3.8-flash-next-iq3_s: up to date
[ok] Strata is updated (engine 0.1.38). Start the model with ./setup.sh when you want it.
```
**Expected:** the engine for this model's actual cards (GPU 0 + GPU 2, both RTX 4090, sm_89) is rebuilt with CUDA 13 and the update finishes with engine 0.1.40.3.
### Why it happens
The PC has three NVIDIA cards: GPU 0 = RTX 4090 (24 GB, sm_89), **GPU 1 = Tesla V100-PCIE-32GB (32 GB, sm_70, never opted in)**, GPU 2 = RTX 4090 (24 GB). The model is installed on GPU 0 + GPU 2 (`strata-iq3_s.json` → `"gpu": [0, 2]`, `"layer_split": "auto"`).
1. `./update.sh` → `setup.py --update` → `update_install()` → `update_installed_engine(url_base, 13)` (`setup.py:2853`) — for the main engine the toolkit is hardcoded to 13.
2. That function takes the card from **`gpu_info()` with no pick** (`setup.py:2909`). `gpu_info()` returns the card with the **most VRAM** (`setup.py:1195`: `g = max(found, key=lambda x: (round(x["vram_gb"]), -x["index"]))`) — here that is the **V100** (32 GB > 24 GB), although it is not one of the model's GPUs and was not admitted as an experimental card (`OLD_GPUS` is only set later, `setup.py:4717`).
3. So `archs` becomes `{70} ∪ BUILD.json's {89}` (`setup.py:2914`), and because the toolkit is 13, `install_build_tools()` computes `old = False` (`setup.py:2984`). The CUDA 12 guard at `setup.py:2991` ("...is compiled here with the NVIDIA CUDA Toolkit 12.x... CUDA 13 cannot compile for these cards") is bypassed, and the build runs with CUDA 13 for `sm_70` — which CUDA 13 cannot compile.
4. `--gpu` / `--gpus` cannot work around it: the `a.update` branch returns (`setup.py:4644`) before `GPU_PICK = a.gpu` is set (`setup.py:4800`). `CUDA_VISIBLE_DEVICES` does not help either, since `gpus()` reads `nvidia-smi`.
### Suggested fix
- In `update_installed_engine()` use the model's own GPU set (the config's `gpu` / `gpus`), or filter `gpu_info()` through `gpu_problem()` / `old_gpus_opt_in()`, so a card that is not part of this model's engine cannot enter the arch list.
- Use `cuda_choice(archs, cuda)` (`setup.py:843`) instead of the hardcoded `13`, or at least let the union of architectures reach `install_build_tools()`'s existing CUDA 12 decision.
- Possibly related: #128 (engine not rebuilt when adding a GPU with a different architecture), #295 (the experimental sm_60/70 build), #236 (community Volta build).
### Workaround
Rebuild through the project's own path with the 4090 only, which writes a correct `engine/BUILD.json` for 0.1.40.3:
```python
import setup as S
g = S.gpu_info(0) # 0 = the RTX 4090
S.build_engine({**g, "archs": [89]}, "gpu", False, S.get_llama_cpp(), toolkit=13)
```
`./setup.sh` then sees the engine as up to date and starts normally.
---
## Strata version, GPU, OS
```
0.1.40.3 (git checkout at d5ea713; engine 0.1.38 stayed because the rebuild failed)
GPU 0: RTX 4090 24 GB (sm_89, used) · GPU 1: Tesla V100-PCIE-32GB 32 GB (sm_70, unused) · GPU 2: RTX 4090 24 GB (sm_89, used)
model IQ3_S on GPU 0,2 with an auto layer split, 262144 context, images on, engine compiled locally
Ubuntu 24.04.4 · driver 580.178.04 · CUDA 13.0.103 (nvcc V13.0.88) · gcc 13.3.0
```
## Engine log (optional)
This bug is in the build step, not in the engine, so the useful log is the compile output above (the full terminal output, or `./update.sh 2>&1 | tee update-fail.log` on the next attempt). The engine log of the previous successful start is attached as `strata-iq3_s.log`.站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。