Pull requests / #81
#81 Accept sm_80+ GPUs at runtime, matching the build guard and the README
closed · merged 2026-09-29 · @mikicvi · 0 评论 · 在 GitHub 查看
NVIDIA / CUDAModels & quantsDocumentation
描述
The CMake arch guard already lets RTX 30 (sm_86) / 40 (sm_89) build — its own comment says "the kernels need sm_80 or newer (tf32 mma in the attention scorer, bf16 math), so RTX 30 (86) and RTX 40 (89) build too" — and the README lists RTX 30/40/50 as supported. But `device_info()` still refused anything with `cc_major != 12`, so a working Ampere/Ada build died at startup with "Strata targets sm_120 (RTX 5000 series / Blackwell) only". This relaxes the runtime check to the kernels' real floor (compute capability 8.0), aligns the stale "120-only" header comment in CMakeLists.txt with the guard below it, and fixes the "no CUDA device" message that also claimed sm_120-only. No behavior change on Blackwell. Verified on Ampere, CUDA 12.0 / gcc-13, engine 0.1.20, GSQ-RCO IQ2_XS and IQ3_XXS — now serving daily on this machine: - RTX A2000 12GB + RTX 3080 10GB (both sm_86), single-GPU on either card - Decode 21-24 t/s flat from 1k to 128k context (13-layer KV + streaming), 29.6 t/s on code generation at 81% draft acceptance - Prefill 375 t/s `@8k`, 445 t/s `@32k` with `--prefill auto` - Quality: 8/8 multi-needle retrieval at 113k context on both quants (real shuffled source code as filler) - This exact branch compiled from a clean checkout: CUDA 12.0, `CMAKE_CUDA_ARCHITECTURES=86` sm_89 is covered by the same reasoning but not measured here; no Blackwell-specific intrinsics are used on these paths, which is what the existing build guard already asserts.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。