Pull requests / #87

#87 Turing port: run on RTX 20 (sm_75) and newer

closed · @hireymage · 0 comments · View on GitHub

Setup & installServer & APINVIDIA / CUDADocumentationWindowsLinux

Description

The rest of the Turing port, rebased onto main (0.1.22).

What main already covers after the rebase (dropped from the PR):
- the portable pre-sm_80 QSA scorer (main has its own fp32-FMA path) and the sm_80 guard on the tensor-core prompt attention (a751715),
- a per-device fused_gr shared-memory opt-in capped at the card's limit.

What the PR adds on top (+43/−20, single commit):
- **the runtime floor**: `device_info` refuses anything below compute capability 7 — main's runtime check still demands sm_120, so a binary built for sm_75 would never start,
- **fused_gr chunking** on top of main's per-device opt-in: the down kernel carries tokens in slices that fit the card's opt-in (Turing, 64 KB → 6-token slices; 4 tokens when no opt-in is reported), so full 8-token windows no longer fail with "invalid argument" on Turing,
- **the setup/CMake floor is 7.5** (RTX 20 / 30 / 40 / 50) without an experimental flag: CMake's arch refusals, setup.py's `gpu_problem` and the GPU hint/error texts, and docs/MULTI_GPU.md updated.

Compiled clean for `CMAKE_CUDA_ARCHITECTURES=75` (121/121, CUDA 12.6, Linux). Runtime verified end-to-end on an RTX 2070 (sm_75) at the 0.1.20 base: model loads, OpenAI/Anthropic serving works, greedy output identical to the sm_120 build.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.