Pull requests / #88

#88 Pascal port: lower the CUDA floor to sm_61 (GTX 10 series)

closed · @hireymage · 0 commentaires · Sur GitHub

Setup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityWindowsLinux

Description

Follow-up to #87 (Turing port, sm_75) - this one is **stacked on that branch**, so the diff below shows both; once #87 merges, this PR's diff shrinks to just the Pascal delta. Files touched by *this* branch: CMakeLists.txt, setup.py, src/core/device.cu, src/prefill/gemm.{hpp,cu}, src/kernels/cuda/{elementwise,verify_kernels}.cu, include/strata/{artifact/gguf_reader,kernels/ngram}.hpp.

## What Pascal (sm_61) is missing, and what this adds

Developed and measured on a **GTX 1080 Ti (sm_61, CUDA 12.6, Linux)**, running the 125B Q2_0 pack end to end. All changes are runtime-guarded, so the sm_80+ path stays bit-identical.

1. **`Gemm::bf16` - cuBLAS has no BF16 GEMM below sm_80.** On cc < 70 both operands are converted to FP32 (a lossless 16-bit shift per element) and the same GEMM runs over `CUDA_R_32F`. Converted weights are cached per device pointer (the arena keeps weight pointers stable); activations are converted per call into a grown-on-demand buffer. Measured prefill slowdown ~2x on Pascal - expected, FP32 GEMM without tensor cores.
2. **`doorbell_wait_kernel` / `wait_flag_ge_kernel`** - `__nanosleep` landed with Volta (sm_70); pre-sm_70 builds spin with a bare `__threadfence()` instead (one thread, waiting on host-published memory).
3. **`GgufFile::open` (Linux)** - `madvise(MADV_RANDOM)` on the GGUF mapping: token gathers 16 rows scattered over the whole file, so read-ahead is pure waste. Guarded by `#ifdef MADV_RANDOM`. The `ngram.hpp` comment now documents the Linux/Windows split instead of claiming there is no Linux.
4. **Floor 75 -> 61** in CMake, the runtime device check, and the setup.py GPU check. Compiling for pre-sm_61 is still refused.

## Verification on the 1080 Ti

- Full engine build with `-DCMAKE_CUDA_ARCHITECTURES=61`: **181/181 targets, 0 errors**.
- `strata-device` on the card: `compute capability 6.1 (sm_61)`.
- **20/20 runnable parity selftests pass** on the real card (iq/native_expert/ple parity binaries are fixture-dependent and not shipped, so skipped).

Same note as #87: the sm_120 card I developed on is the author's documented target, so the Ampere+ path is unchanged by construction - every fallback sits behind a `__CUDA_ARCH__ >= 700` / `cc_major < 70` / `cc_major < 80` guard.

Sur le site

Liens install, modèles, releases.