Pull requests / #136

#136 Linux prebuilt from CI: CUDA 12.8, sm_80/86/89 + PTX, attached to each release

closed · @architectds · 0 Kommentare · Auf GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindowsLinux

Beschreibung

The first piece from #103, as you suggested: a Linux prebuilt, built by CI.

**The workflow** (`.github/workflows/linux-prebuilt.yml`) compiles `strata-linux-x64.zip` on Ubuntu 22.04. It needs no GPU.
- It builds the engine and the image encoder (`strata-vision`) with CUDA 12.8.
- The architectures are `80-real;86-real;89`: sm_80, sm_86 and sm_89 as machine code, plus compute_89 PTX for newer cards.
- The CPU code is at the AVX2 baseline (`STRATA_PORTABLE=ON`). The AVX-512 kernels are still picked at run time where the CPU has them.
- It writes the `BUILD.json` that `get_prebuilt` reads.

When you publish a release, a second job attaches the zip and its SHA-256 to it. `releases/latest/download/strata-linux-x64.zip` then sits next to the Windows zip, about 40 minutes after publishing. Until then, a Linux setup compiles, as it does today. Running the workflow by hand, from Actions > linux-prebuilt with other architectures if you want, builds an artifact only. So does a PR that changes the workflow.

**setup.py uses CUDA 12 on Linux** to match that build:
- CUDA 12.8's pip packages, pinned like the CUDA 13 ones. CUDA 12's wheels keep cuBLAS and the runtime in separate folders, which `cuda_lib_dirs` now finds.
- Driver 525 as the minimum.
- The 12.8 toolkit when a compile has to install one.

Windows is unchanged (CUDA 13). `STRATA_CUDA=12` or `13` picks the other one. A prebuilt built with another CUDA major is not used, because it would look for libraries setup did not install; setup compiles instead.

**Tested:**
- The workflow on my fork: [run 36565974211](https://github.com/architectds/Strata/actions/runs/36565974211), on upstream `main` (0.1.23), 37 minutes. The engine for three architectures took 5 min, the image encoder 29 min. The zip is 99.5 MB:
  - `BUILD.json` is `{"version": "0.1.23", "archs": [80, 86, 89], "ptx": true, "cuda": "12.8", "vision": "gpu", ...}`;
  - both binaries link `libcudart.so.12`, `libcublas.so.12` and `libcublasLt.so.12`, which are the pinned pip wheels.
- The same build setup (sm_80 only, engine 0.1.20) ran Qwen3.8-Flash-Next IQ3_S on a Colab A100-40G through `setup.py --prebuilt` all day: 262K context, images, and decode at 61 tok/s.

The RAM-based context cap is the other piece: #134.

Mehr auf der Site

Links zu Install, Modellen, Releases.