Pull requests / #395

#395 Feature/nvidia p40

closed · @rafal-prasal · 0 Kommentare · Auf GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Beschreibung

# Fix `STRATA_EXPERIMENTAL_SM60`: the existing Pascal option cannot run on a Pascal card

**Branch:** `NVIDIA TESLA P40` · **Base:** `30ec18e` (engine 0.1.30) · **Size:** +356 / −53 across 10 files

---

## The short version

`CMakeLists.txt` already offers `-DSTRATA_EXPERIMENTAL_SM60=ON` and `dp4a.hpp` already ships both
kernel-level fallbacks. The scaffolding is there and it is **dead code** — a build made with it compiles
cleanly and then refuses to start on the only card it was written for. This PR finishes it, corrects it
(`sm_60` → `sm_61`), and adds the one piece that was missing everywhere: a runtime path for the fact that
cuBLAS has no bf16 GEMM below `sm_80`.

Everything is opt-in. Default builds keep the `sm_75` floor and are byte-for-byte unaffected.

## Why the existing option does not work

Four independent reasons, each verified against `30ec18e`:

**1. The option is never propagated to a single source file.**

```
$ git grep -n "STRATA_EXPERIMENTAL_SM60" 30ec18e --
30ec18e:CMakeLists.txt
30ec18e:include/strata/kernels/dp4a.hpp     ← a comment
```

It is never `add_compile_definitions`'d, so it changes nothing for the code.

**2. The runtime floor is unconditional, so the binary refuses its own target card.**
`src/core/device.cu:138` reads `if (d.cc_major * 10 + d.cc_minor < 75)`. A Pascal build therefore
compiles successfully and then throws at startup:

```
device Tesla P40 reports compute capability 6.1; Strata needs compute
capability 7.5 or newer (RTX 20 / 30 / 40 / 50 series)
```

**3. There is no bf16 fallback, and one is required.** `src/prefill/gemm.cu:352` calls `cublasGemmEx` with
bf16 inputs unconditionally. cuBLAS's bf16 GEMM is an `sm_80` feature and there is no flag that gives a
pre-Ampere card one — it is the library, not the compiler. On Pascal it returns `NOT_SUPPORTED`.

**4. The option names the wrong card.** `sm_60` is GP100, which predates `__dp4a` entirely (that arrived in
6.1). The viable datacenter card is `sm_61`.

So today the flag's only observable effect is that CMake stops refusing `CMAKE_CUDA_ARCHITECTURES=61` — and
then the engine refuses to run on it.

## What this PR changes

| Area | Change |
|---|---|
| `CMakeLists.txt` | Rename the option to `STRATA_EXPERIMENTAL_PASCAL`; floor 61 (opt-in) vs 75; refuse a Pascal + `sm_120` mix (no toolkit compiles both); propagate `STRATA_EXPERIMENTAL_PASCAL=1` — but **only** when the arch list actually contains a pre-`sm_75` target, so a build that passes the option with no Pascal arch still refuses 6.1 rather than claiming a card it holds no cubin for. |
| `src/prefill/gemm.cu` | **The substantive fix.** A `bf16_via_f32()` path: both operands widened by an exact shift, then a tiled `cublasSgemm`. Gated on the *runtime* CC, because a fat binary can carry several archs and a Pascal build is expected to run on newer cards too. |
| `src/core/device.cu` | Runtime CC floor becomes 61 under the opt-in, 75 otherwise. Compile-time constant of the build, so a `sm_75` build still refuses a P40 — a Pascal binary cannot be handed to a card nobody measured it on. |
| `setup.py` / `Dockerfile` | Detect 61/62, require CUDA 12.x (CUDA 13 removed offline compilation for Maxwell/Pascal/Volta), always build locally since there is no prebuilt engine, refuse the mixed `61`/`sm_120` build. |
| `dp4a.hpp` | Comment corrected (it described `sm_60`/GP100). |

**Why fp32 and not fp16.** `f32_from_bf16` is a shift, so the widening is *exact*. bf16 has 8 exponent bits
and fp16 has 5, so an fp16 round trip would silently overflow a weight above 65504 and flush one below ~6e-5.
The GEMM also keeps the same `CUBLAS_COMPUTE_32F` accumulator the Ampere path uses — this is not a
lower-precision answer, it is the same answer computed more slowly through more memory.

**Why tiling.** The scratch is sized in fp16 *elements*, so it holds half as many fp32 elements. X is
`[T,K]`, W is `[N,K]`, and a prompt chunk times K easily exceeds the scratch, so both dimensions are tiled.
Each tile is an independent GEMM over the full K, so a tile boundary changes nothing arithmetically.

## Validation

**Numerically, on a real Tesla P40, against a double-precision CPU reference.** The real `Gemm::bf16()` from
this branch was compiled `sm_61`, linked, and run on the card: 22 cases, worst relative error **2.66e-11**.
Covers tiled and untiled shapes, `ldy > N` (row padding must not be touched), `beta` 0 and 1, and
scratch-starved budgets down to `cap/K = 4`.

The suite has teeth — reintroducing the `beta` bug this PR fixed makes it fail at 4.6e18 on the `beta=0`
cases, because the caller's uninitialised `Y` is folded into rows nothing has written yet.

**End-to-end, through the official Dockerfile path.** The image builds an `sm_61` engine from this branch
(`BUILD.json`: `"archs": [61]`, `"source": "local"`) with the stock CUDA 12.9 toolkit, and runs healthy on
the P40 — real 180B-class IQ3_S generation with speculative decoding, ~98% VRAM occupancy.

**The gate, under real CMake.** The architecture-gate block is exercised verbatim: 21/21 cases, including
that `STRATA_EXPERIMENTAL_PASCAL=1` reaches the compile flags for `61`/`62`/`70`/`74` but *not* for an
`sm_120` build that merely passed the option.

## Measured, for calibration

Tesla P40 (sm_61, 22.4 GB), IQ3_S, engine 0.1.30, one engine load. These are the engine's own
llama.cpp-style timings, the same ones `docs/DETAILS.md` reports:

| prompt tokens | prompt tok/s | decode tok/s |
|---:|---:|---:|
| 1,076 | 217.3 / 218.9 | 32.2 / 32.7 |
| 4,148 | 314.8 / 314.0 | 31.5 / 31.7 |
| 16,436 | 373.7 | 30.3 |
| 32,820 | 357.0 | 30.1 |

Against the documented RTX 5070 figures for the same model (prompt 427 / 913 / 1,624; decode 52.4 / 53.3 /
48.3) that is **0.51 / 0.34 / 0.22× on prompt and a steady ~0.60× on decode**. Pascal has no tensor cores
at all, and fp16 buys only 1.26× over fp32 there because it is half-rate FMA rather than `mma`. Decode holds
up better than the raw core ratio (~0.38×) because 22.4 GB of VRAM keeps 8,145 expert slots resident where a
12 GB card cannot.

**This is a slow path and is documented as one.** It is not a recommendation; it is for people who already
own the card.

## Two corrections found on the way

**A card's generation is not a card.** The sm_61 band also contains a 2 GB GT 1030 — measured on this host,
`nvidia-smi` reports compute capability 6.1 for it. It passes any architecture check and cannot hold any part
of the model. `gpu_problem()` therefore gained a 6 GB floor, so it is refused on size with a sentence a user
can act on rather than failing later inside the engine. Note this floor also catches small `sm_75+` cards,
which was a pre-existing hole.

**A comment in `qsa.cu` claimed hardware behaviour that does not hold.** It asserted that a Pascal card
reports 96 KB of opt-in shared memory, accepts `cudaFuncSetAttribute` for it, then fails the launch.
Measured on the P40: `MaxSharedMemoryPerBlockOptin` is 49152 — identical to the non-optin attribute — and a
64 KB request is *refused outright*. I had written a change on the strength of that comment; measuring it
showed the change was a no-op, so **it is not in this PR** and the file is byte-identical to base. The
pre-existing comment in `fused_gr.cu` makes a similar claim and is left alone rather than silently rewritten
— flagging it as a follow-up.

## Safety

- Default builds are unchanged: no Pascal arch → `_min_cc = 75`, no define emitted, no new branch taken.
- The option now *fails to configure* on a Pascal + `sm_120` mix instead of emitting a binary that cannot
  exist.
- `bf16_via_f32` is gated on runtime CC, so it is inert on `sm_80+` and on HIP (which is byte-identical —
  RDNA3/4 have bf16 in hardware).
- Docs mark the path opt-in and not recommended.

## Suggested commit split

Kept separate so the small fix is reviewable on its own:

1. `Propagate STRATA_EXPERIMENTAL_PASCAL and fix the 6.1 runtime floor` — the two-line bug fix for #1/#2 above.
2. `Add a pre-Ampere bf16 GEMM path (exact widen + tiled cublasSgemm)` — the substantive change, with its tests.
3. `setup.py / Dockerfile: detect Pascal, require CUDA 12.x, refuse mixed 61+120`.
4. `docs: P40 instructions, and a 6 GB VRAM floor` — including the `qsa.cu` note as a comment for a follow-up.

## Honest limitations

- No full engine run was done on a *non-native* pack (Q2_0-style), because only an IQ3_S native pack was
  available. IQ packs require `--native SHARD1`, which bypasses `prefill.cpp` — so `bf16_via_f32` is
  **validated numerically but not exercised by the end-to-end benchmark above**. On a native pack what
  actually costs this card its speed is llama.cpp's MMQ on DP4A instead of tensor cores, plus the existing
  `__CUDA_ARCH__ < 800` fp32 fallbacks in `native_qsa_score` and `qsa_prompt_attn`.
- There is no CI in this repo, so nothing will automatically catch a future kernel that breaks Pascal. That
  is the real ongoing cost of shipping this, and the most reasonable reason to decline it.
- The measured rows are IQ3_S only — the heaviest quant — on Zen 3 without AVX-512, where the CPU expert
  pool is a visible share of each decode round.

## Reviewer's guide

The one thing worth reading closely is `bf16_via_f32` in `src/prefill/gemm.cu`. It has three easy-to-get-wrong
details, each of which the tests cover: the `cublasSgemm` is **column-major** (`m` = the N slice, `n` = the T
slice, `C`'s leading dimension still `ldy`); the output offset is `Y + n0 + t0 * ldy` — an N slice moves a
*row*, a T slice moves a *column*; and `beta` goes to **every** tile, because a T tile is a disjoint set of
rows and an N tile a disjoint set of columns, so each element is written exactly once.

Mehr auf der Site

Links zu Install, Modellen, Releases.