Pull requests / #395
#395 Feature/nvidia p40
closed · @rafal-prasal · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
描述
# Fix `STRATA_EXPERIMENTAL_SM60`: the existing Pascal option cannot run on a Pascal card **Branch:** `NVIDIA TESLA P40` · **Base:** `30ec18e` (engine 0.1.30) · **Size:** +356 / −53 across 10 files --- ## The short version `CMakeLists.txt` already offers `-DSTRATA_EXPERIMENTAL_SM60=ON` and `dp4a.hpp` already ships both kernel-level fallbacks. The scaffolding is there and it is **dead code** — a build made with it compiles cleanly and then refuses to start on the only card it was written for. This PR finishes it, corrects it (`sm_60` → `sm_61`), and adds the one piece that was missing everywhere: a runtime path for the fact that cuBLAS has no bf16 GEMM below `sm_80`. Everything is opt-in. Default builds keep the `sm_75` floor and are byte-for-byte unaffected. ## Why the existing option does not work Four independent reasons, each verified against `30ec18e`: **1. The option is never propagated to a single source file.** ``` $ git grep -n "STRATA_EXPERIMENTAL_SM60" 30ec18e -- 30ec18e:CMakeLists.txt 30ec18e:include/strata/kernels/dp4a.hpp ← a comment ``` It is never `add_compile_definitions`'d, so it changes nothing for the code. **2. The runtime floor is unconditional, so the binary refuses its own target card.** `src/core/device.cu:138` reads `if (d.cc_major * 10 + d.cc_minor < 75)`. A Pascal build therefore compiles successfully and then throws at startup: ``` device Tesla P40 reports compute capability 6.1; Strata needs compute capability 7.5 or newer (RTX 20 / 30 / 40 / 50 series) ``` **3. There is no bf16 fallback, and one is required.** `src/prefill/gemm.cu:352` calls `cublasGemmEx` with bf16 inputs unconditionally. cuBLAS's bf16 GEMM is an `sm_80` feature and there is no flag that gives a pre-Ampere card one — it is the library, not the compiler. On Pascal it returns `NOT_SUPPORTED`. **4. The option names the wrong card.** `sm_60` is GP100, which predates `__dp4a` entirely (that arrived in 6.1). The viable datacenter card is `sm_61`. So today the flag's only observable effect is that CMake stops refusing `CMAKE_CUDA_ARCHITECTURES=61` — and then the engine refuses to run on it. ## What this PR changes | Area | Change | |---|---| | `CMakeLists.txt` | Rename the option to `STRATA_EXPERIMENTAL_PASCAL`; floor 61 (opt-in) vs 75; refuse a Pascal + `sm_120` mix (no toolkit compiles both); propagate `STRATA_EXPERIMENTAL_PASCAL=1` — but **only** when the arch list actually contains a pre-`sm_75` target, so a build that passes the option with no Pascal arch still refuses 6.1 rather than claiming a card it holds no cubin for. | | `src/prefill/gemm.cu` | **The substantive fix.** A `bf16_via_f32()` path: both operands widened by an exact shift, then a tiled `cublasSgemm`. Gated on the *runtime* CC, because a fat binary can carry several archs and a Pascal build is expected to run on newer cards too. | | `src/core/device.cu` | Runtime CC floor becomes 61 under the opt-in, 75 otherwise. Compile-time constant of the build, so a `sm_75` build still refuses a P40 — a Pascal binary cannot be handed to a card nobody measured it on. | | `setup.py` / `Dockerfile` | Detect 61/62, require CUDA 12.x (CUDA 13 removed offline compilation for Maxwell/Pascal/Volta), always build locally since there is no prebuilt engine, refuse the mixed `61`/`sm_120` build. | | `dp4a.hpp` | Comment corrected (it described `sm_60`/GP100). | **Why fp32 and not fp16.** `f32_from_bf16` is a shift, so the widening is *exact*. bf16 has 8 exponent bits and fp16 has 5, so an fp16 round trip would silently overflow a weight above 65504 and flush one below ~6e-5. The GEMM also keeps the same `CUBLAS_COMPUTE_32F` accumulator the Ampere path uses — this is not a lower-precision answer, it is the same answer computed more slowly through more memory. **Why tiling.** The scratch is sized in fp16 *elements*, so it holds half as many fp32 elements. X is `[T,K]`, W is `[N,K]`, and a prompt chunk times K easily exceeds the scratch, so both dimensions are tiled. Each tile is an independent GEMM over the full K, so a tile boundary changes nothing arithmetically. ## Validation **Numerically, on a real Tesla P40, against a double-precision CPU reference.** The real `Gemm::bf16()` from this branch was compiled `sm_61`, linked, and run on the card: 22 cases, worst relative error **2.66e-11**. Covers tiled and untiled shapes, `ldy > N` (row padding must not be touched), `beta` 0 and 1, and scratch-starved budgets down to `cap/K = 4`. The suite has teeth — reintroducing the `beta` bug this PR fixed makes it fail at 4.6e18 on the `beta=0` cases, because the caller's uninitialised `Y` is folded into rows nothing has written yet. **End-to-end, through the official Dockerfile path.** The image builds an `sm_61` engine from this branch (`BUILD.json`: `"archs": [61]`, `"source": "local"`) with the stock CUDA 12.9 toolkit, and runs healthy on the P40 — real 180B-class IQ3_S generation with speculative decoding, ~98% VRAM occupancy. **The gate, under real CMake.** The architecture-gate block is exercised verbatim: 21/21 cases, including that `STRATA_EXPERIMENTAL_PASCAL=1` reaches the compile flags for `61`/`62`/`70`/`74` but *not* for an `sm_120` build that merely passed the option. ## Measured, for calibration Tesla P40 (sm_61, 22.4 GB), IQ3_S, engine 0.1.30, one engine load. These are the engine's own llama.cpp-style timings, the same ones `docs/DETAILS.md` reports: | prompt tokens | prompt tok/s | decode tok/s | |---:|---:|---:| | 1,076 | 217.3 / 218.9 | 32.2 / 32.7 | | 4,148 | 314.8 / 314.0 | 31.5 / 31.7 | | 16,436 | 373.7 | 30.3 | | 32,820 | 357.0 | 30.1 | Against the documented RTX 5070 figures for the same model (prompt 427 / 913 / 1,624; decode 52.4 / 53.3 / 48.3) that is **0.51 / 0.34 / 0.22× on prompt and a steady ~0.60× on decode**. Pascal has no tensor cores at all, and fp16 buys only 1.26× over fp32 there because it is half-rate FMA rather than `mma`. Decode holds up better than the raw core ratio (~0.38×) because 22.4 GB of VRAM keeps 8,145 expert slots resident where a 12 GB card cannot. **This is a slow path and is documented as one.** It is not a recommendation; it is for people who already own the card. ## Two corrections found on the way **A card's generation is not a card.** The sm_61 band also contains a 2 GB GT 1030 — measured on this host, `nvidia-smi` reports compute capability 6.1 for it. It passes any architecture check and cannot hold any part of the model. `gpu_problem()` therefore gained a 6 GB floor, so it is refused on size with a sentence a user can act on rather than failing later inside the engine. Note this floor also catches small `sm_75+` cards, which was a pre-existing hole. **A comment in `qsa.cu` claimed hardware behaviour that does not hold.** It asserted that a Pascal card reports 96 KB of opt-in shared memory, accepts `cudaFuncSetAttribute` for it, then fails the launch. Measured on the P40: `MaxSharedMemoryPerBlockOptin` is 49152 — identical to the non-optin attribute — and a 64 KB request is *refused outright*. I had written a change on the strength of that comment; measuring it showed the change was a no-op, so **it is not in this PR** and the file is byte-identical to base. The pre-existing comment in `fused_gr.cu` makes a similar claim and is left alone rather than silently rewritten — flagging it as a follow-up. ## Safety - Default builds are unchanged: no Pascal arch → `_min_cc = 75`, no define emitted, no new branch taken. - The option now *fails to configure* on a Pascal + `sm_120` mix instead of emitting a binary that cannot exist. - `bf16_via_f32` is gated on runtime CC, so it is inert on `sm_80+` and on HIP (which is byte-identical — RDNA3/4 have bf16 in hardware). - Docs mark the path opt-in and not recommended. ## Suggested commit split Kept separate so the small fix is reviewable on its own: 1. `Propagate STRATA_EXPERIMENTAL_PASCAL and fix the 6.1 runtime floor` — the two-line bug fix for #1/#2 above. 2. `Add a pre-Ampere bf16 GEMM path (exact widen + tiled cublasSgemm)` — the substantive change, with its tests. 3. `setup.py / Dockerfile: detect Pascal, require CUDA 12.x, refuse mixed 61+120`. 4. `docs: P40 instructions, and a 6 GB VRAM floor` — including the `qsa.cu` note as a comment for a follow-up. ## Honest limitations - No full engine run was done on a *non-native* pack (Q2_0-style), because only an IQ3_S native pack was available. IQ packs require `--native SHARD1`, which bypasses `prefill.cpp` — so `bf16_via_f32` is **validated numerically but not exercised by the end-to-end benchmark above**. On a native pack what actually costs this card its speed is llama.cpp's MMQ on DP4A instead of tensor cores, plus the existing `__CUDA_ARCH__ < 800` fp32 fallbacks in `native_qsa_score` and `qsa_prompt_attn`. - There is no CI in this repo, so nothing will automatically catch a future kernel that breaks Pascal. That is the real ongoing cost of shipping this, and the most reasonable reason to decline it. - The measured rows are IQ3_S only — the heaviest quant — on Zen 3 without AVX-512, where the CPU expert pool is a visible share of each decode round. ## Reviewer's guide The one thing worth reading closely is `bf16_via_f32` in `src/prefill/gemm.cu`. It has three easy-to-get-wrong details, each of which the tests cover: the `cublasSgemm` is **column-major** (`m` = the N slice, `n` = the T slice, `C`'s leading dimension still `ldy`); the output offset is `Y + n0 + t0 * ldy` — an N slice moves a *row*, a T slice moves a *column*; and `beta` goes to **every** tile, because a T tile is a disjoint set of rows and an N tile a disjoint set of columns, so each element is written exactly once.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。