Issues / #1470
#1470 Pascal (sm_61): decode halves since 0.1.40.2. 2e4ddf6 drops `__restrict__` and peels the MMVQ loop on every CUDA arch
closed · @lineape · 1 Kommentare · Auf GitHub
Setup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityWindowsLinux
Beschreibung
## Summary
On a GTX 1080 + GTX 1070 layer split, decode goes from about 33 t/s on 0.1.40.1 to about 18 t/s on 0.1.40.2 and
0.1.40.3. Prefill drops about 3%. Expert-cache slot counts, draft acceptance and hit rates are the same in both arms.
The cause is the compile-time part of 2e4ddf6 ("verify window: PDL on sm_90+, …", #904). In
`include/strata/kernels/pdl.hpp`, `STRATA_PDL_RESTRICT` became empty and `kPdlPrefetch` became `true` on **all** CUDA
architectures, not just sm_90+. On sm_61 the multi-column MMVQ kernel loses its read-only (`LDG.nc`) loads. Restoring
both (a 2-line change, below) brings 0.1.40.3 back to 0.1.40.1's speed on Pascal. On an RTX 3080 it makes no
measurable difference.
## Hardware and config
**Box 1 (regressed):**
- GPUs: GTX 1080 8 GB (CUDA0, layers 0–25) + GTX 1070 8 GB (CUDA1, layers 26–47 + head), both GP104, sm_61, PCIe 3.0 x16.
- Host: Xeon Gold 6230. The engine log says "no AVX-512: the expert kernels run on AVX-2". Linux (NixOS), driver 580.178.04.
- Build: CUDA 12.8, gcc 14.4, `STRATA_EXPERIMENTAL_SM60=1`, archs `[61]`, `setup.sh --build`. The CMakeCache of
the two versions is identical apart from the source.
- Model: Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS, native pack.
- Engine flags: `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <rt> --max-context 524288 --kv int8 --kv-resident 32768 --conversation-cache-mib 16384 --conversation-cache-slots 8 --layer-split 26 --trim-stage-weights --pipeline-windows 2 --pcie-frac 0.10 --rope-scaling yarn --rope-scale 2`.
- Env: `STRATA_IQ256_GATHER=1 STRATA_RING_BYTES=0`.
- Expert cache: PROFILE policy, pre-filled from `data/expert-profile.bin`, no eviction.
**Box 2 (control):**
- GPU: one RTX 3080 (sm_86). Host: Ryzen 9 7945HX. Windows.
- Build: local, CUDA 13.0 + VS 2022 Build Tools, archs `[86]`.
- Model: Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS.
- Engine flags: `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.70 --mtp <rt> --max-context 524288 --kv int8 --kv-resident 32768 --rope-scaling yarn --rope-scale 2 --conversation-cache-mib 12288 --conversation-cache-slots 6 --vision --pcie-frac 0.35 --prefix-cache-dir <one dir per arm>`.
- Our engine there also carries #1090 (prefix cache) in all three arms.
## Method
This follows the paired method from #1373: arms alternate in one session, nothing else changes between them, and
each arm's `expert cache N slots` line is listed.
- **Order:** A, B, A, B (or A, B, C, A, C, B). The same serve layer and the same config throughout; only the engine
binary changes, or one env var.
- **Load:** each round loads the engine fresh. With the PROFILE policy there is no adaptive warm-up of the cache,
but the first decode after a load is still 10–20% slower. So each round reports the median of three decodes, and
the raw runs are listed too.
- **Decode:** 400 tokens, greedy (`temperature: 0`), from a ~115-token prompt (a fresh UUID prefix each request, so
no cache reuse), three requests per round. The figure is Strata's own `timings.predicted_per_second`, median of three.
- **Prefill:** a ~15.6k-token prompt, fresh UUID prefix, three requests, median `prompt_per_second`.
- **Switch-back:** two ~11k-token conversations in the order A, B, A, B. The figure is the wall time of the last two
requests. Each restores a parked conversation (`cache_n` ≈ 10,965), so it times the snapshot restore plus a
short verify.
## Results: box 1 (GTX 1080 + GTX 1070, IQ2_XS)
**Paired run 1: 0.1.40.1 vs stock 0.1.40.3**
| round | engine | expert cache slots (CUDA0 + CUDA1) | prefill 16k (t/s) | decode median (runs) | switch-back (s) |
|---|---|---|---:|---|---|
| 1 | 0.1.40.1 | 3275 + 2574 | 389.9 | **34.3** (27.6 34.6 34.3) | 1.48 / 1.28 |
| 2 | 0.1.40.3 | 3274 + 2572 | 375.2 | **17.7** (17.1 17.7 19.3) | 2.19 / 1.93 |
| 3 | 0.1.40.1 | 3275 + 2574 | 386.5 | **34.2** (28.8 34.8 34.2) | 1.49 / 1.18 |
| 4 | 0.1.40.3 | 3274 + 2572 | 375.9 | **17.7** (17.7 17.3 18.5) | 2.28 / 1.97 |
`conversation_snapshot_test` passes on 0.1.40.3 on these GPUs (3901 checks).
**Paired run 2: 0.1.40.2 bisects the range**
| round | engine | slots | prefill | decode median (runs) |
|---|---|---|---:|---|
| 1 | 0.1.40.1 | 3275 + 2574 | 387.4 | 32.7 (29.6 32.7 35.1) |
| 2 | 0.1.40.2 | 3274 + 2572 | 375.2 | **18.0** (16.6 18.0 18.5) |
| 4 | 0.1.40.1 | 3275 + 2574 | 385.9 | 34.5 (34.2 38.1 34.5) |
(Rounds 3 and 5 of that session tested another model.)
**Paired run 3: the 2-line fix**
| round | engine | slots | prefill | decode median (runs) | switch-back (s) |
|---|---|---|---:|---|---|
| 1 | 0.1.40.1 | 3275 + 2574 | 389.5 | 32.1 (28.2 32.1 33.2) | 1.66 / 1.35 |
| 2 | 0.1.40.3 | 3274 + 2572 | 375.9 | **18.2** (16.4 18.2 20.1) | 2.18 / 1.95 |
| 3 | 0.1.40.3 + fix | 3275 + 2574 | 386.3 | **33.0** (28.7 33.0 35.1) | 1.69 / 1.35 |
| 4 | 0.1.40.1 | 3275 + 2574 | 386.2 | 36.2 (36.2 34.7 36.9) | 1.36 / 1.24 |
| 5 | 0.1.40.3 + fix | 3275 + 2574 | 386.3 | **31.8** (31.1 31.8 34.2) | 1.69 / 1.43 |
A later session with the fix built into the full tree (not a probe build) gave 32.4 and 33.9 t/s decode, and
388.7 / 385.8 t/s prefill. **The fix recovers decode, prefill and switch-back fully on this box.**
Across all of these sessions, MTP draft acceptance per 400-token decode is 177–264 of 274–345 in every arm, with no
pattern by engine.
## Results: box 2 (RTX 3080, IQ3_XXS), the control
| round | engine | expert cache slots | prefill 16k (t/s) | decode median (runs) | switch-back (s) |
|---|---|---|---:|---|---|
| 1 | 0.1.40.1 | 2193 | 2140.6 | 73.6 (75.5 70.9 73.6) | 0.67 / 0.57 |
| 2 | 0.1.40.3 | 2191 | 2102.3 | 73.5 (71.8 73.5 77.0) | 0.62 / 0.67 |
| 3 | 0.1.40.3 + fix | 2192 | 2101.6 | 75.9 (72.0 75.9 77.0) | 0.64 / 0.71 |
| 4 | 0.1.40.1 | 2193 | 2105.9 | 71.0 (70.4 73.6 71.0) | 0.63 / 0.68 |
| 5 | 0.1.40.3 + fix | 2192 | 2092.7 | 73.1 (72.6 73.1 75.5) | 0.67 / 0.72 |
| 6 | 0.1.40.3 | 2191 | 2106.8 | 74.6 (72.2 74.6 76.4) | 0.64 / 0.63 |
Mean decode: 0.1.40.1 72.3, 0.1.40.3 74.1, 0.1.40.3 + fix 74.5. All three are within this box's spread.
**The fix is neutral on Ampere,** and 0.1.40.3 did not regress there.
## Ruled out on box 1 (each against stock 0.1.40.3 in the same paired session)
| arm | decode median (t/s) | vs 0.1.40.1 in the same session |
|---|---:|---|
| 0.1.40.3 + `STRATA_STAGER_SLEEP=0` | 18.7 | 32.5 |
| 0.1.40.3 + `STRATA_MMVQ_IL=0` | 18.4 | 32.5 |
| 0.1.40.3 + both | 18.2 | 32.5 |
| 0.1.40.3 `--pipeline-windows 1` | 17.2 | 30.8 (0.1.40.1, also `--pipeline-windows 1`) |
- **`STRATA_MMVQ_IL`:** `il_arch_ok()` already requires sm_80+, so this is expected.
- **`--pipeline-windows 1`:** regresses just as much, so the cause is a kernel, not the window scheduling.
- **The same session's slots with `--pipeline-windows 1`:** 3367 + 2574 on 0.1.40.1, 3365 + 2572 on 0.1.40.3.
- **Code reading of v0.1.40.1..v0.1.40.2** (69 engine commits) found no other default-on change on this path that
applies to sm_61. dp4a is gated `< 610`, the IL path is sm_80+, PDL launches are sm_90+, and the file-tier I/O
changes don't apply (this box uses the in-RAM expert arena).
- **Two CPU expert-path commits,** d6850f3 (IQ2_S/IQ3_S AVX2 decode) and 5208b32 (AVX2 Q8_K quantizer), were not
tested on the GPU. The fix alone recovers the full gap, so they are not needed to explain it.
## Root cause
2e4ddf6 changed `include/strata/kernels/pdl.hpp` for every CUDA build:
```c++
// not __CUDA_ARCH__-dependent: the host pass must see the same kernel signature (MSVC rejects the template stubs
// otherwise), so on CUDA these parameters lose __restrict__ for every architecture
#if defined(__HIPCC__)
#define STRATA_PDL_RESTRICT __restrict__
#else
#define STRATA_PDL_RESTRICT
#endif
...
inline constexpr bool kPdlPrefetch = true; // CUDA: the weights a kernel can load before pdl_wait() are loaded there
```
**The SASS.** `cuobjdump -sass` of the shipped sm_61 engines, summed over all 334 `native_mmvq_multi_kernel`
instances. This is the 2–8-token verify-window projection, so it runs in every speculative decode step:
| engine | instructions | read-only loads (`LDG.E.CI` / `.CONSTANT`) | plain `LDG` |
|---|---:|---:|---:|
| 0.1.40.1 | 231,348 | 19,600 | 400 |
| 0.1.40.2 / 0.1.40.3 | 373,044 (+61%) | 7,020 | **32,636** |
| `2e4ddf6^` (`native_mmvq.cu` built alone) | 231,348 | 19,600 | 400 |
| `2e4ddf6` (`native_mmvq.cu` built alone) | 373,044 | 7,020 | 32,636 |
| 0.1.40.2, `kPdlPrefetch = false` only | 231,348 | 14,098 | 5,902 |
| 0.1.40.3 + the fix below (full build) | 231,348 | 19,600 | 400 |
`native_quantize_q8_1_kernel` also goes from `LDG.nc` to plain `LDG`.
**Why this hurts GP104 and not newer cards.** On GP104 (GTX 1070/1080), as on Maxwell, global loads are cached in
L1/tex **only** through the read-only `LDG.nc` path. A plain `LDG` goes to L2. With `__restrict__` gone, the compiler
can no longer prove the q8_1 activations are read-only, so every block re-reads them from L2. GP100 and sm_70+ cache
ordinary global loads in L1 too. That is consistent with the RTX 3080 result above, and with nobody noticing on
current cards.
**Not covered by the fix:** 2e4ddf6 also replaced the q/gate split's `cudaMemcpy2DAsync` with a `copy_rows_strided`
kernel in `verify.cpp`. The fix doesn't touch that, and it doesn't need to: the gap closes fully without it.
## The fix we run (2 lines)
```diff
diff --git a/include/strata/kernels/pdl.hpp b/include/strata/kernels/pdl.hpp
--- a/include/strata/kernels/pdl.hpp
+++ b/include/strata/kernels/pdl.hpp
@@ -35,16 +35,13 @@ bool pdl_supported();
// not __CUDA_ARCH__-dependent: the host pass must see the same kernel signature (MSVC rejects the template stubs
// otherwise), so on CUDA these parameters lose __restrict__ for every architecture
-#if defined(__HIPCC__)
+// __restrict__ kept on CUDA too (PDL is never enabled below sm_90; STRATA_DF_PDL must stay unset)
#define STRATA_PDL_RESTRICT __restrict__
-#else
-#define STRATA_PDL_RESTRICT
-#endif
#if defined(__HIPCC__)
inline constexpr bool kPdlPrefetch = false; // HIP: no PDL, the kernels keep their plain loops
#else
-inline constexpr bool kPdlPrefetch = true; // CUDA: the weights a kernel can load before pdl_wait() are loaded there
+inline constexpr bool kPdlPrefetch = false; // CUDA: the weights a kernel can load before pdl_wait() are loaded there
/// `kernel` may be launched with the PDL attribute into `stream` now (see the header comment).
bool pdl_launch_ok(const void* kernel, cudaStream_t stream);
#endif
```
**Caveat: this is a workaround, not a proposed upstream patch.**
- **What it does:** it puts 0.1.40.1's code generation back on **every** CUDA arch, sm_90+ included.
- **What it costs on sm_90+:** the PDL prefetch. We have not checked whether `__restrict__` is safe there with
`STRATA_DF_PDL` on, so we keep that unset. Neither of our cards is sm_90.
- **What an upstream fix needs:** to keep 0.1.40.1's loads below sm_90 without changing the host-visible kernel
signature, which is what the header comment says ruled out `__CUDA_ARCH__`.
- **Two possible directions:**
- choose restrict and the plain loop per instantiation (a template parameter set from the launch site's arch
check), rather than per translation unit;
- or read `x` through `__ldg` after `pdl_wait()`. The `asm volatile` wait carries a memory clobber, so the load
can't be hoisted above it.
## Related
- **#1373** (closed by its author; Windows, 2-GPU layer split with an RTX 4080 SUPER): a reported 0.1.40.2 decode
drop that turned out to be a cold-engine measurement artifact. This report follows the paired method and slot
reporting asked for there. Unlike that case, the gap here is about 2× rather than 10–20%, it is the same in every
pair, and a one-file rebuild removes it.
- **#1428** (P40 + RTX 3070, 0.1.40.2): about the auto layer split's choice, not decode. The P40 is also sm_61
consumer-class Pascal (GP102), so its stage runs the same `native_mmvq_multi_kernel` code. It may be affected
Mehr auf der Site
Links zu Install, Modellen, Releases.