Issues / #1470

#1470 Pascal (sm_61): decode halves since 0.1.40.2. 2e4ddf6 drops `__restrict__` and peels the MMVQ loop on every CUDA arch

closed · @lineape · 1 comentários · No GitHub

Setup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityWindowsLinux

Descrição

## Summary

On a GTX 1080 + GTX 1070 layer split, decode goes from about 33 t/s on 0.1.40.1 to about 18 t/s on 0.1.40.2 and
0.1.40.3. Prefill drops about 3%. Expert-cache slot counts, draft acceptance and hit rates are the same in both arms.

The cause is the compile-time part of 2e4ddf6 ("verify window: PDL on sm_90+, …", #904). In
`include/strata/kernels/pdl.hpp`, `STRATA_PDL_RESTRICT` became empty and `kPdlPrefetch` became `true` on **all** CUDA
architectures, not just sm_90+. On sm_61 the multi-column MMVQ kernel loses its read-only (`LDG.nc`) loads. Restoring
both (a 2-line change, below) brings 0.1.40.3 back to 0.1.40.1's speed on Pascal. On an RTX 3080 it makes no
measurable difference.

## Hardware and config

**Box 1 (regressed):**
- GPUs: GTX 1080 8 GB (CUDA0, layers 0–25) + GTX 1070 8 GB (CUDA1, layers 26–47 + head), both GP104, sm_61, PCIe 3.0 x16.
- Host: Xeon Gold 6230. The engine log says "no AVX-512: the expert kernels run on AVX-2". Linux (NixOS), driver 580.178.04.
- Build: CUDA 12.8, gcc 14.4, `STRATA_EXPERIMENTAL_SM60=1`, archs `[61]`, `setup.sh --build`. The CMakeCache of
  the two versions is identical apart from the source.
- Model: Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS, native pack.
- Engine flags: `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <rt> --max-context 524288 --kv int8 --kv-resident 32768 --conversation-cache-mib 16384 --conversation-cache-slots 8 --layer-split 26 --trim-stage-weights --pipeline-windows 2 --pcie-frac 0.10 --rope-scaling yarn --rope-scale 2`.
- Env: `STRATA_IQ256_GATHER=1 STRATA_RING_BYTES=0`.
- Expert cache: PROFILE policy, pre-filled from `data/expert-profile.bin`, no eviction.

**Box 2 (control):**
- GPU: one RTX 3080 (sm_86). Host: Ryzen 9 7945HX. Windows.
- Build: local, CUDA 13.0 + VS 2022 Build Tools, archs `[86]`.
- Model: Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS.
- Engine flags: `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.70 --mtp <rt> --max-context 524288 --kv int8 --kv-resident 32768 --rope-scaling yarn --rope-scale 2 --conversation-cache-mib 12288 --conversation-cache-slots 6 --vision --pcie-frac 0.35 --prefix-cache-dir <one dir per arm>`.
- Our engine there also carries #1090 (prefix cache) in all three arms.

## Method

This follows the paired method from #1373: arms alternate in one session, nothing else changes between them, and
each arm's `expert cache N slots` line is listed.

- **Order:** A, B, A, B (or A, B, C, A, C, B). The same serve layer and the same config throughout; only the engine
  binary changes, or one env var.
- **Load:** each round loads the engine fresh. With the PROFILE policy there is no adaptive warm-up of the cache,
  but the first decode after a load is still 10–20% slower. So each round reports the median of three decodes, and
  the raw runs are listed too.
- **Decode:** 400 tokens, greedy (`temperature: 0`), from a ~115-token prompt (a fresh UUID prefix each request, so
  no cache reuse), three requests per round. The figure is Strata's own `timings.predicted_per_second`, median of three.
- **Prefill:** a ~15.6k-token prompt, fresh UUID prefix, three requests, median `prompt_per_second`.
- **Switch-back:** two ~11k-token conversations in the order A, B, A, B. The figure is the wall time of the last two
  requests. Each restores a parked conversation (`cache_n` ≈ 10,965), so it times the snapshot restore plus a
  short verify.

## Results: box 1 (GTX 1080 + GTX 1070, IQ2_XS)

**Paired run 1: 0.1.40.1 vs stock 0.1.40.3**

| round | engine | expert cache slots (CUDA0 + CUDA1) | prefill 16k (t/s) | decode median (runs) | switch-back (s) |
|---|---|---|---:|---|---|
| 1 | 0.1.40.1 | 3275 + 2574 | 389.9 | **34.3** (27.6 34.6 34.3) | 1.48 / 1.28 |
| 2 | 0.1.40.3 | 3274 + 2572 | 375.2 | **17.7** (17.1 17.7 19.3) | 2.19 / 1.93 |
| 3 | 0.1.40.1 | 3275 + 2574 | 386.5 | **34.2** (28.8 34.8 34.2) | 1.49 / 1.18 |
| 4 | 0.1.40.3 | 3274 + 2572 | 375.9 | **17.7** (17.7 17.3 18.5) | 2.28 / 1.97 |

`conversation_snapshot_test` passes on 0.1.40.3 on these GPUs (3901 checks).

**Paired run 2: 0.1.40.2 bisects the range**

| round | engine | slots | prefill | decode median (runs) |
|---|---|---|---:|---|
| 1 | 0.1.40.1 | 3275 + 2574 | 387.4 | 32.7 (29.6 32.7 35.1) |
| 2 | 0.1.40.2 | 3274 + 2572 | 375.2 | **18.0** (16.6 18.0 18.5) |
| 4 | 0.1.40.1 | 3275 + 2574 | 385.9 | 34.5 (34.2 38.1 34.5) |

(Rounds 3 and 5 of that session tested another model.)

**Paired run 3: the 2-line fix**

| round | engine | slots | prefill | decode median (runs) | switch-back (s) |
|---|---|---|---:|---|---|
| 1 | 0.1.40.1 | 3275 + 2574 | 389.5 | 32.1 (28.2 32.1 33.2) | 1.66 / 1.35 |
| 2 | 0.1.40.3 | 3274 + 2572 | 375.9 | **18.2** (16.4 18.2 20.1) | 2.18 / 1.95 |
| 3 | 0.1.40.3 + fix | 3275 + 2574 | 386.3 | **33.0** (28.7 33.0 35.1) | 1.69 / 1.35 |
| 4 | 0.1.40.1 | 3275 + 2574 | 386.2 | 36.2 (36.2 34.7 36.9) | 1.36 / 1.24 |
| 5 | 0.1.40.3 + fix | 3275 + 2574 | 386.3 | **31.8** (31.1 31.8 34.2) | 1.69 / 1.43 |

A later session with the fix built into the full tree (not a probe build) gave 32.4 and 33.9 t/s decode, and
388.7 / 385.8 t/s prefill. **The fix recovers decode, prefill and switch-back fully on this box.**

Across all of these sessions, MTP draft acceptance per 400-token decode is 177–264 of 274–345 in every arm, with no
pattern by engine.

## Results: box 2 (RTX 3080, IQ3_XXS), the control

| round | engine | expert cache slots | prefill 16k (t/s) | decode median (runs) | switch-back (s) |
|---|---|---|---:|---|---|
| 1 | 0.1.40.1 | 2193 | 2140.6 | 73.6 (75.5 70.9 73.6) | 0.67 / 0.57 |
| 2 | 0.1.40.3 | 2191 | 2102.3 | 73.5 (71.8 73.5 77.0) | 0.62 / 0.67 |
| 3 | 0.1.40.3 + fix | 2192 | 2101.6 | 75.9 (72.0 75.9 77.0) | 0.64 / 0.71 |
| 4 | 0.1.40.1 | 2193 | 2105.9 | 71.0 (70.4 73.6 71.0) | 0.63 / 0.68 |
| 5 | 0.1.40.3 + fix | 2192 | 2092.7 | 73.1 (72.6 73.1 75.5) | 0.67 / 0.72 |
| 6 | 0.1.40.3 | 2191 | 2106.8 | 74.6 (72.2 74.6 76.4) | 0.64 / 0.63 |

Mean decode: 0.1.40.1 72.3, 0.1.40.3 74.1, 0.1.40.3 + fix 74.5. All three are within this box's spread.
**The fix is neutral on Ampere,** and 0.1.40.3 did not regress there.

## Ruled out on box 1 (each against stock 0.1.40.3 in the same paired session)

| arm | decode median (t/s) | vs 0.1.40.1 in the same session |
|---|---:|---|
| 0.1.40.3 + `STRATA_STAGER_SLEEP=0` | 18.7 | 32.5 |
| 0.1.40.3 + `STRATA_MMVQ_IL=0` | 18.4 | 32.5 |
| 0.1.40.3 + both | 18.2 | 32.5 |
| 0.1.40.3 `--pipeline-windows 1` | 17.2 | 30.8 (0.1.40.1, also `--pipeline-windows 1`) |

- **`STRATA_MMVQ_IL`:** `il_arch_ok()` already requires sm_80+, so this is expected.
- **`--pipeline-windows 1`:** regresses just as much, so the cause is a kernel, not the window scheduling.
- **The same session's slots with `--pipeline-windows 1`:** 3367 + 2574 on 0.1.40.1, 3365 + 2572 on 0.1.40.3.
- **Code reading of v0.1.40.1..v0.1.40.2** (69 engine commits) found no other default-on change on this path that
  applies to sm_61. dp4a is gated `< 610`, the IL path is sm_80+, PDL launches are sm_90+, and the file-tier I/O
  changes don't apply (this box uses the in-RAM expert arena).
- **Two CPU expert-path commits,** d6850f3 (IQ2_S/IQ3_S AVX2 decode) and 5208b32 (AVX2 Q8_K quantizer), were not
  tested on the GPU. The fix alone recovers the full gap, so they are not needed to explain it.

## Root cause

2e4ddf6 changed `include/strata/kernels/pdl.hpp` for every CUDA build:

```c++
// not __CUDA_ARCH__-dependent: the host pass must see the same kernel signature (MSVC rejects the template stubs
// otherwise), so on CUDA these parameters lose __restrict__ for every architecture
#if defined(__HIPCC__)
#define STRATA_PDL_RESTRICT __restrict__
#else
#define STRATA_PDL_RESTRICT
#endif
...
inline constexpr bool kPdlPrefetch = true;    // CUDA: the weights a kernel can load before pdl_wait() are loaded there
```

**The SASS.** `cuobjdump -sass` of the shipped sm_61 engines, summed over all 334 `native_mmvq_multi_kernel`
instances. This is the 2–8-token verify-window projection, so it runs in every speculative decode step:

| engine | instructions | read-only loads (`LDG.E.CI` / `.CONSTANT`) | plain `LDG` |
|---|---:|---:|---:|
| 0.1.40.1 | 231,348 | 19,600 | 400 |
| 0.1.40.2 / 0.1.40.3 | 373,044 (+61%) | 7,020 | **32,636** |
| `2e4ddf6^` (`native_mmvq.cu` built alone) | 231,348 | 19,600 | 400 |
| `2e4ddf6` (`native_mmvq.cu` built alone) | 373,044 | 7,020 | 32,636 |
| 0.1.40.2, `kPdlPrefetch = false` only | 231,348 | 14,098 | 5,902 |
| 0.1.40.3 + the fix below (full build) | 231,348 | 19,600 | 400 |

`native_quantize_q8_1_kernel` also goes from `LDG.nc` to plain `LDG`.

**Why this hurts GP104 and not newer cards.** On GP104 (GTX 1070/1080), as on Maxwell, global loads are cached in
L1/tex **only** through the read-only `LDG.nc` path. A plain `LDG` goes to L2. With `__restrict__` gone, the compiler
can no longer prove the q8_1 activations are read-only, so every block re-reads them from L2. GP100 and sm_70+ cache
ordinary global loads in L1 too. That is consistent with the RTX 3080 result above, and with nobody noticing on
current cards.

**Not covered by the fix:** 2e4ddf6 also replaced the q/gate split's `cudaMemcpy2DAsync` with a `copy_rows_strided`
kernel in `verify.cpp`. The fix doesn't touch that, and it doesn't need to: the gap closes fully without it.

## The fix we run (2 lines)

```diff
diff --git a/include/strata/kernels/pdl.hpp b/include/strata/kernels/pdl.hpp
--- a/include/strata/kernels/pdl.hpp
+++ b/include/strata/kernels/pdl.hpp
@@ -35,16 +35,13 @@ bool pdl_supported();
 
 // not __CUDA_ARCH__-dependent: the host pass must see the same kernel signature (MSVC rejects the template stubs
 // otherwise), so on CUDA these parameters lose __restrict__ for every architecture
-#if defined(__HIPCC__)
+// __restrict__ kept on CUDA too (PDL is never enabled below sm_90; STRATA_DF_PDL must stay unset)
 #define STRATA_PDL_RESTRICT __restrict__
-#else
-#define STRATA_PDL_RESTRICT
-#endif
 
 #if defined(__HIPCC__)
 inline constexpr bool kPdlPrefetch = false;   // HIP: no PDL, the kernels keep their plain loops
 #else
-inline constexpr bool kPdlPrefetch = true;    // CUDA: the weights a kernel can load before pdl_wait() are loaded there
+inline constexpr bool kPdlPrefetch = false;   // CUDA: the weights a kernel can load before pdl_wait() are loaded there
 /// `kernel` may be launched with the PDL attribute into `stream` now (see the header comment).
 bool pdl_launch_ok(const void* kernel, cudaStream_t stream);
 #endif
```

**Caveat: this is a workaround, not a proposed upstream patch.**
- **What it does:** it puts 0.1.40.1's code generation back on **every** CUDA arch, sm_90+ included.
- **What it costs on sm_90+:** the PDL prefetch. We have not checked whether `__restrict__` is safe there with
  `STRATA_DF_PDL` on, so we keep that unset. Neither of our cards is sm_90.
- **What an upstream fix needs:** to keep 0.1.40.1's loads below sm_90 without changing the host-visible kernel
  signature, which is what the header comment says ruled out `__CUDA_ARCH__`.
- **Two possible directions:**
  - choose restrict and the plain loop per instantiation (a template parameter set from the launch site's arch
    check), rather than per translation unit;
  - or read `x` through `__ldg` after `pdl_wait()`. The `asm volatile` wait carries a memory clobber, so the load
    can't be hoisted above it.

## Related

- **#1373** (closed by its author; Windows, 2-GPU layer split with an RTX 4080 SUPER): a reported 0.1.40.2 decode
  drop that turned out to be a cold-engine measurement artifact. This report follows the paired method and slot
  reporting asked for there. Unlike that case, the gap here is about 2× rather than 10–20%, it is the same in every
  pair, and a one-file rebuild removes it.
- **#1428** (P40 + RTX 3070, 0.1.40.2): about the auto layer split's choice, not decode. The P40 is also sm_61
  consumer-class Pascal (GP102), so its stage runs the same `native_mmvq_multi_kernel` code. It may be affected
  

No site

Links install, modelos, releases.