Issues / #542

#542 sm_120 + IQ packs: ~4x prefill regression 0.1.34-0.1.36 when libcudart's ABI is mismatched - the #420 fits() gate silently disables MMQ (smpbo reads 1)

closed · @gravitomagnetic · 3 评论 · 在 GitHub 查看

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quants

描述

# sm_120 + IQ packs: ~4x prefill regression 0.1.34–0.1.36 — the #420 `fits()` gate silently disables MMQ when libcudart's ABI is mismatched (`smpbo` reads 1)

Reported by [gravitomagnetic](https://github.com/gravitomagnetic). Same machine and config as my community benchmark (RTX 5090, 1M-context agent workload). This one comes with a root cause and a reproduction that needs no GPU workload at all.

## Symptom

Cold prefill drops ~4x from 0.1.32 onward; decode is unaffected. Engine-reported `timings`, medians of 3 cold runs (unique-salted prompts), same box, same config:

| engine | 4k | 32k | 128k | decode |
|---|---:|---:|---:|---:|
| 0.1.31 | 2,532 | 3,285 | 3,384 | ~110 |
| 0.1.34 | — | 1,087 | — | ~107 |
| 0.1.36 | 595 | 1,083 | 1,124 | ~130 |
| **0.1.36, link fixed (below)** | **2,560** | **3,281** | **3,379** | ~120 |

The startup log on the slow engines prints, per IQ type:

```
strata: prompt kernels: llama.cpp's MMQ has no tile for iq3_s (1280 rows) on GPU 0
(cc 1200, 1 bytes of shared memory per block): that product takes the non-MMQ path (#420)
```

**"1 bytes of shared memory per block"** — an RTX 5090 reports 101,376. That impossible value is the whole story.

## Setup

| | |
|---|---|
| Model | Qwen3.8-Flash-Next IQ3_S (GSQ-RCO split GGUF) |
| GPU | single RTX 5090 (sm_120), driver 595.91.07 |
| Build | local, `-DCMAKE_CUDA_ARCHITECTURES=120`, nvcc from a CUDA 13.4 kit |
| Config | `--max-context 1048576 --rope-scaling yarn --rope-scale 4 --kv int8 --kv-resident 32768 --expert-cache 7500 --mtp --spec 4`, prefill `auto` |

## Root cause

My 13.4 kit had a **dangling `libcudart.so` symlink** (the kit shipped only `libcudart_static.a`). CMake's FindCUDAToolkit therefore resolved cudart to the system's **CUDA 12.4** `libcudart.so.12` — and every engine binary I built (0.1.31 through 0.1.36) carries `NEEDED libcudart.so.12` while its host code was compiled against **13.x headers**.

`cudaGetDeviceProperties` fills a struct whose layout changed between 12.4 and 13.x. Reads land at wrong offsets. A probe linked against the engine's own `libstrata_mmq.a` (which contains `src/prefill/ggml_cuda_host.cu`'s `ggml_cuda_info()`) shows:

```
SHIM:  cc=1200 nsm=1   smpb=49152 smpbo=1        # what the engine sees
FRESH: cc=12.0 nsm=1   smpb=49152 smpbo=4294967297  # 0x1_00000001 — a shifted struct
```

(`nsm=1`, `integrated=1` — all garbage. The enum-based `cudaDeviceGetAttribute` calls on the *same* old runtime return the truth: `smpbo=101376`, `nsm=170`.)

`mmq::fits()` (#420, commit `8404117`) compares `mmq_get_nbytes_shared(config)` (~41–48 KB) against `info.devices[id].smpbo`. With `smpbo` reading **1**, it returns false for every IQ type, and all expert products silently take the dequant-FP16 fallback — the 4x prefill regression.

**Why 0.1.31 was unaffected:** its MMQ decision is the compile-time type list in `supported()`; it never reads `smpbo`. The ABI mismatch was latent in all my builds; the #420 runtime gate is what weaponized it.

## Reproduction (no model, no workload)

Compile a probe that calls `ggml_cuda_info()` and prints `devices[0].smpbo`, linked against `libstrata_mmq.a` + `libggml-base.a`, built with a toolkit whose `libcudart.so` resolves to an older runtime than the headers. `smpbo` reads 1; replicating the `fits()` loop returns false for iq3_s/iq3_xxs/iq2_s/iq4_xs at 1280 and 5120 rows. Relinked against the matching `libcudart.so.13`: `smpbo=101376`, `fits()=true` for all twelve combinations.

## The fix, locally

Repaired the kit symlink and rebuilt with `-DCUDAToolkit_ROOT=<kit>` (without it CMake still finds the system 12.4 dev symlink). `NEEDED libcudart.so.13` → full speed restored (last table row).

## Why this matters upstream

The trigger here was my broken kit, but the failure mode is nasty in general:

1. **Silent.** Nothing says "your runtime and headers disagree." The engine just runs 4x slower on prompts, and the only clue is a log line that *prints the garbage value* ("1 bytes of shared memory") as if it were a real GPU property.
2. **Easy to hit.** Mixed CUDA toolkits are common; llama.cpp upstream itself moved to `cudaDeviceGetAttribute` for exactly this ABI fragility.
3. Suggested hardening in `src/prefill/ggml_cuda_host.cu` / `moe_mmq.cu`:
   - read device properties via `cudaDeviceGetAttribute` (ABI-stable) instead of `cudaGetDeviceProperties`, or at least for `cc`, `nsm`, `smpbo`;
   - sanity-check the result: `smpbo < 8192` or `nsm == 0` on any CUDA card is physically impossible — log a loud "compiled against newer CUDA headers than the libcudart you linked (ABI mismatch?)" warning instead of the misleading "MMQ has no tile" line;
   - optionally print the actual `mmq_get_nbytes_shared` value in the fallback message so the comparison is auditable.

Happy to test any candidate patch on this box — the probe and the 3-run cold-prefill harness are reproducible in minutes.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。