Issues / #1757
#1757 Layer split (RTX 3090 + 2x Tesla P40, Windows): every batched prefill chunk aborts with `prefill gemm: cublasGemmEx: cuBLAS status 14` — 100% on 0.1.40.4 / 0.1.41, single-GPU unaffected
open · @vgame86 · 0 コメント · GitHub で見る
Setup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux
本文
## Summary With a 3-GPU layer split (RTX 3090 + 2x Tesla P40, Pascal sm_61) on Windows, **every request whose rendered prompt exceeds the short-read threshold (~64 tokens) aborts the engine on its first batched prefill chunk** with: ``` prefill gemm: cublasGemmEx: cuBLAS status 14 ``` (`CUBLAS_STATUS_INTERNAL_ERROR`; the engine then `std::exit(1)`s via `ck()` in `src/prefill/gemm.cu:81`, and the server answers the client with `the engine stopped unexpectedly (exit code 1); the next request restarts it`.) This is **not** tool/function-calling related and **not** a small-chunk issue: a plain multi-hundred-token Chinese prompt **without** `tools` crashes identically. Tools only made the visible difference because they render a 3-token message into a 277-token prompt, crossing the 64-token short-read threshold. ## Reproduction - Windows 11 (WDDM), NVIDIA driver **551.61**, CUDA 12.4 (driver-reported) - GPUs: 0 = RTX 3090 (sm_86, display), 1-2 = Tesla P40 24 GB (sm_61) - Engine: ready-made **0.1.41** release (CUDA 12.9 experimental build, `BUILD.json`: archs 60/61/70/75/80/86/89, `"experimental": true`) - Model: Qwen3.8-Flash-Next, IQ3_S (GSQ-RCO), 48 layers Serve config (argv as logged): ``` strata.exe --serve --pack .../iq3_s --native ...-00001-of-00002.gguf --ple-gguf ...-00002-of-00002.gguf --expert-profile .../expert-profile.bin --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp .../rt --max-context 65536 --kv int8 --resident-experts --resident-budget-gib 16 --layer-split 32,40 --vram-reserve-mib 400 --vision ``` (i.e. `3090: layers 0-31`, `P40 #1: 32-39`, `P40 #2: 40-47 + head`). Then send any chat request whose rendered prompt is > 64 tokens, e.g. a ~170-token question. With `STRATA_TRACE=1` in the config's `env`: ``` strata trace: request 174 0 strata trace: prompt start 173 -1 strata trace: lent 149 slots for 256 tokens in 0.8 ms strata trace: prompt chunk 0 of 169 prefill gemm: cublasGemmEx: cuBLAS status 14 ``` Second session, identical: ``` strata trace: request 277 0 strata trace: prompt start 276 -1 strata trace: prompt chunk 0 of 272 prefill gemm: cublasGemmEx: cuBLAS status 14 ``` ## What I can rule out / confirm | Configuration | Result | |---|---| | 3-GPU split (`32,40`), prompt ≤ 64 tokens (verify-window path) | ✅ works, repeatedly | | 3-GPU split (`32,40`), prompt > 64 tokens (batched prefill) | ❌ **100% abort on first chunk, every time, every session** (0.1.40.4 and 0.1.41) | | 3-GPU split, 0.1.40.3, same shape | ⚠️ intermittent — some sessions completed full conversations (1000+ token generations) before aborting; most aborted after 1-2 requests | | Single GPU (3090 only, no `--layer-split`), same model/pack, long prompts (≥ 277 tokens, incl. multi-page answers) | ✅ works, 0.1.40.3 and 0.1.41 | Additional checks: - The aborting call is the **BF16** `Gemm::bf16` path (`cublasGemmEx`, `CUDA_R_16BF`, no `f16` suffix in the error line). The Pascal fp32 fallback in the same function (`"cublasSgemm (bf16 on Pascal)"`) **never appears in any of the logs** (0 occurrences across ~25 crash sessions), i.e. when the batched prefill runs on the P40 stages the code is reaching the raw BF16 `cublasGemmEx` — the path that `#395` measured as unsupported on a P40. If that reading is right, the question is why `bf16_path()`/`grow()` silently falls through to the BF16 call instead of the `cublasSgemm` fp32 route on sm_61. - Free VRAM at crash time was low but not zero in most sessions (100/84/118 MiB; one session 0 MiB — the engine itself printed the `LOW: requests may stall` warning). Short reads (≤64 tokens) on the same sessions work, so it doesn't look like a plain OOM; but I have not isolated this (see "not yet tried"). - No `nvlddmkm`/Display events in the Windows System log around the crash times — no TDR, no driver fault. It is a clean cuBLAS error return. - Unrelated to `--resident-experts` per se: 0.1.40.3 completed long replies with the same resident-RAM mode before aborting later. ## Not yet tried (happy to run) 1. `STRATA_BF16_TC=2` (force the FP16 test path) on the split, to see whether the crash follows the BF16 call or the split itself. 2. 2-GPU split (3090 + one P40) to see whether a single Pascal stage is enough to trigger it. 3. 3-GPU split with `--short-read 0` and `--short-read 2048` to move the threshold. 4. Same setup on a P40-only pair (no 3090) — I don't have a second sm_86 card, but community testers in #875/#1239 do. 5. Linux, if you think WDDM is involved. Full per-line logs of two sessions (startup → crash, `STRATA_TRACE=1`) are available; ask and I'll paste/attach. ## Related - #1650 — same abort signature (`cublasGemmEx ... status 14`) but on sub-32-token chunks with `--short-read 0`, reported on RTX 2080 Ti + P100. Possibly the same underlying GEMM-path bug, different trigger. - #395 — "cuBLAS has no BF16 GEMM for Pascal, measured NOT_SUPPORTED on a P40"; the code comment in `src/prefill/gemm.cu` routes Pascal through `cublasSgemm` fp32 — which is exactly the path we never observe taken. - #875 / #929 / #1239 / #1469 — other P40-in-a-layer-split reports for reference. ## Workaround Run the model on the 3090 alone (drop `--layer-split` and the P40s from `gpu`). Verified stable for long prompts and function-calling requests. Cost: no speed benefit from the P40s.
関連リンク
インストール・モデル・リリースへの站内リンク。