Issues / #929

#929 `--batch` with `--layer-split` exits ("verify batch: layer 34 never rang") with a 3070 first and a P40 second; works with the P40 first (CUDA, v0.1.39)

open · @paulhothersall · 3 コメント · GitHub で見る

Setup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants

本文

**On v0.1.39 (CUDA 12.9), `--layer-split 34 --batch 4 --batch-groups 1` on an RTX 3070 (GPU 0, layers 0-33) + Tesla P40 (GPU 1, layers 34-47) exits when a second request arrives: `verify batch: layer 34 never rang (graph finished)`, exit code 1. With the cards in the other order (P40 first) the same command works. The batch-groups flag, the slot count, the weight carve and the context size do not change it.**

Setup: v0.1.39 (6f32ec0) built from source, CUDA 12.9, `-DCMAKE_CUDA_ARCHITECTURES="61;86"`, RTX 3070 8 GB (cc 8.6) + Tesla P40 24 GB (cc 6.1), PCIe 3.0 x8 each, no P2P, 64 GB RAM, Qwen3.8-Flash-Next IQ3_XXS, `--spec 4 --mtp`, `--expert-cache auto --prefill auto`, served through `serve/server.py`.

Repro: start the server with `--layer-split 34 --batch 4 --batch-groups 1` and send two chat requests at once (about 600-token prompts, 200 new tokens). The first request gets 2 tokens, then the engine reports `verify batch: layer 34 never rang (graph finished)` and exits (code 1). Last engine log line before the exit: `strata verify: captured the batch window over slots 0`.

| change | result |
|---|---|
| 3070 first, `--batch 4 --batch-groups 1`, context 36,864, `STRATA_STAGE_TRIM=1` | crashes |
| same, context 8,192, no carve | crashes |
| same, `--batch 4` without `--batch-groups` | crashes |
| same, `--batch 2 --batch-groups 1` | crashes |
| **P40 first (P40 as GPU 0, 3070 as GPU 1), `--layer-split 34 --batch 4 --batch-groups 1`, context 8,192** | **works**: 2 streams, 386 chunks in 17.1 s (20.0 / 14.5 chunks/s per stream) |
| P40 alone, `--batch 4` | works |

Related: #890 (HIP, 2x R9700, `parallel: 2` + `--layer-split`) ends at the same point (`captured the batch window over slots 0,1`, exit code 1, no error text) and reports the same setup works on 3x RTX 5070 Ti with CUDA. #867 has the same message text on a SYCL port, where the issue traces it to a `cudaStreamQuery` mapping; I am not claiming the same cause. My reading, not checked in the code: the stage on the Pascal card does not complete its part of the batch window when it is the second stage.

Not tested: `--spec 0` (my attempt failed to load), other card pairs, P2P, vision on. I can send the engine logs or run a debug build.

関連リンク

インストール・モデル・リリースへの站内リンク。