Issues / #1153
#1153 [Volta] STRATA_GDN_CHUNK=1 (PR #1098 chain-split) segfaults the engine on a 4-GPU layer split (4×V100)
closed · @justxiami · 1 コメント · GitHub で見る
BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quants
本文
**Summary.** Applying PR #1098 to v0.1.40.1 and setting `STRATA_GDN_CHUNK=1` on our 4×V100 layer-split setup makes the engine die with SIGSEGV (exit code -11) on any prompt we tested from 8k tokens up. Every request then returns empty (the server restarts the engine, the request fails with 503 / `done: 0 tokens … (error)`). Without the env var the same binary is healthy, and the base v0.1.40/0.1.40.1 engine is unaffected — so this is the opt-in Volta chain-split path itself, not a packaging issue. **Setup (as reported in the PR family, but on multi-GPU).** - v0.1.40.1 source + PR #1098 patch applied as-is (applies cleanly except the bench wiring hunk; `STRATA_BUILD_TESTS=OFF`). - Build: CUDA 12.8, `-DCMAKE_CUDA_ARCHITECTURES=70`, `-DSTRATA_EXPERIMENTAL_SM60=ON`, Release. - Hardware: 4× V100 16 GB (full NVLink2 mesh), one NUMA node, 59 GB host RAM. - Engine args: model Qwen3.8-Flash-Next (IQ3_XXS native pack), `--kv int8 --max-context 262144 --layer-split 12,24,36 --split-device 1,2,3 --spec 3 --spec-min-p 0.35 --mtp …` (the usual resident-split config); served through `serve/server.py` at `--port 8095`. - Only change vs. the healthy control: environment `STRATA_GDN_CHUNK=1`. **Observed, with `STRATA_GDN_CHUNK=1`:** ``` the engine stopped unexpectedly (exit code -11). … The next request starts the engine again. done: 0 tokens in 5 s (0.0 tok/s) (error, cancel=False) ``` - First crash already at an 8,074-token prompt; at 11k the HTTP request returns 503 while the server relaunches the engine. - Full prompt ladder 8k/16k/32k/64k/120k, 3 reps per tier, two fresh engine starts: **15/15 requests produce no completion and no usage** (the answer-needle check consequently misses at every tier). Engine log during the window contains only the repeated `done: 0 tokens … (error)` lines. - Not the kernel OOM killer: `dmesg` has no OOM record for these events (the last OOM kill on this host was a different, unrelated unit days earlier); exit -11 = SIGSEGV. **Controls, same day, same tree and binary:** - `arm-bin` rebuilt from v0.1.40.1 + #1098, `STRATA_GDN_CHUNK` unset: healthy (`done: 12 tokens in 8 s (122.0 tok/s)` on the probe that 503'd a minute earlier after a restart with the env var set; expert cache 99.2%). - Base v0.1.40 engine: the identical needle ladder is 10/10 with normal timings. **Not tested:** shorter prompts (<8k), a single-GPU setup (our box is always a 4-card layer split; the PR's bench numbers look single-GPU), and `STRATA_GDN_CHUNK=2` (the FP32-only FMA scan variant). **Questions.** Is the chain-split path assumed single-device — e.g. does the chunked recurrence kernel assume the whole prompt path lives on one card, while with `--layer-split` the prompt recurrence runs across stage hand-offs? Happy to test a candidate patch or run `STRATA_GDN_CHUNK=2` / a single-GPU config if that helps localize it. Engine logs and the raw ladder JSONs are available.
関連リンク
インストール・モデル・リリースへの站内リンク。