Issues / #1153

#1153 [Volta] STRATA_GDN_CHUNK=1 (PR #1098 chain-split) segfaults the engine on a 4-GPU layer split (4×V100)

closed · @justxiami · 1 comments · View on GitHub

BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quants

Description


**Summary.** Applying PR #1098 to v0.1.40.1 and setting `STRATA_GDN_CHUNK=1` on our 4×V100 layer-split setup makes the engine die with SIGSEGV (exit code -11) on any prompt we tested from 8k tokens up. Every request then returns empty (the server restarts the engine, the request fails with 503 / `done: 0 tokens … (error)`). Without the env var the same binary is healthy, and the base v0.1.40/0.1.40.1 engine is unaffected — so this is the opt-in Volta chain-split path itself, not a packaging issue.

**Setup (as reported in the PR family, but on multi-GPU).**

- v0.1.40.1 source + PR #1098 patch applied as-is (applies cleanly except the bench wiring hunk; `STRATA_BUILD_TESTS=OFF`).
- Build: CUDA 12.8, `-DCMAKE_CUDA_ARCHITECTURES=70`, `-DSTRATA_EXPERIMENTAL_SM60=ON`, Release.
- Hardware: 4× V100 16 GB (full NVLink2 mesh), one NUMA node, 59 GB host RAM.
- Engine args: model Qwen3.8-Flash-Next (IQ3_XXS native pack), `--kv int8 --max-context 262144 --layer-split 12,24,36 --split-device 1,2,3 --spec 3 --spec-min-p 0.35 --mtp …` (the usual resident-split config); served through `serve/server.py` at `--port 8095`.
- Only change vs. the healthy control: environment `STRATA_GDN_CHUNK=1`.

**Observed, with `STRATA_GDN_CHUNK=1`:**

```
the engine stopped unexpectedly (exit code -11). … The next request starts the engine again.
done: 0 tokens in 5 s (0.0 tok/s) (error, cancel=False)
```

- First crash already at an 8,074-token prompt; at 11k the HTTP request returns 503 while the server relaunches the engine.
- Full prompt ladder 8k/16k/32k/64k/120k, 3 reps per tier, two fresh engine starts: **15/15 requests produce no completion and no usage** (the answer-needle check consequently misses at every tier). Engine log during the window contains only the repeated `done: 0 tokens … (error)` lines.
- Not the kernel OOM killer: `dmesg` has no OOM record for these events (the last OOM kill on this host was a different, unrelated unit days earlier); exit -11 = SIGSEGV.

**Controls, same day, same tree and binary:**

- `arm-bin` rebuilt from v0.1.40.1 + #1098, `STRATA_GDN_CHUNK` unset: healthy (`done: 12 tokens in 8 s (122.0 tok/s)` on the probe that 503'd a minute earlier after a restart with the env var set; expert cache 99.2%).
- Base v0.1.40 engine: the identical needle ladder is 10/10 with normal timings.

**Not tested:** shorter prompts (<8k), a single-GPU setup (our box is always a 4-card layer split; the PR's bench numbers look single-GPU), and `STRATA_GDN_CHUNK=2` (the FP32-only FMA scan variant).

**Questions.** Is the chain-split path assumed single-device — e.g. does the chunked recurrence kernel assume the whole prompt path lives on one card, while with `--layer-split` the prompt recurrence runs across stage hand-offs? Happy to test a candidate patch or run `STRATA_GDN_CHUNK=2` / a single-GPU config if that helps localize it. Engine logs and the raw ladder JSONs are available.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.