Issues / #1153

#1153 [Volta] STRATA_GDN_CHUNK=1 (PR #1098 chain-split) segfaults the engine on a 4-GPU layer split (4×V100)

closed · @justxiami · 1 comentarios · En GitHub

BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quants

Descripción


**Summary.** Applying PR #1098 to v0.1.40.1 and setting `STRATA_GDN_CHUNK=1` on our 4×V100 layer-split setup makes the engine die with SIGSEGV (exit code -11) on any prompt we tested from 8k tokens up. Every request then returns empty (the server restarts the engine, the request fails with 503 / `done: 0 tokens … (error)`). Without the env var the same binary is healthy, and the base v0.1.40/0.1.40.1 engine is unaffected — so this is the opt-in Volta chain-split path itself, not a packaging issue.

**Setup (as reported in the PR family, but on multi-GPU).**

- v0.1.40.1 source + PR #1098 patch applied as-is (applies cleanly except the bench wiring hunk; `STRATA_BUILD_TESTS=OFF`).
- Build: CUDA 12.8, `-DCMAKE_CUDA_ARCHITECTURES=70`, `-DSTRATA_EXPERIMENTAL_SM60=ON`, Release.
- Hardware: 4× V100 16 GB (full NVLink2 mesh), one NUMA node, 59 GB host RAM.
- Engine args: model Qwen3.8-Flash-Next (IQ3_XXS native pack), `--kv int8 --max-context 262144 --layer-split 12,24,36 --split-device 1,2,3 --spec 3 --spec-min-p 0.35 --mtp …` (the usual resident-split config); served through `serve/server.py` at `--port 8095`.
- Only change vs. the healthy control: environment `STRATA_GDN_CHUNK=1`.

**Observed, with `STRATA_GDN_CHUNK=1`:**

```
the engine stopped unexpectedly (exit code -11). … The next request starts the engine again.
done: 0 tokens in 5 s (0.0 tok/s) (error, cancel=False)
```

- First crash already at an 8,074-token prompt; at 11k the HTTP request returns 503 while the server relaunches the engine.
- Full prompt ladder 8k/16k/32k/64k/120k, 3 reps per tier, two fresh engine starts: **15/15 requests produce no completion and no usage** (the answer-needle check consequently misses at every tier). Engine log during the window contains only the repeated `done: 0 tokens … (error)` lines.
- Not the kernel OOM killer: `dmesg` has no OOM record for these events (the last OOM kill on this host was a different, unrelated unit days earlier); exit -11 = SIGSEGV.

**Controls, same day, same tree and binary:**

- `arm-bin` rebuilt from v0.1.40.1 + #1098, `STRATA_GDN_CHUNK` unset: healthy (`done: 12 tokens in 8 s (122.0 tok/s)` on the probe that 503'd a minute earlier after a restart with the env var set; expert cache 99.2%).
- Base v0.1.40 engine: the identical needle ladder is 10/10 with normal timings.

**Not tested:** shorter prompts (<8k), a single-GPU setup (our box is always a 4-card layer split; the PR's bench numbers look single-GPU), and `STRATA_GDN_CHUNK=2` (the FP32-only FMA scan variant).

**Questions.** Is the chain-split path assumed single-device — e.g. does the chunked recurrence kernel assume the whole prompt path lives on one card, while with `--layer-split` the prompt recurrence runs across stage hand-offs? Happy to test a candidate patch or run `STRATA_GDN_CHUNK=2` / a single-GPU config if that helps localize it. Engine logs and the raw ladder JSONs are available.

En el sitio

Enlaces a install, modelos, releases.