Issues / #964

#964 SM75 (2080Ti) dual-GPU: "layer X never rang (illegal memory access)" and verify-window hang on IQ3_S

open · @dehuiyede · 2 Kommentare · Auf GitHub

BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux

Beschreibung


GPU: 2x NVIDIA GeForce RTX 2080 Ti 22GB (compute capability 7.5, Turing / SM75)
RAM: 63 GB
CPU: Intel Xeon E5-2696 v3 @ 2.30GHz (AVX2)
Driver: 580.178.04
CUDA: 13.0
OS: Linux (Ubuntu-based)
Strata version: v0.1.38 (commit 99f3dbd)
Model: Qwen3.8-Flash-Next IQ3_S
Context: 262144 (256K), KV int8

Running Strata with a dual-GPU layer split (--gpu 0,1) on IQ3_S (Qwen3.8-Flash-Next). 
The engine starts successfully, but crashes during inference with:
  verify: layer X never rang (an illegal memory access was encountered)

The layer number varies between runs (10, 14, 20, 23, 24), suggesting a multi-GPU 
synchronization/communication issue rather than a layer-specific bug.

With CUDA_LAUNCH_BLOCKING=1, a long prompt (133,853 tokens) hangs at 
"reading the prompt (verify windows), from token 133848" and the engine watchdog 
kills itself after 60s of no progress (exit code -6).

IQ2_XS runs stably on the same hardware (57-71 tok/s, zero crashes over thousands of 
tokens), suggesting the issue is specific to the IQ3_S inference path or activation 
sizes that trigger a memory access pattern not present in IQ2_XS.



1. ./setup.sh --setup, select Qwen3.8-Flash-Next → IQ3_S → 256K context → KV int8 → dual GPU
2. Start with: python serve/server.py --engine strata --config strata-iq3_s.json --port 8080
3. Send a long prompt (100K+ tokens) or run multiple chat requests
4. Engine crashes with "layer X never rang (illegal memory access)"






[strata] reading the prompt: 133,848 of 133,853 tokens, 138 s so far
[strata] reading the prompt: 133,848 of 133,853 tokens, 148 s so far
[strata] reading the prompt: 133,848 of 133,853 tokens, 158 s so far
[strata] reading the prompt: 133,848 of 133,853 tokens, 168 s so far
[strata] reading the prompt: 133,848 of 133,853 tokens, 178 s so far
[strata] reading the prompt: 133,848 of 133,853 tokens, 188 s so far
[strata] reading the prompt: 133,848 of 133,853 tokens, 198 s so far
[strata] the engine stopped unexpectedly (exit code -6). The engine stopped itself because 
it had stopped making progress - a hang it caught. Its log line: strata serve: no progress 
for 60 s during a request (reading the prompt (verify windows), from token 133848) - 
stopping the engine so the server starts it again (issue #29) - please report it at 
github.com/Niko1221/Strata/issues.
[strata] done: 0 tokens in 205 s (0.0 tok/s) (error, cancel=False)





[strata] the engine reported an error: verify: layer 24 never rang (an illegal memory access was encountered)
[strata] done: 3087 tokens in 56 s (56.5 tok/s) (error, cancel=False)
[strata] the engine had stopped (exit code 1); starting it again (a minute or two) ...



CUDA_ERROR_OUT_OF_MEMORY (error 2) due to "out of memory" on CUDA API call to cuLibraryLoadData.
Host Frame: strata::prefill::Gemm::init_external
Host Frame: strata::prefill::Prefill::init
cublasLtCtxInit → cublasCreate_v2 → Gemm::init_external



[strata] layer split across GPUs [0, 1] (auto)
[strata] experts loaded: 46.84 GiB at 1.33 GiB/s (60 s so far)
[strata] filling the GPU's expert cache (8231 experts, 14.82 GiB of VRAM) ...
[strata] layer split: layers 0-23 (CUDA0), 24-47 (CUDA1), one hand-off per window
[strata] verify: window up to 6 tokens, 75.4 MiB of device buffers
[strata] serve: mtp: the draft head does not fit




Mehr auf der Site

Links zu Install, Modellen, Releases.