Issues / #964
#964 SM75 (2080Ti) dual-GPU: "layer X never rang (illegal memory access)" and verify-window hang on IQ3_S
open · @dehuiyede · 2 comments · View on GitHub
BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux
Description
GPU: 2x NVIDIA GeForce RTX 2080 Ti 22GB (compute capability 7.5, Turing / SM75) RAM: 63 GB CPU: Intel Xeon E5-2696 v3 @ 2.30GHz (AVX2) Driver: 580.178.04 CUDA: 13.0 OS: Linux (Ubuntu-based) Strata version: v0.1.38 (commit 99f3dbd) Model: Qwen3.8-Flash-Next IQ3_S Context: 262144 (256K), KV int8 Running Strata with a dual-GPU layer split (--gpu 0,1) on IQ3_S (Qwen3.8-Flash-Next). The engine starts successfully, but crashes during inference with: verify: layer X never rang (an illegal memory access was encountered) The layer number varies between runs (10, 14, 20, 23, 24), suggesting a multi-GPU synchronization/communication issue rather than a layer-specific bug. With CUDA_LAUNCH_BLOCKING=1, a long prompt (133,853 tokens) hangs at "reading the prompt (verify windows), from token 133848" and the engine watchdog kills itself after 60s of no progress (exit code -6). IQ2_XS runs stably on the same hardware (57-71 tok/s, zero crashes over thousands of tokens), suggesting the issue is specific to the IQ3_S inference path or activation sizes that trigger a memory access pattern not present in IQ2_XS. 1. ./setup.sh --setup, select Qwen3.8-Flash-Next → IQ3_S → 256K context → KV int8 → dual GPU 2. Start with: python serve/server.py --engine strata --config strata-iq3_s.json --port 8080 3. Send a long prompt (100K+ tokens) or run multiple chat requests 4. Engine crashes with "layer X never rang (illegal memory access)" [strata] reading the prompt: 133,848 of 133,853 tokens, 138 s so far [strata] reading the prompt: 133,848 of 133,853 tokens, 148 s so far [strata] reading the prompt: 133,848 of 133,853 tokens, 158 s so far [strata] reading the prompt: 133,848 of 133,853 tokens, 168 s so far [strata] reading the prompt: 133,848 of 133,853 tokens, 178 s so far [strata] reading the prompt: 133,848 of 133,853 tokens, 188 s so far [strata] reading the prompt: 133,848 of 133,853 tokens, 198 s so far [strata] the engine stopped unexpectedly (exit code -6). The engine stopped itself because it had stopped making progress - a hang it caught. Its log line: strata serve: no progress for 60 s during a request (reading the prompt (verify windows), from token 133848) - stopping the engine so the server starts it again (issue #29) - please report it at github.com/Niko1221/Strata/issues. [strata] done: 0 tokens in 205 s (0.0 tok/s) (error, cancel=False) [strata] the engine reported an error: verify: layer 24 never rang (an illegal memory access was encountered) [strata] done: 3087 tokens in 56 s (56.5 tok/s) (error, cancel=False) [strata] the engine had stopped (exit code 1); starting it again (a minute or two) ... CUDA_ERROR_OUT_OF_MEMORY (error 2) due to "out of memory" on CUDA API call to cuLibraryLoadData. Host Frame: strata::prefill::Gemm::init_external Host Frame: strata::prefill::Prefill::init cublasLtCtxInit → cublasCreate_v2 → Gemm::init_external [strata] layer split across GPUs [0, 1] (auto) [strata] experts loaded: 46.84 GiB at 1.33 GiB/s (60 s so far) [strata] filling the GPU's expert cache (8231 experts, 14.82 GiB of VRAM) ... [strata] layer split: layers 0-23 (CUDA0), 24-47 (CUDA1), one hand-off per window [strata] verify: window up to 6 tokens, 75.4 MiB of device buffers [strata] serve: mtp: the draft head does not fit
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.