Issues / #1643

#1643 [Bug] v0.1.41 no-P2P peer prefill crashes on long prompts (RTX 2080 Ti + RTX 5050, CUDA Xid 31)

open · @GarminChen · 0 comentarios · En GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsLinux

Descripción

<html>
<body>
<!--StartFragment--><html><head></head><body><h2><span>Summary</span></h2><p><span>Strata v0.1.41 successfully initializes the new no-P2P peer expert tier on my heterogeneous NVIDIA GPU pair (RTX 2080 Ti 22GB + RTX 5050 8GB).</span></p><p><span>However, </span><strong><span>a long prompt of 31,358 tokens consistently crashes the engine when peer prefill is enabled</span></strong><span>.</span></p><p><span>Short prompts work. The same 31K prompt also works with </span><code><span>--peer-prefill-rows 0</span></code><span>, while keeping the peer tier enabled for decode.</span></p><p><span>Disabling </span><code><span>STRATA_BF16_TC</span></code><span> and </span><code><span>STRATA_PF_PEER_XSTREAM</span></code><span> did not fix the long-prefill crash.</span></p><h2><span>System</span></h2><ul><li><p><span>Strata: v0.1.41, locally compiled, SM75 + SM120</span></p></li><li><p><span>OS: Linux x86_64 (NAS)</span></p></li><li><p><span>CPU: Intel Core i3-8100, AVX2</span></p></li><li><p><span>RAM: 48 GB physical (46 GiB reported by Linux)</span></p></li><li><p><span>GPU0: NVIDIA RTX 2080 Ti 22GB (modified VRAM, SM75)</span></p></li><li><p><span>GPU1: NVIDIA RTX 5050 8GB (SM120)</span></p></li><li><p><span>NVIDIA driver: 580.142</span></p></li><li><p><span>CUDA Toolkit: 13.0.88</span></p></li><li><p><span>GPU communication: No P2P; Strata reports the mapped-host-memory peer prefill path</span></p></li></ul><h2><span>Model and configuration</span></h2><ul><li><p><span>Model: Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS</span></p></li><li><p><span>Context: 131072 tokens</span></p></li><li><p><span>KV cache: INT8</span></p></li><li><p><span>MTP: </span><code><span>--spec 4 --spec-min-p 0.5</span></code></p></li><li><p><span>Expert arena: </span><code><span>STRATA_ARENA_MMAP=1</span></code></p></li><li><p><code><span>--pcie-frac 0</span></code></p></li><li><p><code><span>--expert-profile</span></code><span> enabled</span></p></li><li><p><code><span>--expert-cache auto</span></code></p></li><li><p><code><span>--prefill auto</span></code></p></li><li><p><code><span>--peer-device 1</span></code></p></li><li><p><code><span>--peer-reserve-mib 1024</span></code></p></li><li><p><code><span>STRATA_PREFILL_CPU_SHARE=0</span></code></p></li><li><p><span>CPU affinity: cores 0, 1, 3</span></p></li></ul><p><span>The old remote helper flags (</span><code><span>--expert-cache-device1</span></code><span> and </span><code><span>--remote-expert-opt</span></code><span>) are not used with the peer configuration.</span></p><p><span>The peer tier initializes with approximately 4,765 experts / 6.39 GiB on GPU1. The primary GPU holds approximately 10,709 expert slots.</span></p><h2><span>Steps to reproduce</span></h2><ol start="1"><li><p><span>Launch Strata v0.1.41 with the configuration above, enabling no-P2P peer prefill (default </span><code><span>--peer-prefill-rows -1</span></code><span>).</span></p></li><li><p><span>Send a short prompt of about 55 tokens. It works and can generate over 1,000 tokens.</span></p></li><li><p><span>Send a fixed synthetic archive prompt of 31,358 tokens with </span><code><span>max_tokens=256</span></code><span>, using </span><code><span>/v1/chat/completions</span></code><span>.</span></p></li><li><p><span>The engine exits with code 1. The server responds with </span><code><span>server_error</span></code><span> and attempts to restart the engine.</span></p></li></ol><p><span>The failing long prompt has </span><code><span>cached_tokens=0</span></code><span>.</span></p><h2><span>Actual behavior</span></h2><p><span>Engine log:</span></p><pre><code><span>strata prefill: peer GPU 1 computes its experts' rows of each prompt chunk
(up to 81920 rows per layer, no P2P: through mapped host memory in FP16,
one weighted sum per token back, compact group buffers);
321 MiB left free on it

strata serve: 493 MiB of VRAM free with everything loaded

prefill gemm: cublasGemmEx f16: cuBLAS status 13</span></code></pre><p><span>NVIDIA kernel log, GPU0 (RTX 2080 Ti):</span></p><pre><code><span>NVRM: Xid (PCI:0000:01:00): 31, pid=2491110, name=strata
MMU Fault: ENGINE GRAPHICS GPC5 GPCCLIENT_T1_2
faulted @ 0x7f46_31ded000
FAULT_PDE ACCESS_TYPE_VIRT_READ</span></code></pre><p><span>A separate NVIDIA driver log also reported </span><code><span>NV_ERR_NO_MEMORY</span></code><span>, but I have not established that it belongs to the same failure.</span></p><h2><span>Isolation results</span></h2>
Configuration | Short prompt | 31,358-token prompt
-- | -- | --
Peer prefill enabled, STRATA_BF16_TC=1 | Not established | Crash
Peer prefill enabled, BF16_TC disabled | Pass | Crash
Peer prefill enabled, BF16_TC disabled, STRATA_PF_PEER_XSTREAM=0 | Pass | Crash
--peer-prefill-rows 0, BF16_TC enabled | Not separately tested | Pass
Old remote helper 5000 architecture, v0.1.41 | Not separately tested | Pass

<p><span>Successful peer-tier run with prefill disabled:</span></p><ul><li><p><span>Cold prompt: 31,358 tokens</span></p></li><li><p><span>Prefill: 1,070.6 tokens/s</span></p></li><li><p><span>Decode: 55.8 tokens/s</span></p></li><li><p><span>Drafts accepted: 182/221</span></p></li><li><p><span>Decode expert cache hit rate: 91.1%</span></p></li><li><p><span>No engine crash</span></p></li></ul><p><span>These are single-run performance observations, not repeated benchmark medians.</span></p><h2><span>Expected behavior</span></h2><p><span>The no-P2P peer prefill path should complete the 31K-token prompt without CUDA memory faults or engine termination.</span></p><h2><span>Additional observations</span></h2><ul><li><p><span>The host system has approximately 34 GiB of available RAM after initialization.</span></p></li><li><p><code><span>STRATA_ARENA_MMAP=1</span></code><span> successfully maps the expert arena without locking it.</span></p></li><li><p><span>The failure is associated with enabling peer prefill for long prompts, not with peer-tier initialization.</span></p></li><li><p><span>Disabling the BF16 tensor-core optimization does not resolve the failure.</span></p></li><li><p><span>Disabling the independent peer transfer stream does not resolve the failure.</span></p></li><li><p><span>I have stopped further long-prefill tests to avoid repeatedly triggering GPU faults.</span></p></li></ul><p><span>Could you investigate whether the no-P2P peer prefill path has a buffer allocation, memory lifetime, or synchronization issue with this SM75 + SM120 GPU combination?</span></p><p><span>I can provide the complete engine logs, configuration, and synthetic benchmark payload if useful.</span></p><p><span>Thanks for the work on Strata! The remote helper configuration has been performing very well on this system, and I am excited to test the new peer prefill support.</span></p></body></html><!--EndFragment-->
</body>
</html>

En el sitio

Enlaces a install, modelos, releases.