反馈 / #1740
#1740 [Bug]: Two-GPU pipelined decode (--pipeline-windows 2) stalls on Windows: the serving thread blocks in cudaSetDevice while both stages wait on it
open · @architectds · 0 评论 · 去 GitHub 看
BenchmarksServer & APINVIDIA / CUDAModels & quantsWindows
说明
### What happened Two-GPU pipelined decode (`--pipeline-windows 2`) stopped in the middle of a long agent request and never resumed: no token for 60 s, then the server's stall report, the released GPU waits (#267) and an engine restart (#29). It is rare here - once in 51 requests, about 12 hours of agent use - but each hit is an outage: the request fails, the parked conversations are gone, and the client's retry has to read its whole prompt again (398K tokens, about 4 minutes on this PC), so Strata is unusable for about 5 minutes. **The state at the stall** (the report below, position 398,917 of a 512K context): both pipeline windows in flight, and each stage's GPU has rung for a layer the host has not served. - stage 0 (RTX 3060, layers 0-11), even window: `served 3/12; GPU rang 4` - stage 1 (RTX 5070 Ti, layers 12-47), odd window: `served 11/36; GPU rang 12` - the expert pool done with its 48 jobs, its 15 workers asleep, the host idle 61 s - the loop counter at 919,907,381 in both reports, 2 s apart: the serving thread was blocked, not spinning - nvidia-smi meanwhile: RTX 3060 100% at 51 W (P2, the spin kernel), RTX 5070 Ti 0% at P8 **Where the host was.** The stall dump's thread stacks, symbolized with the build's PDB (dbghelp): the serving thread sat in the CUDA runtime under `Verifier::service`, in the `OnDevice` guard's `cudaSetDevice` at the end of `service` (verify.cpp:3084 in v0.1.41), called from the loop's stage-1 `service` (generate.cpp:10308). No other engine thread was in CUDA: the 15 pool workers parked, the 8 stagers idle, the PLE reader, the stdin reader and the watchdog. So the GPUs wait for the host's flags while the host waits inside a driver call - it looks like the (WDDM) driver waiting for work that cannot move on. Every `service` call switches the device (`cudaGetDevice` and two `cudaSetDevice`) even with nothing to serve, and the loop calls it for both stages on every pass: about a billion passes in that session. **Before:** the same shape once on 2026-10-06 in a test run of a fork build of 0.1.39 with Hardin22's pipelined windows: position 102,464, stage 0's odd window `served 2/12; GPU rang 3`, the loop counter at 103,183,883 in both reports, host idle 61 s. So it does not need a long context. **A proposed fix** is in the PR linked below: `service` returns before any CUDA call while nothing has rung, which takes the driver calls out of the waiting loop (a WDDM flush every 2 ms stays). It cannot prove itself quickly at one stall in ~50 requests; we run it from today and will report here. ### Strata version, GPU, OS 0.1.41 (stock fb58e0d, built locally for sm_86 + sm_120, CUDA 13.0) · RTX 3060 12 GB (stage 0, also the display) + RTX 5070 Ti 16 GB (stage 1), both PCIe 3.0 x8 · Windows 11 · Ryzen 9 5900XT, 96 GB DDR4-3200 Config: `"gpu": [1, 0]`, `"layer_split": "12"`, IQ3_S (GSQ-RCO), `STRATA_PREFILL_CPU_SHARE=auto`, `--expert-cache auto --prefill auto:16384 --spec 4 --spec-min-p 0.70 --mtp <rt> --max-context 512000 --kv q4_0 --kv-resident 32768 --vision --vram-reserve-mib 1200 --vram-reserve-later-mib 500 --rope-scaling yarn --rope-scale 1.953125 --pcie-frac 0.35 --mmap-experts --conversation-cache-mib 8192 --conversation-cache-slots 4 --pipeline-windows 2 --trim-stage-weights` ### Engine log ``` strata serve: prompt 397822 tokens = 397793 reused + 29 read in 799 ms (36.3 tok/s), 1343 generated in 23822 ms (56.4 tok/s), drafts accepted 649 of 1009, 6 checkpoints strata serve: conversation cache: parked 399164 tokens in 2459.3 ms; parked=1 bytes=4450367480 evictions=5 snapshot_bytes=4450367480 reused_kv_bytes=0 strata serve: no progress for 60 s during a request (verify window (pipelined): the CPU experts of layer 22) - stopping the engine so the server starts it again (issue #29) strata serve: stall report (engine 0.1.41): stage "verify window (pipelined): the CPU experts of layer 22" for 61 s; 0 layers served since the last finished step (0 = stopped, more = slow) expert pool: epoch 1339538, batch epoch 1339538: 48 of 48 jobs claimed, 48 done; 15 of 15 workers parked, 15 sleeping; mode 0 expert pool threads: w0..w14=sleeping; host idle for 61301 ms verify window: 4 tokens at position 398917, host at layer step 3; the GPU rang 4; flags: served 3, plan (A) 3, copies (B) 3 pipelined loop: 919907381 iterations; A seq 533 [RLFCS-] B [RL----] D [RLF---] (Ready Launched Finished Committed S1-launched S1-done); chain kind 0 live 0; doomed 0 ending 0 stage 0 even: IN FLIGHT T=4 pos 398917 served 3/12; GPU rang 4, flags served 3 A 3 B 3; window event PENDING, commit event done (commit launched 1) stage 0 odd: idle T=4 pos 398913 served 12/12; GPU rang 12, flags served 12 A 12 B 12; window event done, commit event done (commit launched 1) stage 1 even: idle T=2 pos 398912 served 36/36; GPU rang 36, flags served 36 A 36 B 36; window event done, commit event done (commit launched 1) stage 1 odd: IN FLIGHT T=4 pos 398913 served 11/36; GPU rang 12, flags served 11 A 11 B 11; window event PENDING, commit event done (commit launched 1) memory: 54226 MiB resident, 39264 MiB committed, 18271 MiB RAM available; 56129065 page faults so far 2 s later: (the same, the loop still at 919907381 iterations, host idle 63316 ms) wrote the thread stacks to D:\Strata\strata-stall-50252.dmp (attach it to the issue) strata: released the verify window's GPU waits (#267): the GPU did not finish within 5 s ``` The dump (373 KB) is available on request; it may hold prompt text on the stacks, so it is not attached.
本站相关内容
相关页面的快捷入口。