Pull requests / #1741
#1741 verify: service returns before any CUDA call while nothing has rung (#1740)
open · @architectds · 0 Kommentare · Auf GitHub
BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
Beschreibung
## Title Issue: #1740 (the stall this addresses; one stall in ~50 requests cannot prove it gone quickly) ## Summary The pipelined decode loop asks both stages' windows for service on every pass. While a window waits on its doorbell, `Verifier::service` still switched the device - `OnDevice`: `cudaGetDevice` and two `cudaSetDevice` - before finding nothing to do: a few driver calls a pass, about a billion passes in a long session. In #1740 the serving thread blocked inside one of those `cudaSetDevice` calls while both stages' windows waited on it, and Strata was down until the engine restarted. Now `service` returns 0 before any CUDA call when nothing has rung and neither the WDDM flush (2 ms) nor the timeout (20 s) is due - the answer the loop below gives in that case - so the waiting loop calls the driver only to serve a layer, to flush, or to time out. Measured on the PC of the issue (RTX 3060 + RTX 5070 Ti, layers 0-11 / 12-47, IQ3_S, `--pipeline-windows 2`, `STRATA_PREFILL_CPU_SHARE=auto`): a coding agent's recorded conversation (a 100K-token start, then 8 turns of 1-6K tokens of code, 128 tokens written a turn), A B A B in one window: 130.7 / 128.6 s before, 128.0 / 128.3 s with this; writing 69.2 / 70.5 against 70.8 / 69.7 tok/s. With `STRATA_PREFILL_CPU_SHARE=auto` the answers are not repeatable from run to run (the share is balanced by measured times), before and after alike; all four runs' answers read normally. ## What changed - `src/core/verify.cpp`, `Verifier::service`: the early return above, before the `OnDevice` guard. ## Extra Notes - Nothing else changes: when something has rung, or a flush or the timeout is due, the code below runs as before. - Not proven: the stall came once in 51 requests here. If it comes back, the flush's `cudaEventQuery` / `cudaStreamQuery` every 2 ms is the driver call left in the waiting loop. - Tested: Windows 11, CUDA 13.0, sm_86 + sm_120. Not tested: Linux, HIP. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Mehr auf der Site
Links zu Install, Modellen, Releases.