Pull requests / #1741

#1741 verify: service returns before any CUDA call while nothing has rung (#1740)

open · @architectds · 0 commentaires · Sur GitHub

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Description

## Title
Issue: #1740 (the stall this addresses; one stall in ~50 requests cannot prove it gone quickly)

## Summary
The pipelined decode loop asks both stages' windows for service on every pass. While a window waits on its doorbell,
`Verifier::service` still switched the device - `OnDevice`: `cudaGetDevice` and two `cudaSetDevice` - before finding
nothing to do: a few driver calls a pass, about a billion passes in a long session. In #1740 the serving thread
blocked inside one of those `cudaSetDevice` calls while both stages' windows waited on it, and Strata was down until the
engine restarted.

Now `service` returns 0 before any CUDA call when nothing has rung and neither the WDDM flush (2 ms) nor the timeout
(20 s) is due - the answer the loop below gives in that case - so the waiting loop calls the driver only to serve a
layer, to flush, or to time out.

Measured on the PC of the issue (RTX 3060 + RTX 5070 Ti, layers 0-11 / 12-47, IQ3_S, `--pipeline-windows 2`,
`STRATA_PREFILL_CPU_SHARE=auto`): a coding agent's recorded conversation (a 100K-token start, then 8 turns of 1-6K tokens
of code, 128 tokens written a turn), A B A B in one window: 130.7 / 128.6 s before, 128.0 / 128.3 s with this; writing
69.2 / 70.5 against 70.8 / 69.7 tok/s. With `STRATA_PREFILL_CPU_SHARE=auto` the answers are not repeatable from run to
run (the share is balanced by measured times), before and after alike; all four runs' answers read normally.

## What changed
- `src/core/verify.cpp`, `Verifier::service`: the early return above, before the `OnDevice` guard.

## Extra Notes
- Nothing else changes: when something has rung, or a flush or the timeout is due, the code below runs as before.
- Not proven: the stall came once in 51 requests here. If it comes back, the flush's `cudaEventQuery` /
  `cudaStreamQuery` every 2 ms is the driver call left in the waiting loop.
- Tested: Windows 11, CUDA 13.0, sm_86 + sm_120. Not tested: Linux, HIP.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Sur le site

Liens install, modèles, releases.