Pull requests / #1201

#1201 verify batch: service the doorbell graph per-window (ar_on), not per-…

closed · @wOvAN · 0 comentarios · En GitHub

NVIDIA / CUDA

Descripción

A stage that was 100% resident at init hangs once a prompt loan/shrink/swap puts it on the doorbell graph: run_slot_rows, batch_poll and service() key the all-resident service skip off the init-time all_resident_ flag instead of the per-window ar_on(). The GPU then waits on host doorbells nobody raises (one GPU at 100%, pool idle) while the host sits in cudaStreamSynchronize, until the watchdog fires with the previous stage's stale label (the CPU experts of layer 7 — always stage 0's last serviced layer, 0 ticks).

Evidence: host backtrace of a stuck engine shows the main thread in cuStreamSynchronize called directly from run_slot_rows (post-loop sync, verified by disassembly); pool trace shows layers of stages 1+ never serviced; nvidia-smi shows the loan stage's GPU spinning while the others are idle. The token path run() already uses ar_on() — this extends the same rule to the three batch paths.

Fix: 4 one-word changes, all_resident_ → ar_on(). Verified with a loan repro (parallel 80k-token prefills in borrow mode): 11 watchdog kills before, 0 after, no content regression.

En el sitio

Enlaces a install, modelos, releases.