Pull requests / #1383

#1383 generate: say that CUDA_LAUNCH_BLOCKING=1 hangs the verify windows (#1341 #964)

open · @ischencheng · 0 comentários · No GitHub

Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Descrição

The 2/2 hang reproduced in #1341 and the hang in #964 were both run with `CUDA_LAUNCH_BLOCKING=1`, set to find a faulting kernel. That variable alone hangs the engine. A verify window's graph waits on the GPU for flags the host raises after `cudaGraphLaunch` returns, and with blocking launches `cudaGraphLaunch` returns only when the graph is done. So the first window after any prompt waits forever and the #29 watchdog ends the engine, whatever the prompt length, GPU count or `--pipeline-windows`.

This prints a warning at start when the variable is set, and the watchdog's line names it. Without the variable nothing changes. The server console still shows the `#267` ERR line that follows the release, so the hint is in the engine log, which is what both reports quoted.

Modal L40S 46 GB, Linux, driver 610.57, sm_89 portable build, IQ2_XS, upstream e8ca9af, a fresh server per row:

| setup | CUDA_LAUNCH_BLOCKING | prompt tokens | result |
|---|---|---|---|
| 1 GPU | 1 | 2,428 | hangs at `reading the prompt (verify windows), from token 2420` |
| 2 GPUs, `--pipeline-windows 0` | 1 | 2,428 | hangs at the same stage (token 2421) |
| 2 GPUs, `--pipeline-windows 2` | 1 | 2,428 | hangs at `(pipelined verify windows), from token 2421` |
| 2 GPUs, `--pipeline-windows 2`, ~12 GB expert cache per card | unset | 120,790 | read in 16.6 s |
| 2 GPUs, `--pipeline-windows 0`, ~12 GB expert cache per card | unset | 120,790 | read in 16.8 s |
| 2 GPUs, `--pipeline-windows 2`, all experts in VRAM | unset | 120,790 | read in 17.3 s |

The 2-GPU rows use `"layer_split": "26"` with `--max-context 262144 --kv int8 --kv-resident 32768 --prefill auto:32768 --trim-stage-weights`; the 12 GB rows add `--vram-reserve-mib 30000 --vram-reserve-later-mib 30000`.

With `STRATA_VERIFY_TRACE=1` the stall report shows why: the host's last event is `WINDOW` with no `LAUNCHED` after it (it is still inside `cudaGraphLaunch`), and the GPU breadcrumbs stop at `layer 1 group 0: pre(0)`, waiting for the host.

Before/after on 1 GPU with the variable set, main logs nothing about it. This branch logs:

```
warning: CUDA_LAUNCH_BLOCKING=1: the verify windows cannot run with blocking launches, so the first one after a prompt hangs until the watchdog ends the engine (#1341). Unset it; STRATA_PF_STEP_SYNC=1 narrows down a failing prompt step without it
strata serve: no progress for 45 s during a request (reading the prompt (verify windows), from token 2391) - stopping the engine so the server starts it again (issue #29); CUDA_LAUNCH_BLOCKING=1 is set, and no verify window can finish with it (#1341)
```

Without the variable, three fixed prompts at temperature 0 gave identical replies on main and this branch, and the engine logs match apart from the binary's path.

Not checked: Windows (the #1341 machine) and HIP. The check is skipped on HIP builds; `HIP_LAUNCH_BLOCKING` or `AMD_SERIALIZE_KERNEL` may do the same there, I have not tried. This does not explain the #1341 failures that ran without the variable (the 129k request that ended with an error, and the 107k `prefill copy_i32` illegal access in the batched reader). I could not get a hang without it on two L40S.

Related to #1341, #964.

No site

Links install, modelos, releases.