Pull requests / #1185

#1185 verify: a window graph that finds no VRAM frees older batch slot layouts' graphs instead of failing (#997)

closed · @ischencheng · 0 评论 · 在 GitHub 查看

Setup & installAMD / HIPNVIDIA / CUDAWindowsLinux

描述

With `"parallel": N`, every set of slots that decodes together gets its own window and commit graph pair (`exec_bm_` / `commit_bm_`, keyed by the row layout), and they stay for the life of the engine unless `--batch-mtp` sets its limit of 64. On an L40S one pair takes 20-30 MiB (more rows, more MiB), and 8 slots have 255 possible sets. The expert cache sizes itself to leave only `--vram-reserve-mib` free, so after a dozen or a few dozen layouts the next `cudaGraphInstantiate` fails and the engine exits with `verify: batch instantiate: out of memory`. That's what Jackwwg83 found in #997 (379 MiB free after load, fails when 7-8 slots decode together, `--vram-reserve-mib 1700` avoids it). The one-request windows captured on first use (2-5 tokens) hit the same wall as `verify: instantiate: out of memory`.

Now an instantiate that runs out of VRAM frees the least recently used batch layouts' graphs (never the layout being captured) and tries again. A freed layout is captured again when it comes back. While everything fits nothing is freed, so the default path doesn't change. The `--batch-mtp` limit goes through the same eviction helper now. I also fixed the batch commit graph's message, which said `verify: batch commit capture: no error` when its instantiate failed.

Measured on an L40S (46068 MiB, sm_89, Linux), IQ2_XS, setup's config plus `"parallel": 8`, 12 clients sending random short prompts with random `max_tokens` for 7 minutes. For the low-free-VRAM case in the issue, another process held 12000 MiB before the engine started, so the cache sized itself to ~440 MiB free:

| | free after load | requests ok | engine restarts | batch captures |
|---|---|---|---|---|
| main, 12000 MiB held | 440 MiB | 45 / 142 | 8 | 102 |
| this PR, 12000 MiB held | 442 MiB | 623 / 623 | 0 | 152 (140 evictions) |
| main, nothing held | 1334 MiB | 762 / 762 | 0 | 47, 184 MiB left at the end |
| this PR, nothing held | 1334 MiB | 786 / 786 | 0 | 47, no eviction |

So the plain 48 GB setup was a few more layouts away from the same crash. Three fixed waves of 8 concurrent requests at temperature 0 gave the same 24 replies byte for byte on main (two runs) and on this PR. The per-layout MiB come from a debug-only log line (free VRAM after each capture) that isn't in this PR.

Not tested: Windows, HIP (`cudaErrorMemoryAllocation` maps to `hipErrorOutOfMemory` there), `--batch` on a layer split, and `--batch-mtp` (its eviction only moved into the helper). It also doesn't touch #1102, which applies cleanly on top.

Fixes the batch instantiate crash from Jackwwg83's comment in #997. The original report there (RTX 3090, 360 MiB free) may be the same thing, but I can't tell without its ERR line.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。