Pull requests / #947

#947 fix: avoid CPU doorbell waits in fully resident batch decoding

closed · @CC-David-CC · 0 评论 · 在 GitHub 查看

Multi-GPUNVIDIA / CUDADocumentation

描述

## Draft: small all-resident batch completion fix

Fully GPU-resident batch graphs do not emit CPU expert-work doorbells, but the batch host paths still wait for them. Batch PLE rows are already gathered before launch.

Set the PLE-ready flag for this path and skip the absent CPU doorbell steps in run_slot_rows and batch_launch. Partially resident paths retain their previous flag/step values. No new feature, policy or preset.

Code delta: one file, +5/-3 lines, based on current upstream 6f32ec0. A second commit adds documentation and a compact test receipt.

Completed integration screen on RTX PRO 6000 Blackwell 96 GB: five models, 15 requests; each 8K input + 512 output, FP16 KV. All five two-request target-only batches completed. Testing used integration commit 66b9474, not a fresh standalone build. No speedup claim.

Before ready for review: isolated build/reproducer, partially resident control, GPU sanitizer and cancellation/recovery checks, and separate pipeline-path coverage. Multi-GPU execution remains untested.

[Details and receipt](https://github.com/CC-David-CC/Strata-a5500/blob/fix/all-resident-batch/docs/ALL_RESIDENT_BATCH_FIX.md)

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。