Pull requests / #1240

#1240 Avoid unused MTP branch stream on HIP

closed · @rkcth · 0 评论 · 在 GitHub 查看

BenchmarksAMD / HIPNVIDIA / CUDAModels & quants

描述

On HIP, fresh long prompts with MTP can lose prompt-processing speed after the first request. `MtpDrafter::load` creates a branch stream and two events even though the shared-expert branch is disabled on HIP. Guard those unused resources for CUDA, leaving the CUDA branch path unchanged.

On a Radeon AI PRO R9700 with IQ3_S, v0.1.40.1 processed a fresh ~30K-token prompt at 1,161 tok/s, then the next at 842 tok/s. With this fix, five successive fresh ~30K/~33K prompts measured 1,155, 1,188, 1,190, 1,142, and 1,139 tok/s. All generated the requested 301–305 output tokens. Two concurrent 2K/128-token requests also completed, and the HIP build passed.

The regression bisects to `4a5713f`, which introduced the branch resource allocation. The experiment guarded the stream and events together; it did not isolate which allocation causes the HIP runtime effect. NVIDIA hardware has not been tested with this patch.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。