Pull requests / #1240

#1240 Avoid unused MTP branch stream on HIP

closed · @rkcth · 0 comments · View on GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quants

Description

On HIP, fresh long prompts with MTP can lose prompt-processing speed after the first request. `MtpDrafter::load` creates a branch stream and two events even though the shared-expert branch is disabled on HIP. Guard those unused resources for CUDA, leaving the CUDA branch path unchanged.

On a Radeon AI PRO R9700 with IQ3_S, v0.1.40.1 processed a fresh ~30K-token prompt at 1,161 tok/s, then the next at 842 tok/s. With this fix, five successive fresh ~30K/~33K prompts measured 1,155, 1,188, 1,190, 1,142, and 1,139 tok/s. All generated the requested 301–305 output tokens. Two concurrent 2K/128-token requests also completed, and the HIP build passed.

The regression bisects to `4a5713f`, which introduced the branch resource allocation. The experiment guarded the stream and events together; it did not isolate which allocation causes the HIP runtime effect. NVIDIA hardware has not been tested with this patch.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.