Pull requests / #1240
#1240 Avoid unused MTP branch stream on HIP
closed · @rkcth · 0 comentarios · En GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quants
Descripción
On HIP, fresh long prompts with MTP can lose prompt-processing speed after the first request. `MtpDrafter::load` creates a branch stream and two events even though the shared-expert branch is disabled on HIP. Guard those unused resources for CUDA, leaving the CUDA branch path unchanged. On a Radeon AI PRO R9700 with IQ3_S, v0.1.40.1 processed a fresh ~30K-token prompt at 1,161 tok/s, then the next at 842 tok/s. With this fix, five successive fresh ~30K/~33K prompts measured 1,155, 1,188, 1,190, 1,142, and 1,139 tok/s. All generated the requested 301–305 output tokens. Two concurrent 2K/128-token requests also completed, and the HIP build passed. The regression bisects to `4a5713f`, which introduced the branch resource allocation. The experiment guarded the stream and events together; it did not isolate which allocation causes the HIP runtime effect. NVIDIA hardware has not been tested with this patch.
En el sitio
Enlaces a install, modelos, releases.