Issues / #1304

#1304 Prefill/Reading the prompt with multi-gpu cards is slow

open · @sense1024 · 2 评论 · 在 GitHub 查看

Multi-GPUModels & quants

描述

I have 2x 2080TI 22GB with NVLink. 
I use them to run llama.cpp with Qwen 3.8 27B Q6. The prefill speed is between 2300 and 700 t/s. 

Now, I set up Strata with 3.8 Flash Next IQ3_S. The prefill speed is always less than 400 t/s. 
Why and how to improve prefill speed?

I know that the current version of Strata does not support P2P/NCCL communication.
Do I need to remove NVLink to run Strata?

Thanks.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。