Issues / #1304

#1304 Prefill/Reading the prompt with multi-gpu cards is slow

open · @sense1024 · 2 コメント · GitHub で見る

Multi-GPUModels & quants

本文

I have 2x 2080TI 22GB with NVLink. 
I use them to run llama.cpp with Qwen 3.8 27B Q6. The prefill speed is between 2300 and 700 t/s. 

Now, I set up Strata with 3.8 Flash Next IQ3_S. The prefill speed is always less than 400 t/s. 
Why and how to improve prefill speed?

I know that the current version of Strata does not support P2P/NCCL communication.
Do I need to remove NVLink to run Strata?

Thanks.

関連リンク

インストール・モデル・リリースへの站内リンク。