Issues / #1403

#1403 Almost no changes from 2 recent updates,

open · @Noctorngaming · 0 评论 · 在 GitHub 查看

Multi-GPUModels & quantsLinux

描述

I was using 0.1.39 for a while, and was getting near 20-35tps decode, and around 150 pp, I do not know if its because the 2nd gpu is on pcie 3.0 x4 or not.

Hardware -
4070 12gb - running on 4.0 x16
v100 16gb - running on 3.0 x4
32gb 3200mhz ram

Model + para-
rco iq3_xxs, q8 kv, 128k context, with vision.

os-
ubuntu 24.04

speed -
25-30~~ tps on decode
150 ~ tps on prefill

I have seen many speedup in releases but I have not seen much speed difference, other than my speed going from 30tps top's to around 35tps tops on decode, the decode are always all over the place my best guess is again the pcie lane on the v100, as it goes from 20-60 anywhere in between that,  I have seen pp going up, speedup for multigpu, speedup for linux on 32gb ram, and such, but I dont see much changes on mine, am I missing something?

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。