Issues / #642
#642 A dual-GPU fork for 32 GB PCs (resident mode on a layer split, pipelined windows)
closed · @Hardin22 · 10 评论 · 在 GitHub 查看
BenchmarksSetup & installDocumentationWindows
描述
Hi Niko, thanks for Strata, great work. Really. I have a 5080 + 4060 Ti with 32 GB of RAM, where setup recommends one card because the resident low-RAM mode has no layer split. I made it work on a split and then pipelined the two stages, so the 4060 Ti runs the next window while the 5080 verifies the current one. On my PC that took Swift 1.5 IQ2_XS from 29 tok/s (5080 alone) to 161tok/s on code (200+ peaks), with the same perplexity as 0.1.38 on a teacher-forced comparison. It's a fork on top of 0.1.38: https://github.com/Hardin22/Strata-DualGPU What changed and what each part measured: docs/DUAL_GPU.md Some of it isn't specific to two GPUs, and I'd be glad to send PRs if you want any of it: - the hybrid-CPU pool default (P-cores plus half of the E-cores): on an i9-14900KF, 15 workers decoded much faster than all 23 - copying an evicted expert back from the card that owns its layer (resident mode on a split; with one card it doesn't matter) - --vram-reserve-later-mib, a separate reserve for the cards that don't drive a display No pressure either way. I'm going to post the fork on r/LocalLLaMA and wanted you to hear about it here first. PS i do believe this can be pushed even further.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。