Issues / #642

#642 A dual-GPU fork for 32 GB PCs (resident mode on a layer split, pipelined windows)

closed · @Hardin22 · 10 comentários · No GitHub

BenchmarksSetup & installDocumentationWindows

Descrição

Hi Niko, thanks for Strata, great work. Really. 

I have a 5080 + 4060 Ti with 32 GB of RAM, where setup recommends one card because the resident low-RAM mode has no layer split. I made it work on a split and then pipelined the two stages, so the 4060 Ti runs the next window while the 5080 verifies the current one. On my PC that took Swift 1.5 IQ2_XS from 29 tok/s (5080 alone) to 161tok/s on code (200+ peaks), with the same perplexity as 0.1.38 on a teacher-forced comparison.

It's a fork on top of 0.1.38: https://github.com/Hardin22/Strata-DualGPU
What changed and what each part measured: docs/DUAL_GPU.md

Some of it isn't specific to two GPUs, and I'd be glad to send PRs if you want any of it:
- the hybrid-CPU pool default (P-cores plus half of the E-cores): on an i9-14900KF, 15 workers decoded much faster than all 23
- copying an evicted expert back from the card that owns its layer (resident mode on a split; with one card it doesn't matter)
- --vram-reserve-later-mib, a separate reserve for the cards that don't drive a display

No pressure either way. I'm going to post the fork on r/LocalLLaMA and wanted you to hear about it here first.
PS i do believe this can be pushed even further.

No site

Links install, modelos, releases.