Issues / #1352
#1352 Multi-GPU performance is ~2x lower than Francesco Albano fork despite the changes being merged into v0.1.40
open · @chemical12 · 5 comentários · No GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux
Descrição
Hi, I noticed in the v0.1.40 release notes that several of the multi-GPU improvements from the Hardin22/Strata-DualGPU fork were incorporated into Strata. However, I am seeing a significant performance difference when running the same model and the same two GPUs. Hardware: [0] RTX 4070 Ti SUPER 16 GB [2] RTX 5060 Ti 16 GB 128 GB system RAM Linux Debian Model: Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S Strata configuration I am using: --expert-cache auto --prefill auto:32768 --spec 4 --spec-min-p 0.5 --mtp ... --max-context 262144 --kv int8 --kv-resident 32768 --vision --vram-reserve-mib 700 --remote-expert-opt --pipeline-windows 2 GPU: [0, 2] layer_split: auto With Niko1221/Strata v0.1.40, I get approximately: ~40 tok/s decode With the Hardin22/Strata-DualGPU fork, using the same model and the same two GPUs, I get approximately: ~80 tok/s decode (110 tok/s when coding) I also tried explicitly enabling: --resident-experts --adapt-async 1 --pipeline-windows 2 in v0.1.40, but this did not reproduce the Hardin22 performance. There is also a significant difference in prompt processing speed depending on the GPU order. With: GPU: [0, 2] a prompt of approximately 6k tokens takes around 50 seconds to process. If I change the order to: GPU: [2, 0] the same ~6k-token prompt takes only around 15 seconds. However, the decode speed is essentially identical in both configurations (around 30-40 tok/s).
No site
Links install, modelos, releases.