Issues / #1352

#1352 Multi-GPU performance is ~2x lower than Francesco Albano fork despite the changes being merged into v0.1.40

open · @chemical12 · 5 comentários · No GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux

Descrição

Hi,

I noticed in the v0.1.40 release notes that several of the multi-GPU improvements from the Hardin22/Strata-DualGPU fork were incorporated into Strata.
However, I am seeing a significant performance difference when running the same model and the same two GPUs.

Hardware:
[0] RTX 4070 Ti SUPER 16 GB
[2] RTX 5060 Ti 16 GB

128 GB system RAM
Linux Debian

Model:
Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S

Strata configuration
I am using:

--expert-cache auto
--prefill auto:32768
--spec 4
--spec-min-p 0.5
--mtp ...
--max-context 262144
--kv int8
--kv-resident 32768
--vision
--vram-reserve-mib 700
--remote-expert-opt
--pipeline-windows 2

GPU: [0, 2]
layer_split: auto

With Niko1221/Strata v0.1.40, I get approximately: ~40 tok/s decode

With the Hardin22/Strata-DualGPU fork, using the same model and the same two GPUs, I get approximately: ~80 tok/s decode (110 tok/s when coding)

I also tried explicitly enabling:
--resident-experts
--adapt-async 1
--pipeline-windows 2
in v0.1.40, but this did not reproduce the Hardin22 performance.

There is also a significant difference in prompt processing speed depending on the GPU order.
With: GPU: [0, 2] a prompt of approximately 6k tokens takes around 50 seconds to process.
If I change the order to: GPU: [2, 0] the same ~6k-token prompt takes only around 15 seconds.

However, the decode speed is essentially identical in both configurations (around 30-40 tok/s).

No site

Links install, modelos, releases.