Issues / #1631

#1631 [Feature Request]: Support RAM-budgeted SSD-backed experts with multi-GPU layer splitting

open · @KapellMeister61 · 0 コメント · GitHub で見る

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quants

本文

Hi!

First off, really appreciate the work going into Strata. I've been running Qwen3.8 Flash Next with it for a while now, and the performance has been surprisingly good considering my hardware.

I'm currently getting around **1,650 tok/s prefill and consistently 50+ tok/s decoding**, using just an RTX 3090 Ti, system RAM, and SSD-backed cold experts.

I also have an RTX 5060 Ti 16GB in the same machine, but I've run into a limitation when trying to use both GPUs.

### What I'm trying to do

My current setup uses `--resident-budget-gib 40`, which works really well. It keeps the RAM usage under control while allowing the remaining cold experts to stay on storage.

However, when I enable multi-GPU layer splitting, Strata no longer allows me to use the resident RAM budget. Instead, it wants to load the remaining experts into system RAM.

That's a bit of a problem since my inference LXC only has 64 GiB allocated, even though the Proxmox host itself has 128GB.

What I'd love to see is the ability to combine both features:

- RTX 3090 Ti + RTX 5060 Ti using layer splitting
- A configurable RAM budget for resident experts (`--resident-budget-gib`)
- Remaining cold experts staying SSD-backed and loaded when needed

Basically, the same tiered memory setup that's already working with one GPU, but with both GPUs participating.

### My hardware

- **CPU:** Intel Core i7-14700K
- **RAM:** 128GB DDR4-3000, with 64 GiB allocated to the inference LXC
- **GPU 0:** RTX 3090 Ti 24GB (PCIe 4.0 x16)
- **GPU 1:** RTX 5060 Ti 16GB (PCIe 3.0 x4)
- **Storage:** Local ZFS filesystem with SSD-backed expert storage
- **NVIDIA driver:** 580.82.07
- **Environment:** Proxmox LXC, kernel 6.14.11-1-pve
- **Model:** Qwen3.8 Flash Next UD-IQ4_XS

I'm currently running **Strata v0.1.40.3**, commit `d5ea7133741e67743c0e886bb426c0ce8d69cf6c`.

### Current working configuration

These are my actual engine arguments, with local paths shortened:

```text
--pack <pack-path>
--native <model-path>
--expert-profile <profile-path>
--expert-cache auto
--prefill auto
--spec 4
--spec-min-p 0.5
--mtp <mtp-path>
--max-context 204800
--kv int8
--kv-resident 32768
--resident-budget-gib 40
--pool-workers 13
```

The server config currently uses `"gpu": 0` with no layer split.

With this configuration, I'm seeing approximately:

- **Prefill:** 1,650 tok/s
- **Decode:** 50+ tok/s
- **3090 Ti VRAM:** ~23.4 GiB used
- **5060 Ti VRAM:** unused
- **LXC RAM:** ~47 GiB used out of 64 GiB
- **Swap:** disabled

These are observations from regular usage, not controlled benchmark results.

### Is this something that could be supported?

I noticed that PR #848 added multi-GPU support for resident experts, which looks like a step in this direction. But from what I understand, that still works differently from the budgeted RAM mode, where cold experts can remain SSD-backed.

Would it be possible to extend that functionality to support `--resident-budget-gib` alongside layer splitting?

I'm not expecting the second GPU to automatically double performance. With my 5060 Ti on a PCIe x4 connection, there might even be some trade-offs. But having another 16GB of VRAM available for caching and compute, without losing SSD-backed cold experts, would be really interesting to experiment with.

I imagine this could also be useful for people running larger MoE models on systems with limited RAM, especially as newer models come out.

I found a few related discussions:

- #498 — Multi-GPU setup and RAM budget
- #384 — Resident experts with layer splitting
- #604 — Expert placement on mixed GPUs
- PR #848 — Resident RAM mode with layer splitting

I'm happy to test an experimental implementation or provide additional logs and benchmarks if that would help.

Thanks again for the project!

関連リンク

インストール・モデル・リリースへの站内リンク。