Issues / #1631
#1631 [Feature Request]: Support RAM-budgeted SSD-backed experts with multi-GPU layer splitting
open · @KapellMeister61 · 0 comments · View on GitHub
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quants
Description
Hi! First off, really appreciate the work going into Strata. I've been running Qwen3.8 Flash Next with it for a while now, and the performance has been surprisingly good considering my hardware. I'm currently getting around **1,650 tok/s prefill and consistently 50+ tok/s decoding**, using just an RTX 3090 Ti, system RAM, and SSD-backed cold experts. I also have an RTX 5060 Ti 16GB in the same machine, but I've run into a limitation when trying to use both GPUs. ### What I'm trying to do My current setup uses `--resident-budget-gib 40`, which works really well. It keeps the RAM usage under control while allowing the remaining cold experts to stay on storage. However, when I enable multi-GPU layer splitting, Strata no longer allows me to use the resident RAM budget. Instead, it wants to load the remaining experts into system RAM. That's a bit of a problem since my inference LXC only has 64 GiB allocated, even though the Proxmox host itself has 128GB. What I'd love to see is the ability to combine both features: - RTX 3090 Ti + RTX 5060 Ti using layer splitting - A configurable RAM budget for resident experts (`--resident-budget-gib`) - Remaining cold experts staying SSD-backed and loaded when needed Basically, the same tiered memory setup that's already working with one GPU, but with both GPUs participating. ### My hardware - **CPU:** Intel Core i7-14700K - **RAM:** 128GB DDR4-3000, with 64 GiB allocated to the inference LXC - **GPU 0:** RTX 3090 Ti 24GB (PCIe 4.0 x16) - **GPU 1:** RTX 5060 Ti 16GB (PCIe 3.0 x4) - **Storage:** Local ZFS filesystem with SSD-backed expert storage - **NVIDIA driver:** 580.82.07 - **Environment:** Proxmox LXC, kernel 6.14.11-1-pve - **Model:** Qwen3.8 Flash Next UD-IQ4_XS I'm currently running **Strata v0.1.40.3**, commit `d5ea7133741e67743c0e886bb426c0ce8d69cf6c`. ### Current working configuration These are my actual engine arguments, with local paths shortened: ```text --pack <pack-path> --native <model-path> --expert-profile <profile-path> --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <mtp-path> --max-context 204800 --kv int8 --kv-resident 32768 --resident-budget-gib 40 --pool-workers 13 ``` The server config currently uses `"gpu": 0` with no layer split. With this configuration, I'm seeing approximately: - **Prefill:** 1,650 tok/s - **Decode:** 50+ tok/s - **3090 Ti VRAM:** ~23.4 GiB used - **5060 Ti VRAM:** unused - **LXC RAM:** ~47 GiB used out of 64 GiB - **Swap:** disabled These are observations from regular usage, not controlled benchmark results. ### Is this something that could be supported? I noticed that PR #848 added multi-GPU support for resident experts, which looks like a step in this direction. But from what I understand, that still works differently from the budgeted RAM mode, where cold experts can remain SSD-backed. Would it be possible to extend that functionality to support `--resident-budget-gib` alongside layer splitting? I'm not expecting the second GPU to automatically double performance. With my 5060 Ti on a PCIe x4 connection, there might even be some trade-offs. But having another 16GB of VRAM available for caching and compute, without losing SSD-backed cold experts, would be really interesting to experiment with. I imagine this could also be useful for people running larger MoE models on systems with limited RAM, especially as newer models come out. I found a few related discussions: - #498 — Multi-GPU setup and RAM budget - #384 — Resident experts with layer splitting - #604 — Expert placement on mixed GPUs - PR #848 — Resident RAM mode with layer splitting I'm happy to test an experimental implementation or provide additional logs and benchmarks if that would help. Thanks again for the project!
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.