反馈 / #1631
#1631 [Feature Request]: Support RAM-budgeted SSD-backed experts with multi-GPU layer splitting
open · @KapellMeister61 · 0 评论 · 去 GitHub 看
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quants
说明
Hi! First off, really appreciate the work going into Strata. I've been running Qwen3.8 Flash Next with it for a while now, and the performance has been surprisingly good considering my hardware. I'm currently getting around **1,650 tok/s prefill and consistently 50+ tok/s decoding**, using just an RTX 3090 Ti, system RAM, and SSD-backed cold experts. I also have an RTX 5060 Ti 16GB in the same machine, but I've run into a limitation when trying to use both GPUs. ### What I'm trying to do My current setup uses `--resident-budget-gib 40`, which works really well. It keeps the RAM usage under control while allowing the remaining cold experts to stay on storage. However, when I enable multi-GPU layer splitting, Strata no longer allows me to use the resident RAM budget. Instead, it wants to load the remaining experts into system RAM. That's a bit of a problem since my inference LXC only has 64 GiB allocated, even though the Proxmox host itself has 128GB. What I'd love to see is the ability to combine both features: - RTX 3090 Ti + RTX 5060 Ti using layer splitting - A configurable RAM budget for resident experts (`--resident-budget-gib`) - Remaining cold experts staying SSD-backed and loaded when needed Basically, the same tiered memory setup that's already working with one GPU, but with both GPUs participating. ### My hardware - **CPU:** Intel Core i7-14700K - **RAM:** 128GB DDR4-3000, with 64 GiB allocated to the inference LXC - **GPU 0:** RTX 3090 Ti 24GB (PCIe 4.0 x16) - **GPU 1:** RTX 5060 Ti 16GB (PCIe 3.0 x4) - **Storage:** Local ZFS filesystem with SSD-backed expert storage - **NVIDIA driver:** 580.82.07 - **Environment:** Proxmox LXC, kernel 6.14.11-1-pve - **Model:** Qwen3.8 Flash Next UD-IQ4_XS I'm currently running **Strata v0.1.40.3**, commit `d5ea7133741e67743c0e886bb426c0ce8d69cf6c`. ### Current working configuration These are my actual engine arguments, with local paths shortened: ```text --pack <pack-path> --native <model-path> --expert-profile <profile-path> --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <mtp-path> --max-context 204800 --kv int8 --kv-resident 32768 --resident-budget-gib 40 --pool-workers 13 ``` The server config currently uses `"gpu": 0` with no layer split. With this configuration, I'm seeing approximately: - **Prefill:** 1,650 tok/s - **Decode:** 50+ tok/s - **3090 Ti VRAM:** ~23.4 GiB used - **5060 Ti VRAM:** unused - **LXC RAM:** ~47 GiB used out of 64 GiB - **Swap:** disabled These are observations from regular usage, not controlled benchmark results. ### Is this something that could be supported? I noticed that PR #848 added multi-GPU support for resident experts, which looks like a step in this direction. But from what I understand, that still works differently from the budgeted RAM mode, where cold experts can remain SSD-backed. Would it be possible to extend that functionality to support `--resident-budget-gib` alongside layer splitting? I'm not expecting the second GPU to automatically double performance. With my 5060 Ti on a PCIe x4 connection, there might even be some trade-offs. But having another 16GB of VRAM available for caching and compute, without losing SSD-backed cold experts, would be really interesting to experiment with. I imagine this could also be useful for people running larger MoE models on systems with limited RAM, especially as newer models come out. I found a few related discussions: - #498 — Multi-GPU setup and RAM budget - #384 — Resident experts with layer splitting - #604 — Expert placement on mixed GPUs - PR #848 — Resident RAM mode with layer splitting I'm happy to test an experimental implementation or provide additional logs and benchmarks if that would help. Thanks again for the project!
本站相关内容
相关页面的快捷入口。