Issues / #514
#514 Feature Request: use AMD iGPU instead of CPU for offloading
closed · @Vitaliy86 · 1 comentários · No GitHub
Multi-GPUAMD / HIPNVIDIA / CUDA
Descrição
I would like to suggest an idea for systems with an AMD Ryzen processor with an integrated GPU and an NVIDIA discrete GPU.
Would it be possible to use the AMD integrated GPU as an alternative to CPU offloading in Strata?
The idea is not to use NVIDIA and AMD as two equal GPUs for the same inference workload.
Instead, the desired configuration would be:
NVIDIA GPU
↓
main inference
AMD iGPU
↓
parts that would otherwise be CPU-offloaded
System RAM
↓
additional model storage
In other words:
Current:
NVIDIA GPU + CPU
Potential:
NVIDIA GPU + AMD iGPU
The AMD iGPU would effectively replace the CPU as the secondary compute device.
Why this could be useful
I have personally tested my AMD integrated GPU with image-generation workloads.
In my measurements, the AMD iGPU was approximately 40% faster than CPU-only execution for the workload I tested.
This is not intended as a general claim about AMD GPUs being 40% faster for AI workloads. It is simply the result of my own testing with image generation.
This made me wonder whether the same approach could be useful for LLM inference.
Instead of leaving the integrated GPU unused while CPU offloading is performed, Strata could potentially use the AMD iGPU as an additional compute device.
Possible architecture
Conceptually:
Strata
│
NVIDIA CUDA
│
primary inference
│
overflow/offload
│
AMD iGPU
│
system memory
The goal would therefore not be a general heterogeneous multi-GPU system.
The goal would be much simpler:
«Can the existing CPU offload path be extended so that an AMD iGPU can perform the offloaded computation instead of the CPU?»
I noticed that Strata already has experimental AMD/HIP support, so perhaps some of the necessary infrastructure already exists.
I also understand that integrated AMD GPUs are different from discrete Radeon GPUs because they use shared system memory.
Possible first implementation
Even an experimental implementation would be interesting.
For example:
NVIDIA RTX
├── main model computation
└── VRAM-resident data
AMD iGPU
└── CPU-offloaded layers / experts
System RAM
└── model weights and shared memory
The CPU could then remain primarily responsible for orchestration, data preparation and other non-GPU tasks.
This could be especially useful for desktop systems with:
- AMD Ryzen CPU
- AMD integrated GPU
- NVIDIA RTX discrete GPU
Would it be technically feasible to investigate using the AMD iGPU as a replacement for CPU offloading, rather than treating it as a separate primary GPU backend?
Even if this is currently difficult due to CUDA/HIP/backend limitations, I think it could be an interesting direction for future Strata development.No site
Links install, modelos, releases.