Pull requests / #16

#16 Add optional multi-GPU expert caches and compact transfers

closed · draft · @Daerdaal · 0 Kommentare · Auf GitHub

Setup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsWindows

Beschreibung

Add optional multi-GPU expert caches and compact transfers. Tested on my setup (4xeGPU 3060 12Gb (1oculink / 3usb4) )
option to add multi GPU choice at startup (1, 2 or 4 GPU) : GPU use is chosen when configuring Strata:
--gpu1-experts N enables CUDA1 alongside CUDA0
adding --gpu2-experts N --gpu3-experts N enables all four GPUs.
A single-GPU configuration uses no secondary expert caches. The generated run-iq3_xxs.bat (or any other compatible model launcher) then starts the saved configuration. Switching later requires editing its JSON configuration or running setup again. I tested IQ3_XXS. Other model sizes still need to be validated on hardware.
note : my test setup have an igpu (780m) for monitor, so NO VRAM is used on the CUDA devices when server starts.
The CUDA0 GPU must be at least 12Gb, the others too or bigger. (Not tested with less VRAM)

Mehr auf der Site

Links zu Install, Modellen, Releases.