Issues / #679
#679 Feature Request: Load/Offload/Swap mmproj dynamically from RAM to VRAM when used
closed · @Interpause · 12 comentarios · En GitHub
Descripción
mmproj is only used when images are attached, and AFAIK all images can be encoded to image tokens first before prompt processing with the rest of the text tokens. As such, swapping mmproj from RAM to VRAM only when used is a way to reduce VRAM use and hence keep more experts in VRAM. See discussion on llama.cpp for ideas on how this can be done optimally for max speed: https://github.com/ggml-org/llama.cpp/discussions/20246
En el sitio
Enlaces a install, modelos, releases.