Issues / #679

#679 Feature Request: Load/Offload/Swap mmproj dynamically from RAM to VRAM when used

closed · @Interpause · 12 comentarios · En GitHub

Descripción

mmproj is only used when images are attached, and AFAIK all images can be encoded to image tokens first before prompt processing with the rest of the text tokens. As such, swapping mmproj from RAM to VRAM only when used is a way to reduce VRAM use and hence keep more experts in VRAM.

See discussion on llama.cpp for ideas on how this can be done optimally for max speed: https://github.com/ggml-org/llama.cpp/discussions/20246

En el sitio

Enlaces a install, modelos, releases.