Issues / #679

#679 Feature Request: Load/Offload/Swap mmproj dynamically from RAM to VRAM when used

closed · @Interpause · 12 comments · View on GitHub

Description

mmproj is only used when images are attached, and AFAIK all images can be encoded to image tokens first before prompt processing with the rest of the text tokens. As such, swapping mmproj from RAM to VRAM only when used is a way to reduce VRAM use and hence keep more experts in VRAM.

See discussion on llama.cpp for ideas on how this can be done optimally for max speed: https://github.com/ggml-org/llama.cpp/discussions/20246

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.