Issues / #679

#679 Feature Request: Load/Offload/Swap mmproj dynamically from RAM to VRAM when used

closed · @Interpause · 12 评论 · 在 GitHub 查看

描述

mmproj is only used when images are attached, and AFAIK all images can be encoded to image tokens first before prompt processing with the rest of the text tokens. As such, swapping mmproj from RAM to VRAM only when used is a way to reduce VRAM use and hence keep more experts in VRAM.

See discussion on llama.cpp for ideas on how this can be done optimally for max speed: https://github.com/ggml-org/llama.cpp/discussions/20246

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。