Pull requests / #50
#50 tools: add mmproj quantization utility
closed · @code-martin · 0 コメント · GitHub で見る
Setup & installAMD / HIPModels & quants
本文
### Summary Adds a standalone utility script ( ools/quantize_mmproj.py) to quantize multimodal projector (mmproj) GGUF files for Strata. ### Motivation By default, multimodal projectors distributed for vision models (such as Swift-Qwen) are shipped in unquantized 16-bit float formats (BF16 or F16), consuming ~907 MB of RAM/VRAM. On memory-constrained GPUs (8 GB – 12 GB), this footprint directly reduces the VRAM available for the MoE expert cache by over 300–600 slots. Quantizing the projector to Q4_0 reduces its size from ~907 MB down to ~461 MB (~50% reduction) without noticeable loss in grounding and visual perception quality. ### Key Changes - ** ools/quantize_mmproj.py**: - Automatically locates standard llama-quantize binaries in PATH or standard system installation directories. - Automatically derives output filename from source format tag (e.g. mmproj-...-BF16.gguf -> mmproj-...-Q4_0.gguf). - Supports --type Q4_0 (default), Q8_0, or any supported GGML quantization format. - Reports byte savings and reduction percentages. ### Testing - Tested on mmproj-Swift-Qwen3.8-Flash-Next-BF16.gguf (907 MB) -> successfully generated mmproj-Swift-Qwen3.8-Flash-Next-Q4_0.gguf (461 MB). - Verified end-to-end with strata-vision.exe on live image queries.
関連リンク
インストール・モデル・リリースへの站内リンク。