Pull requests / #50

#50 tools: add mmproj quantization utility

closed · @code-martin · 0 评论 · 在 GitHub 查看

Setup & installAMD / HIPModels & quants

描述

### Summary
Adds a standalone utility script (	ools/quantize_mmproj.py) to quantize multimodal projector (mmproj) GGUF files for Strata.

### Motivation
By default, multimodal projectors distributed for vision models (such as Swift-Qwen) are shipped in unquantized 16-bit float formats (BF16 or F16), consuming ~907 MB of RAM/VRAM. On memory-constrained GPUs (8 GB – 12 GB), this footprint directly reduces the VRAM available for the MoE expert cache by over 300–600 slots.

Quantizing the projector to Q4_0 reduces its size from ~907 MB down to ~461 MB (~50% reduction) without noticeable loss in grounding and visual perception quality.

### Key Changes
- **	ools/quantize_mmproj.py**:
  - Automatically locates standard llama-quantize binaries in PATH or standard system installation directories.
  - Automatically derives output filename from source format tag (e.g. mmproj-...-BF16.gguf -> mmproj-...-Q4_0.gguf).
  - Supports --type Q4_0 (default), Q8_0, or any supported GGML quantization format.
  - Reports byte savings and reduction percentages.

### Testing
- Tested on mmproj-Swift-Qwen3.8-Flash-Next-BF16.gguf (907 MB) -> successfully generated mmproj-Swift-Qwen3.8-Flash-Next-Q4_0.gguf (461 MB).
- Verified end-to-end with strata-vision.exe on live image queries.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。