Issues / #625

#625 Vision (CPU path): offer a Q8_0 mmproj and a higher image-token cap

closed · @Thxeverybody · 4 commentaires · Sur GitHub

Setup & installServer & APIAMD / HIPModels & quants

Description

Context
setup.py hardcodes VISION = {"gpu": {"max_tokens": 1024}, "cpu": {"max_tokens": 300}} and always downloads mmproj-Qwen3.8-Flash-Next-BF16.gguf (~0.9 GB). On single-GPU AMD setups (e.g. an RX 9070 XT, 16 GB, also driving the desktop) vision runs on the CPU per --vision cpu — so the encoder is RAM/CPU-bound, and the 300-token image cap limits detail for charts, small text and screenshots.

Observation
strata-vision passes the mmproj straight to mtmd_init_from_file(), and the vendored mtmd supports quantized vision weights — the F32-only check in clip.cpp covers only vector/scalar metadata, not the tower weights. I run this model's mmproj quantized to Q8_0 under llama.cpp at higher image resolutions with good results. On x86 there is no native BF16 GEMM in ggml (BF16 → F32 conversion + F32 matmul), while Q8_0 has optimized AVX-512 dot-product kernels — so a Q8_0 tower should actually encode faster on the CPU at half the memory (~460 MB vs ~908 MB). Q8_0 is generally treated as near-lossless for vision encoders. The engine side is unaffected: it only reads the f32 embeddings (.sve), whose shape doesn't change with the mmproj quantization.

Proposal
Offer a Q8_0 mmproj alongside the BF16 one in the HF repo (or a documented one-liner with llama-quantize, or a --mmproj q8_0 setup flag).
Raise/make configurable the CPU vision token cap. serve/server.py already forwards vision.max_tokens to --max-tokens; only the preset hardcodes 300. The code's own comment measures one CPU encode at “~6 s at 1,024 tokens” — with a Q8_0 tower that seems acceptable as a documented option (e.g. --vision cpu-high, or an image_tokens key).
Optional: expose mtmd's image_min_tokens next to image_max_tokens for upscaling small images.
Suggested validation
Same image set at 300 / 768 / 1024 image tokens, BF16 vs Q8_0 mmproj: CPU encode time, peak RSS, and a small accuracy check (chart reading, OCR). Happy to test on an RX 9070 XT (gfx1201) / Ryzen 9700X and report numbers.

This issue was drafted with support of qwen3.8-flash-next-iq3_s under Strata and Hermes 

Sur le site

Liens install, modèles, releases.