Issues / #767
#767 --vision cpu caps images at 300 tokens; llama.cpp's mtmd asks for >=1024 for Qwen-VL grounding
closed · @miskahm · 3 comentarios · En GitHub
Setup & installAMD / HIPModels & quantsDocumentation
Descripción
## What On the AMD backend `--vision cpu` is the only image path (#304), and `VISION["cpu"]["max_tokens"]` in `setup.py` caps an image at **300** tokens (the GPU path gets 1024). llama.cpp's mtmd prints this at encoder start: ``` load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks load_hparams: if you encounter problems with accuracy, try adding --image-min-tokens 1024 load_hparams: more info: https://github.com/ggml-org/llama.cpp/issues/16842 ``` `strata_vision.cpp` exposes `--max-tokens` but not `--image-min-tokens`, so that advice cannot be followed from Strata's config — the 300-token cap is also the *minimum* mtmd will use here, so the two settings are tied. ## Why it matters Fine-grained visual questions are where a 300-token grid is weakest: pointing at an object, counting small items, reading small text inside a large scene, OCR-ish tasks. On this machine the encoder produced 260 tokens for a 640x400 image (grid 20x13) in 2.73 s, so a larger budget is affordable when the image warrants it — the cost is linear in tokens, and it is paid once per image (results are cached by image hash, so a conversation sends it once). ## Suggestion (not a patch) Either of these would do, and neither changes the wire format: - pass `--image-min-tokens` through to mtmd alongside the existing `--max-tokens`, and pick the pair per path (e.g. CPU `max-tokens 1024` / `image-min-tokens 300`, so small images stay cheap and large ones get the resolution the model wants); or - make `VISION["cpu"]["max_tokens"]` configurable, the way `VISION["gpu"]` is a table today. ## Also: the documented CPU encode cost looks pessimistic `docs/INSTALL.md` and `docs/TROUBLESHOOTING.md` say a picture takes **10-30 s** on the CPU encoder. Measured here (RX 7900 XTX host, engine 0.1.38, Swift 1.5 IQ3_XXS, 8 threads, `VISION["cpu"]["max_tokens"]=300`, 640x400 PNG): **2.73 s**, and 2.74 s on the repeat: ``` $ printf 'ENC shot.png /tmp/o.sve\nQUIT\n' | engine/strata-vision --mmproj … --model … --max-tokens 300 --threads 8 READY 2560 OK 260 20 13 2727 ``` End to end, a request with an image took 7.0 s wall (2.7 s encode + 321 prompt tokens prefill + ~120 tokens generated). If 10-30 s was measured on a slower CPU or at the GPU path's 1024 tokens, the docs could say so — the current figure sets the wrong expectation for anyone deciding whether `--vision cpu` is usable. Environment: RX 7900 XTX (gfx1100), Ryzen 7 9800X3D, 60 GB RAM, engine 0.1.38 (`99f3dbd`), inside `rocm/dev-ubuntu-24.04:10.0.0-full`.
En el sitio
Enlaces a install, modelos, releases.