Issues / #767

#767 --vision cpu caps images at 300 tokens; llama.cpp's mtmd asks for >=1024 for Qwen-VL grounding

closed · @miskahm · 3 commentaires · Sur GitHub

Setup & installAMD / HIPModels & quantsDocumentation

Description

## What

On the AMD backend `--vision cpu` is the only image path (#304), and `VISION["cpu"]["max_tokens"]` in
`setup.py` caps an image at **300** tokens (the GPU path gets 1024). llama.cpp's mtmd prints this at encoder
start:

```
load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks
load_hparams: if you encounter problems with accuracy, try adding --image-min-tokens 1024
load_hparams: more info: https://github.com/ggml-org/llama.cpp/issues/16842
```

`strata_vision.cpp` exposes `--max-tokens` but not `--image-min-tokens`, so that advice cannot be followed
from Strata's config — the 300-token cap is also the *minimum* mtmd will use here, so the two settings are
tied.

## Why it matters

Fine-grained visual questions are where a 300-token grid is weakest: pointing at an object, counting small
items, reading small text inside a large scene, OCR-ish tasks. On this machine the encoder produced 260
tokens for a 640x400 image (grid 20x13) in 2.73 s, so a larger budget is affordable when the image warrants
it — the cost is linear in tokens, and it is paid once per image (results are cached by image hash, so a
conversation sends it once).

## Suggestion (not a patch)

Either of these would do, and neither changes the wire format:

- pass `--image-min-tokens` through to mtmd alongside the existing `--max-tokens`, and pick the pair per path
  (e.g. CPU `max-tokens 1024` / `image-min-tokens 300`, so small images stay cheap and large ones get the
  resolution the model wants); or
- make `VISION["cpu"]["max_tokens"]` configurable, the way `VISION["gpu"]` is a table today.

## Also: the documented CPU encode cost looks pessimistic

`docs/INSTALL.md` and `docs/TROUBLESHOOTING.md` say a picture takes **10-30 s** on the CPU encoder. Measured
here (RX 7900 XTX host, engine 0.1.38, Swift 1.5 IQ3_XXS, 8 threads, `VISION["cpu"]["max_tokens"]=300`,
640x400 PNG): **2.73 s**, and 2.74 s on the repeat:

```
$ printf 'ENC shot.png /tmp/o.sve\nQUIT\n' | engine/strata-vision --mmproj … --model … --max-tokens 300 --threads 8
READY 2560
OK 260 20 13 2727
```

End to end, a request with an image took 7.0 s wall (2.7 s encode + 321 prompt tokens prefill + ~120 tokens
generated). If 10-30 s was measured on a slower CPU or at the GPU path's 1024 tokens, the docs could say so —
the current figure sets the wrong expectation for anyone deciding whether `--vision cpu` is usable.

Environment: RX 7900 XTX (gfx1100), Ryzen 7 9800X3D, 60 GB RAM, engine 0.1.38 (`99f3dbd`), inside
`rocm/dev-ubuntu-24.04:10.0.0-full`.

Sur le site

Liens install, modèles, releases.