Pull requests / #1014

#1014 vision: downscale large images before CPU encoding (12MP: >8 min → 20 s)

closed · @Dylan997S · 0 Kommentare · Auf GitHub

Setup & installServer & APIModels & quantsDocumentation

Beschreibung

### Background
When the vision encoder runs on the CPU (`--vision` without `--vision gpu`), sending a phone photo (12 MP, 4032×3024) takes many minutes to encode — and since the server handles one request at a time, everything else queues behind it. The docs list "Pictures are slow (10-30 s)" as expected for CPU mode, but real-world phone photos are far worse than that range.

### Root cause
`Vision.normalize()` passes JPEG/PNG/BMP/GIF through untouched, so the full-resolution image goes to `strata-vision`. But the vision encoder works on a fixed token budget (about 1024 image tokens max, fewer in CPU mode) — mtmd scales the image down to that budget internally anyway. The extra megapixels never reach the model; the CPU just burns minutes preprocessing and encoding detail that gets discarded.

### Change
In `normalize()`, downscale pictures larger than 1024 px on the long side (PIL `LANCZOS`) before handing them to the encoder. Deliberately conservative:
- Images at or below 1024 px keep the exact old passthrough path — **byte-identical output**, verified by test.
- Without Pillow installed, behavior is completely unchanged (graceful fallback to the old path).
- `max_side` is a parameter (default 1024) so it can be tuned later.

### A/B test
Machine: Ryzen 7 9800X3D (8 cores), CPU vision encoder, 12 MP JPEG test photo (4032×3024, 280 KB). Timed at the `strata-vision` binary level (`ENC`), so the numbers isolate exactly what this patch changes:

| | before (upstream) | after (this PR) |
|---|---|---|
| `normalize()` | 0.00 s → 280 KB passthrough | 0.10 s → 72 KB (1024×768) |
| binary `ENC` | **>513 s, still not done** (killed) | **20.0 s** → `OK`, 768 image tokens |
| speedup | — | **>25×** |

Regression check: a 640×480 JPEG produces byte-identical output before/after the patch.

End-to-end through the server (`/v1/chat/completions`): 21.3 s first request, 1.1 s on cache hit, and the model describes the test image correctly.

### Quality
No meaningful quality loss: the encoder's token budget discards the extra detail either way — this patch just does the downscale in 0.1 s with PIL instead of letting the CPU chew on 12 MP for minutes. The 1024 px cap matches the encoder's ~1024-token budget, so the model sees the same information.

Mehr auf der Site

Links zu Install, Modellen, Releases.