Pull requests / #1014
#1014 vision: downscale large images before CPU encoding (12MP: >8 min → 20 s)
closed · @Dylan997S · 0 コメント · GitHub で見る
Setup & installServer & APIModels & quantsDocumentation
本文
### Background When the vision encoder runs on the CPU (`--vision` without `--vision gpu`), sending a phone photo (12 MP, 4032×3024) takes many minutes to encode — and since the server handles one request at a time, everything else queues behind it. The docs list "Pictures are slow (10-30 s)" as expected for CPU mode, but real-world phone photos are far worse than that range. ### Root cause `Vision.normalize()` passes JPEG/PNG/BMP/GIF through untouched, so the full-resolution image goes to `strata-vision`. But the vision encoder works on a fixed token budget (about 1024 image tokens max, fewer in CPU mode) — mtmd scales the image down to that budget internally anyway. The extra megapixels never reach the model; the CPU just burns minutes preprocessing and encoding detail that gets discarded. ### Change In `normalize()`, downscale pictures larger than 1024 px on the long side (PIL `LANCZOS`) before handing them to the encoder. Deliberately conservative: - Images at or below 1024 px keep the exact old passthrough path — **byte-identical output**, verified by test. - Without Pillow installed, behavior is completely unchanged (graceful fallback to the old path). - `max_side` is a parameter (default 1024) so it can be tuned later. ### A/B test Machine: Ryzen 7 9800X3D (8 cores), CPU vision encoder, 12 MP JPEG test photo (4032×3024, 280 KB). Timed at the `strata-vision` binary level (`ENC`), so the numbers isolate exactly what this patch changes: | | before (upstream) | after (this PR) | |---|---|---| | `normalize()` | 0.00 s → 280 KB passthrough | 0.10 s → 72 KB (1024×768) | | binary `ENC` | **>513 s, still not done** (killed) | **20.0 s** → `OK`, 768 image tokens | | speedup | — | **>25×** | Regression check: a 640×480 JPEG produces byte-identical output before/after the patch. End-to-end through the server (`/v1/chat/completions`): 21.3 s first request, 1.1 s on cache hit, and the model describes the test image correctly. ### Quality No meaningful quality loss: the encoder's token budget discards the extra detail either way — this patch just does the downscale in 0.1 s with PIL instead of letting the CPU chew on 12 MP for minutes. The 1024 px cap matches the encoder's ~1024-token budget, so the model sees the same information.
関連リンク
インストール・モデル・リリースへの站内リンク。