Issues / #1404

#1404 [Feature]: run the image encoder on a second PC

open · @TraceRecursion · 2 comments · View on GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentation

Description

**Goal:** read pictures at GPU speed without spending ~1.4 GB of VRAM on the image encoder.

**What I use today:** `--vision cpu`, because the troubleshooting table's other option costs the
VRAM. It works, but a 1024x1024 picture takes **15.3 s** on this machine (Ryzen 7 9700X, RTX 5080
16 GB), measured the way a client sees it.

**What I did:** `strata-vision` is already a separate process behind a small line protocol, and it is
not tied to the server's machine — so I run it on a second PC (i7-11800H + RTX 3060 Laptop 12 GB)
over ssh. Same config, same pictures, only the encoder different; median of 5 runs each:

| encoder | one 1024x1024 picture | local VRAM | expert cache |
| --- | ---: | ---: | --- |
| this GPU | 1.63 s | 1383 MiB | 3112 experts, 5.96 GiB |
| **another PC** | **2.83 s** | **0** | 3641 experts, 6.96 GiB |
| this CPU (8 threads) | 15.27 s | 0 | 3641 experts, 6.96 GiB |

That is 5.4x faster than the CPU encoder, and 1.00 GiB more expert cache than the GPU one (from the
engine's own startup line). It is 1.2 s a picture slower than a local GPU encoder, which is the price
of not spending the VRAM.

**What it does not do:** make text faster. The expert cache does grow (3112 to 3641 experts), but
long-context text is identical in the three modes — 3.27/3.30/3.31 K prompt tok/s at 257 K tokens and
3.03/3.05/3.04 K at 463 K, within run-to-run noise. ~1 GiB against a ~47 GiB expert set is ~2%
coverage. It is a picture-latency and VRAM trade, not a text-speed one.

I have it working behind an off-by-default flag, with tests and docs. **Do you want it?** Or is
`--vision cpu` the answer you would rather give people? If you do want it: config key
(`"remote": true`) or `--vision remote` in setup?

(It needs two narrow server-side changes: skip the request FIFO when the encoder is remote — a local
encoder shares the GPU and must be serialised, a remote one shares nothing, and holding the FIFO
across a slow upload froze the whole API, text included — and start serving text when the encoder is
not there, so a vision host that is off does not cost the server. Detail and measurements are ready
for the PR.)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.