Issues / #1404

#1404 [Feature]: run the image encoder on a second PC

open · @TraceRecursion · 2 コメント · GitHub で見る

BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentation

本文

**Goal:** read pictures at GPU speed without spending ~1.4 GB of VRAM on the image encoder.

**What I use today:** `--vision cpu`, because the troubleshooting table's other option costs the
VRAM. It works, but a 1024x1024 picture takes **15.3 s** on this machine (Ryzen 7 9700X, RTX 5080
16 GB), measured the way a client sees it.

**What I did:** `strata-vision` is already a separate process behind a small line protocol, and it is
not tied to the server's machine — so I run it on a second PC (i7-11800H + RTX 3060 Laptop 12 GB)
over ssh. Same config, same pictures, only the encoder different; median of 5 runs each:

| encoder | one 1024x1024 picture | local VRAM | expert cache |
| --- | ---: | ---: | --- |
| this GPU | 1.63 s | 1383 MiB | 3112 experts, 5.96 GiB |
| **another PC** | **2.83 s** | **0** | 3641 experts, 6.96 GiB |
| this CPU (8 threads) | 15.27 s | 0 | 3641 experts, 6.96 GiB |

That is 5.4x faster than the CPU encoder, and 1.00 GiB more expert cache than the GPU one (from the
engine's own startup line). It is 1.2 s a picture slower than a local GPU encoder, which is the price
of not spending the VRAM.

**What it does not do:** make text faster. The expert cache does grow (3112 to 3641 experts), but
long-context text is identical in the three modes — 3.27/3.30/3.31 K prompt tok/s at 257 K tokens and
3.03/3.05/3.04 K at 463 K, within run-to-run noise. ~1 GiB against a ~47 GiB expert set is ~2%
coverage. It is a picture-latency and VRAM trade, not a text-speed one.

I have it working behind an off-by-default flag, with tests and docs. **Do you want it?** Or is
`--vision cpu` the answer you would rather give people? If you do want it: config key
(`"remote": true`) or `--vision remote` in setup?

(It needs two narrow server-side changes: skip the request FIFO when the encoder is remote — a local
encoder shares the GPU and must be serialised, a remote one shares nothing, and holding the FIFO
across a slow upload froze the whole API, text included — and start serving text when the encoder is
not there, so a vision host that is off does not cost the server. Detail and measurements are ready
for the PR.)

関連リンク

インストール・モデル・リリースへの站内リンク。