Pull requests / #666

#666 setup, serve: the image encoder's device as its own role (--vision-device)

closed · @hireymage · 0 评论 · 在 GitHub 查看

Setup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

描述

## What

A third placement role for the picture encoder, beside the engine's GPUs (`--vision-device`):

- `--vision-device auto` — a spare card (the one with the most VRAM) when one exists, else the engine's main card as before;
- `--vision-device cpu` — the pictures read on the CPU (#304's path);
- `--vision-device 1` / `cuda:1` — pinned to a card, numbered like nvidia-smi (the same numbering the engine's device report uses);
- flag absent — nothing changes (the encoder stays on the engine's first card).

## Why this shape

The image encoder is not in the engine: `strata-vision` is a process of its own, and the engine reads the embeddings it wrote from the disk. So the placement lives at the layer that owns the vision compute, and the implementation reuses what is already merged rather than adding a parallel mechanism:

- **#408**: a `cuda_device` in the `"vision"` config — this branch writes that same key, and `vision_env()` does the rest (`CUDA_VISIBLE_DEVICES`, PCI-bus order) unchanged.
- **#304**: the CPU encoder — `--vision-device cpu` simply routes into the existing `--vision cpu` path.
- Setup validates the card against the same GPU list the engine's choices use (and the HIP list when the backend is HIP), warns when the encoder must share a card with the engine (the expert cache there is sized after the encoder's VRAM), and `serve` states the placement at start-up (`the image encoder runs on GPU 1 (its own card)`).

Cross-card, the handoff is by construction the same as every existing vision run: the `.sve` embeddings go through the disk, so no card-to-card path is involved on any mix of GPUs, and P2P is never required.

Nothing changes when the flag is absent; no engine (C++) code is touched.

## Testing

- `tools/test_setup_choices.py`: 10 new tests for the flag (auto/cpu/card/share/missing-card/refusals) — 41 green; `serve/test_server.py`: the new start-up role line — 119 green.
- Validated on a 2× GTX 1080 Ti (Pascal, sm_61) rig against the real encoder binary: with `cuda_device: 1` the encoder's VRAM lands on GPU 1 alone (GPU 0 delta 0 MiB, GPU 1 +1192 MiB through the serve path), the second encode of the same image hits the disk cache, and the CPU run produces the same-size output. Pascal is affected only by proxy — no engine code changed.

## Docs

`docs/AI_SETUP.md`, `docs/DETAILS.md` (0.1.38 paragraph) and `docs/MULTI_GPU.md` describe the flag and the placement rules.


---

## Follow-up in this PR: the encoder on Vulkan or SYCL GPUs (`--vision-backend vulkan|sycl`)

The image encoder is its own llama.cpp process, so unlike the engine it can run on a GPU whose backend the engine
doesn't speak. `--vision-backend vulkan` builds `strata-vision` for Vulkan (any Vulkan GPU, an Intel or AMD iGPU
included) and `--vision-backend sycl` for Intel GPUs through oneAPI; the device is named through that backend's own
variable and its own device list:

| `--vision-backend` | encoder's device variable | where `--vision-device`'s numbers come from |
| --- | --- | --- |
| *(default, CUDA)* | `CUDA_VISIBLE_DEVICES` | `nvidia-smi` |
| `vulkan` | `GGML_VK_VISIBLE_DEVICES` | `vulkaninfo --summary` |
| `sycl` | `ONEAPI_DEVICE_SELECTOR="level_zero:N"` | `sycl-ls` (dGPUs first, iGPUs at the end) |

`tools/vision/CMakeLists.txt` gains the two backend options (one GPU backend at a time), the config keeps
`"backend"` in `"vision"`, and changing the backend rebuilds the encoder. `--vision-device auto` is refused with a
non-CUDA backend: these backends number devices differently than nvidia-smi, so name the device.

**Validation limits:** no Intel iGPU was available to test on. The Vulkan build+pin path was validated on a 2x GTX
1080 Ti rig as a proxy (Vulkan encoder, `GGML_VK_VISIBLE_DEVICES` pinning to card0/card1, VRAM placement via
nvidia-smi); the SYCL path is unit-tested only.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。