Pull requests / #1400

#1400 docs: running Strata behind Open WebUI

open · @fioruccione · 0 评论 · 在 GitHub 查看

Multi-GPUNVIDIA / CUDAModels & quantsDocumentation

描述

Adds `docs/OPEN_WEBUI.md`, a guide like `LLAMA_SWAP.md`: what we needed to run Strata behind Open WebUI for a small office. Docs only.

Tested with Open WebUI 0.11.4 (venv, same machine) and Strata 0.1.40 (Unsloth UD-IQ4_XS, 262K context, CPU vision encoder) on 2x RTX 4060 Ti 16 GB in a layer split with a Threadripper PRO 3975WX.

- **How many people can share it.** One slot, worked out from *your* decode and prompt rates: waits grow in proportion to how much slower your machine is. The worked examples come from the community reports (Haswell + DDR3, ours, 2x RX 6900 XT, RTX 5090). Also what `"parallel"` and `--batch-mtp` change today.
- **The connection.** URL and key, a model on top of Strata's for prompt, tools and settings, sharing it with users, and a pointer to the Host/Origin checks for remote or container setups. We did not test those.
- **Four traps we hit.**
  - Capabilities are all on by default. Without the vision encoder, one pasted picture makes every following message in that chat fail with a 400.
  - The task model (titles, tags) queues behind the answers on one slot.
  - Compaction and attachment limits have to be sized to Strata's context.
  - `reasoning_effort` is not sent, so Strata's Chat settings apply.

Every number in it was measured on this machine or comes from a linked community report. Untested setups are labelled as such.

Measured on my machine, with an AI coding assistant helping on the write-up.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。