Pull requests / #1132

#1132 docs: running Strata behind llama-swap

closed · @christopherrobertbrooks-tech · 0 评论 · 在 GitHub 查看

Setup & installServer & APIModels & quantsDocumentation

描述

Reopened from #829, which GitHub closed when `main` was force-pushed; the page is unchanged. Rechecked against 0.1.40.1: `STRATA_ALLOWED_HOSTS` / `allowed_hosts`, `/health` and `--port` work as the page describes.

A new page, `docs/LLAMA_SWAP.md`: how to run Strata as one of the models behind [llama-swap](https://github.com/mostlygeek/llama-swap). Docs only.

- **The entry:** call `serve/server.py` with llama-swap's `${PORT}` (the generated `run-<model>.sh` fixes 8080), `checkEndpoint: /health`, and a `healthCheckTimeout` that covers Strata's load time.
- **The trap:** the first request through llama-swap fails with `403 Host '...' is not allowed (DNS rebinding protection)` when clients reach the proxy by host name; it works by IP or localhost, so it can pass a local test and fail from another machine. Three fixes: `STRATA_ALLOWED_HOSTS`, `allowed_hosts` in the config, or an API key.
- **Stopping:** llama-swap's SIGTERM stops the server and its engine; the GPU was free within 2 s in our test.
- **`/upstream/<model>/...`** reaches Strata's `/health` and `/v1/status` through the proxy (and loads the model).
- A short, labelled note that Claude Code sends `output_config.effort: "high"` by default and `CLAUDE_CODE_EFFORT_LEVEL` changes it (one measured task, one run per level).

Tested on 2026-10-04 with Strata 0.1.39, Coder IQ1_M on a Tesla V100, clients using the machine's host name: the exact entry on the page loaded the model, answered through the host name, and freed the GPU on unload.

The page was drafted with an AI assistant (Claude) from our setup and tests, and checked by me.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。