Pull requests / #829

#829 docs: running Strata behind llama-swap

closed · @christopherrobertbrooks-tech · 0 comentários · No GitHub

Setup & installServer & APIModels & quantsDocumentation

Descrição

A new page, `docs/LLAMA_SWAP.md`: how to run Strata as one of the models behind [llama-swap](https://github.com/mostlygeek/llama-swap). Docs only.

- **The entry:** call `serve/server.py` with llama-swap's `${PORT}` (the generated `run-<model>.sh` fixes 8080), `checkEndpoint: /health`, and a `healthCheckTimeout` that covers Strata's load time.
- **The trap:** the first request through llama-swap fails with `403 Host '...' is not allowed (DNS rebinding protection)` when clients reach the proxy by host name; it works by IP or localhost, so it can pass a local test and fail from another machine. Three fixes: `STRATA_ALLOWED_HOSTS`, `allowed_hosts` in the config, or an API key.
- **Stopping:** llama-swap's SIGTERM stops the server and its engine; the GPU was free within 2 s in our test.
- **`/upstream/<model>/...`** reaches Strata's `/health` and `/v1/status` through the proxy (and loads the model).
- A short, labelled note that Claude Code sends `output_config.effort: "high"` by default and `CLAUDE_CODE_EFFORT_LEVEL` changes it (one measured task, one run per level).

Tested today with Strata 0.1.39, Coder IQ1_M on a Tesla V100, clients using the machine's host name: the exact entry on the page loaded the model, answered through the host name, and freed the GPU on unload.

The page was drafted with an AI assistant (Claude) from our setup and tests, and checked by me.

No site

Links install, modelos, releases.