Pull requests / #1421

#1421 docs/kubernetes: manifests for one Strata server on a GPU node (kubectl, Kustomize, kapp)

open · @blange48 · 0 コメント · GitHub で見る

BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDADocumentation

本文

## docs/kubernetes: one Strata server on a GPU node

Plain manifests next to `docs/monitoring/`, so the Docker image can run on Kubernetes without anyone rediscovering the
same traps. They run the image built from the `Dockerfile` through `docker-entrypoint.sh`, so every setting is one of
the entrypoint's env vars, exactly as in INSTALL.md's Docker section. They work with `kubectl apply -k`, Kustomize,
Argo CD / Flux, or kapp (`kapp deploy -a strata -f <(kubectl kustomize docs/kubernetes) --wait-timeout 2h`).

| File | What it is |
|---|---|
| `kustomization.yaml` | namespace `strata` and the image name to set |
| `namespace.yaml`, `pvc.yaml` | the namespace and the data volume (model, pack, configs) |
| `deployment.yaml` | one replica, `Recreate`, `runtimeClassName: nvidia`, `IPC_LOCK`, no memory limit, start-up probe and `progressDeadlineSeconds` of 2 h, `terminationGracePeriodSeconds: 90`, a `/dev/shm` emptyDir |
| `service.yaml` | port `http`, label `app: strata`: what `docs/monitoring/servicemonitor.yaml` already selects, and the same `strata-api-key` secret |
| `alerts.yaml` | optional PrometheusRule (not in the kustomization): slow decode, a queue that never empties, first token over a minute, missing batch slots |
| `README.md` | install, kapp, `CONFIG`, several GPUs, memory, stopping, updating, monitoring |

INSTALL.md's Docker section links to it (two lines).

### What the README warns about (each one bit us)

- **The config's `host`.** `CONFIG` (0.1.40.3, thanks for #1244) is the way to run an edited config, but the file's
  own `"host"` must be `"0.0.0.0"`. A config written with `HOST=127.0.0.1` (ours was, behind an nginx sidecar)
  listens on the pod's loopback, so the probes fail and the Service gets "connection refused". `HOST` only applies when
  setup writes a config.
- **The first start's download.** The start-up probe allows 2 h. Without `progressDeadlineSeconds: 7200`,
  `kubectl rollout status` (and Argo / kapp waits) report a first start that is still downloading as failed after
  10 min.
- **A locally imported image.** On a single-node k3s, `docker build` alone leaves the image in Docker's store, and the
  pod sits in `ErrImageNeverPull`. The README gives the `k3s ctr images import` line and the check before deploying.
- **The `Config:` line.** A server that quietly starts the wrong config still answers, only slower (we saw 138
  instead of 390 tok/s at 8 clients). Check the entrypoint's line after every change. `StrataBatchSlotsMissing`
  catches the batch case.
- Page-locked RAM (`IPC_LOCK`, the runtime's `RLIMIT_MEMLOCK`, no cgroup memory limit unless `LOW_RAM=on`), the
  stop that takes longer than 30 s, and `Recreate`, since two servers don't fit on the same GPUs.

### Tested

On k3s v1.36.4+k3s1 (single node, 4× RTX 5080 16 GB, NVIDIA Container Toolkit, `runtimeClassName: nvidia`), image built
from this branch's `main` (0.1.40.3, `CUDA_ARCHITECTURES=120`) and imported with `k3s ctr`. The manifests came from
`kubectl kustomize docs/kubernetes`, with the namespace and names changed for the test, an existing volume (so no
download), `GPUS=0,1,2,3`, 4 GPUs requested, and `CONFIG` pointing to a `--batch 8 --batch-groups 4` config:

- `kubectl apply --dry-run=server`, then `apply`: rolled out, ready in 102 s;
- the log prints `Config: /data/config/...json (from CONFIG)`;
- through the Service: `/health` ok, `/metrics` reports engine 0.1.40.3 and 8 batch slots, a chat completion answers,
  1 / 4 / 8 clients give 123 / 178 / 377 tok/s, and the Prometheus text format serves 124 series;
- `kubectl delete`: the pod stopped cleanly in 8 s, with nothing in the kernel log.

The first run of this test failed on the `host` and the progress deadline above. Both are fixed and documented here.
The first-start download path, kapp itself and the Prometheus Operator objects (`alerts.yaml`; the same rules run
on our cluster, validated against its history) were not part of this run.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

関連リンク

インストール・モデル・リリースへの站内リンク。