Pull requests / #1635

#1635 Add a prometheus metrics endpoint with llama.cpp's names

open · @zeevo · 0 comments · View on GitHub

Server & API

Description

Replaces #877, which GitHub closed when `main` was force-pushed. Same commit, cherry-picked onto the new `main`.

Adds `GET /metrics/prometheus`: request totals and what is running in Prometheus' text format, under llama-server's metric names (`llamacpp:prompt_tokens_seconds`, `llamacpp:predicted_tokens_seconds`, `llamacpp:requests_processing`, ...), plus Strata's drafts and requests as `strata:*`. Tools that read only llama.cpp's `/metrics` (llm-visuals, as described in #877) work unchanged.

It sits next to #793's vLLM-named output on `GET /metrics` and does not change it or the Monitor's JSON.

Tests: `python -m unittest serve.test_server serve.test_prometheus` (278 OK).

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.