Pull requests / #1635

#1635 Add a prometheus metrics endpoint with llama.cpp's names

open · @zeevo · 0 コメント · GitHub で見る

Server & API

本文

Replaces #877, which GitHub closed when `main` was force-pushed. Same commit, cherry-picked onto the new `main`.

Adds `GET /metrics/prometheus`: request totals and what is running in Prometheus' text format, under llama-server's metric names (`llamacpp:prompt_tokens_seconds`, `llamacpp:predicted_tokens_seconds`, `llamacpp:requests_processing`, ...), plus Strata's drafts and requests as `strata:*`. Tools that read only llama.cpp's `/metrics` (llm-visuals, as described in #877) work unchanged.

It sits next to #793's vLLM-named output on `GET /metrics` and does not change it or the Monitor's JSON.

Tests: `python -m unittest serve.test_server serve.test_prometheus` (278 OK).

関連リンク

インストール・モデル・リリースへの站内リンク。