Pull requests / #107
#107 serve: report cached prompt tokens, llama.cpp timings and GET /v1/status
closed · merged 2026-09-29 · @architectds · 0 Kommentare · Auf GitHub
Server & APIModels & quantsLinux
Beschreibung
The conversation cache reuses prompt prefixes, and `/metrics` already records the reuse per request. But the chat APIs don't report it, so clients show a cached prompt as 0% cached. Front-ends that read speed from llama.cpp's `timings` show nothing. This PR adds the missing fields and changes nothing else in the replies: - **OpenAI** (`/v1/chat/completions`): `usage.prompt_tokens_details.cached_tokens` is the part of the prompt the conversation cache held. It's on the last streamed chunk and on a whole reply. - **Anthropic** (`/v1/messages`): the final `usage` splits `input_tokens` from `cache_read_input_tokens`, the way Anthropic's API counts them. That's the stream's `message_delta` and a whole reply. `message_start` still carries the whole prompt, the only count known before the engine reads it. - **llama.cpp's `timings`** go on the last chunk and on a whole reply: `prompt_n`, `cache_n`, `prompt_ms`, `prompt_per_second`, `predicted_n`, `predicted_ms`, `predicted_per_second` and the per-token ms. They come from the engine's own clock, the same numbers `/metrics` shows. - **`GET /v1/status`** is a small read-only summary for front-ends that poll their server rather than guess. It's behind the API key when one is set. It gives: - the model and its context window; - images, and the APIs served; - whether a request is running; - the last request's timings; - the GPU and RAM, from the telemetry. The Monitor's `/metrics` is unchanged. An engine without a clock (the mock engine) reports no `timings` and 0 cached tokens. Tested: - `python -m unittest serve.test_server serve.test_mcp` passes. The new `UsageAndStatus` class covers both APIs, streaming, `/v1/status` and the no-clock case. - Live on an A100-40G on Linux (Colab, a 0.1.20 build with the same change): - a second turn reported `cached_tokens: 2784`, with `timings` `cache_n: 2784, prompt_n: 20`; - a front-end read the new `/v1/status` for its speed panel and for image upload.
Mehr auf der Site
Links zu Install, Modellen, Releases.