Pull requests / #107

#107 serve: report cached prompt tokens, llama.cpp timings and GET /v1/status

closed · merged 2026-09-29 · @architectds · 0 コメント · GitHub で見る

Server & APIModels & quantsLinux

本文

The conversation cache reuses prompt prefixes, and `/metrics` already records the reuse per request. But the chat APIs don't report it, so clients show a cached prompt as 0% cached. Front-ends that read speed from llama.cpp's `timings` show nothing.

This PR adds the missing fields and changes nothing else in the replies:

- **OpenAI** (`/v1/chat/completions`): `usage.prompt_tokens_details.cached_tokens` is the part of the prompt the conversation cache held. It's on the last streamed chunk and on a whole reply.
- **Anthropic** (`/v1/messages`): the final `usage` splits `input_tokens` from `cache_read_input_tokens`, the way Anthropic's API counts them. That's the stream's `message_delta` and a whole reply. `message_start` still carries the whole prompt, the only count known before the engine reads it.
- **llama.cpp's `timings`** go on the last chunk and on a whole reply: `prompt_n`, `cache_n`, `prompt_ms`, `prompt_per_second`, `predicted_n`, `predicted_ms`, `predicted_per_second` and the per-token ms. They come from the engine's own clock, the same numbers `/metrics` shows.
- **`GET /v1/status`** is a small read-only summary for front-ends that poll their server rather than guess. It's behind the API key when one is set. It gives:
  - the model and its context window;
  - images, and the APIs served;
  - whether a request is running;
  - the last request's timings;
  - the GPU and RAM, from the telemetry.

  The Monitor's `/metrics` is unchanged.

An engine without a clock (the mock engine) reports no `timings` and 0 cached tokens.

Tested:
- `python -m unittest serve.test_server serve.test_mcp` passes. The new `UsageAndStatus` class covers both APIs, streaming, `/v1/status` and the no-clock case.
- Live on an A100-40G on Linux (Colab, a 0.1.20 build with the same change):
  - a second turn reported `cached_tokens: 2784`, with `timings` `cache_n: 2784, prompt_n: 20`;
  - a front-end read the new `/v1/status` for its speed panel and for image upload.

関連リンク

インストール・モデル・リリースへの站内リンク。