Pull requests / #114

#114 serve: report prefix reuse as usage.prompt_tokens_details.cached_tokens

closed · @Yunado · 0 评论 · 在 GitHub 查看

Server & API

描述

The engine's per-request prefix reuse count already reaches the web Monitor and the internal DONE line; expose it in the OpenAI usage block under the standard `prompt_tokens_details.cached_tokens` field so any client can show prompt-cache hits. 0 when cold.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。