Pull requests / #114
#114 serve: report prefix reuse as usage.prompt_tokens_details.cached_tokens
closed · @Yunado · 0 comments · View on GitHub
Description
The engine's per-request prefix reuse count already reaches the web Monitor and the internal DONE line; expose it in the OpenAI usage block under the standard `prompt_tokens_details.cached_tokens` field so any client can show prompt-cache hits. 0 when cold.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.