Pull requests / #83

#83 serve: llama-server-style timings on responses (prefill/decode rates for llama-swap & friends)

closed · @mikicvi · 0 comentários · No GitHub

Server & API

Descrição

The engine already reports per-request prompt/decode timings, conversation-cache reuse and draft counts on its `DONE` line, but those numbers stayed inside `server.py`. This emits them everywhere the `usage` object already goes, in llama-server's dialect — a top-level `timings` object on the final streamed chunk, the non-streaming response, and the Anthropic `message_delta`:

```json
"timings": {"prompt_n": 10597, "prompt_ms": 22683.0, "prompt_per_second": 467.3,
            "predicted_n": 64, "predicted_ms": 3248.0, "predicted_per_second": 19.7,
            "cache_n": 0, "draft_n": 59, "draft_n_accepted": 32}
```

(`cache_n`/`draft_*` only when the engine reported them.)

Why: proxies and monitors that chart per-request prefill/decode speed read exactly this shape — llama-swap's activity log (the llama-server, vLLM and TabbyAPI dialects it accepts), and anything else that already understands llama.cpp responses. Without it, Strata requests show token counts but `unknown` rates. Nothing else changes: fields are additive, and clients that don't know `timings` ignore it.

Verified end to end through llama-swap v257: streaming and non-streaming `/v1/chat/completions`, and `/v1/messages`, all produce activity rows with prefill/decode rates, cached-token counts on reused prompts, and draft acceptance — e.g. a 10.6k-token request charts PP 467 t/s, decode 19.7 t/s, drafts 32 of 59 accepted.

No site

Links install, modelos, releases.