Pull requests / #615
#615 fix(serve): correct cache usage and timings across reasoning continuations
closed · @liowald · 0 Kommentare · Auf GitHub
Setup & installServer & APIModels & quants
Beschreibung
When `reasoning_budget_tokens` starts a continuation, the response combines the original input count with the continuation's cache count. In a captured OMP 18.4.10 request, Strata reported **181 input tokens and 244 cached tokens**. OMP consequently calculated 968 total tokens from an API response whose total was 905. Anthropic responses also misclassify original input as cached, though their clamp prevents an out-of-range count. The fix captures each native `DONE` after closing and draining its generation. It preserves the first segment's input/cache counts and sums completed segments' timing, generated-token, draft, and I/O counters. A continuation that ends without a new `DONE` cannot count the preceding segment twice. Input/cache usage describes the original request; `prompt_ms` includes continuation prefill. Their ratio therefore does not measure native prefill throughput. Completion counts retain inserted wrap-up tokens, while decode rates use native generated tokens. Validation: - **133 tests pass:** `python -m unittest serve.test_server serve.test_lifecycle serve.test_request_accounting`. - Seven CPU-only HTTP tests cover OpenAI and Anthropic usage in both response modes, timings, metrics, cancellation during a continuation, and failure without a new `DONE`. Five regression cases fail on unmodified `99f3dbd`; both controls pass before and after. - The same server patch passed eight installed OMP checks with full Flash-Next IQ3_S, including three reasoning-budget probes comparing captured streaming usage with the client's parsed totals. Related to #484, which preserves first-pass reuse in the monitor's history row. This patch fixes API usage and whole-request counters without depending on its additional monitor fields. Implementation and independent code review were AI-assisted. Tests and application checks ran locally.
Mehr auf der Site
Links zu Install, Modellen, Releases.