Pull requests / #278
#278 serve: count_tokens, Anthropic thinking only when asked, STRATA_REQUEST_LINES, a relative exe
closed · @sergqwer · 0 comments · View on GitHub
BenchmarksServer & APIModels & quantsWindows
Description
## Summary Four small server fixes, found by running Claude Code against the server and by starting a config with a relative engine path: - **`POST /v1/messages/count_tokens`** answered 404; Claude Code asks it for its context figures. It now renders and tokenizes the request exactly as `/v1/messages` would, without running the model. - **Anthropic thinking only when asked.** A `/v1/messages` request without `thinking` (and without an effort) got the template's default, which thinks. Claude Code's helper calls (a ~60-token session title) came back with a thinking block and no text. Anthropic's thinking is opt-in, so such a request now renders with `enable_thinking` false. A config's `reasoning_effort` still applies; a request with `thinking` or an effort behaves as before. - **`STRATA_REQUEST_LINES=1`**: a thread follows the engine log and prints one line per request on stdout (`[strata] request prompt P cached C output O ttft T ms total S ms prefill X tok/s decode Y tok/s`), for supervisors that only see the server's stdout (a tray, llama-swap). Off by default. - **A relative `exe`** in the config failed in Windows' `CreateProcess` (WinError 2). It is now made absolute against the config's `cwd`, as the engine's own path already was. ## Tests Two new tests: `count_tokens` equals the `input_tokens` that `/v1/messages` reports for the same request, and the request-line pattern parses the engine's summary. `python -m unittest serve.test_server`: 44 OK. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.