Pull requests / #278

#278 serve: count_tokens, Anthropic thinking only when asked, STRATA_REQUEST_LINES, a relative exe

closed · @sergqwer · 0 comentários · No GitHub

BenchmarksServer & APIModels & quantsWindows

Descrição

## Summary

Four small server fixes, found by running Claude Code against the server and by starting a config with a relative
engine path:

- **`POST /v1/messages/count_tokens`** answered 404; Claude Code asks it for its context figures. It now renders and
  tokenizes the request exactly as `/v1/messages` would, without running the model.
- **Anthropic thinking only when asked.** A `/v1/messages` request without `thinking` (and without an effort) got the
  template's default, which thinks. Claude Code's helper calls (a ~60-token session title) came back with a thinking
  block and no text. Anthropic's thinking is opt-in, so such a request now renders with `enable_thinking` false. A
  config's `reasoning_effort` still applies; a request with `thinking` or an effort behaves as before.
- **`STRATA_REQUEST_LINES=1`**: a thread follows the engine log and prints one line per request on stdout
  (`[strata] request prompt P cached C output O ttft T ms total S ms prefill X tok/s decode Y tok/s`), for supervisors
  that only see the server's stdout (a tray, llama-swap). Off by default.
- **A relative `exe`** in the config failed in Windows' `CreateProcess` (WinError 2). It is now made absolute
  against the config's `cwd`, as the engine's own path already was.

## Tests

Two new tests: `count_tokens` equals the `input_tokens` that `/v1/messages` reports for the same request, and the
request-line pattern parses the engine's summary. `python -m unittest serve.test_server`: 44 OK.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

No site

Links install, modelos, releases.