Pull requests / #970
#970 serve: a tool call written inside the thinking ends it implicitly (#804)
closed · @talisp · 0 Kommentare · Auf GitHub
Setup & installServer & APIModels & quantsLinux
Beschreibung
Fixes #804. ## What Qwen3.8-Flash-Next sometimes writes a complete tool call **inside the thinking, without `</think>` first** (4 of ~1,050 Hermes Agent replies here). `OutputParser` only left the `"reasoning"` state on `</think>`, so the call ended up in `reasoning_content`: no `tool_calls`, `finish_reason=stop`, and the agent stopped. Like vLLM's Qwen3 parser (PR vllm-project/vllm#35687, where `<tool_call>` in the reasoning is an implicit end of it), a generated call now ends the thinking. Unlike vLLM, which fires on any `<tool_call>`, it fires only on a real call, so quoted markup stays reasoning. All three must hold: 1. `<tool_call>` at the **start of a line** (right after `\n`, or as the very first text of the thinking); 2. **outside a ``` or ~~~ code block** of the reasoning; 3. followed, after whitespace only, by **`<function=`** (the same rule #700 uses in the content state). The text before it stays `reasoning_content`, and the call goes through the normal `"call"` state. While the text after the tag is still whitespace or a prefix of `<function=`, it is held back, so streamed and whole outputs are identical (with and without `stream_tools`). An implicit call that never completes is returned as reasoning, as before. Each one is logged (`[strata] implicit end of thinking: ...`) and counted in `GET /metrics` (`totals.implicit_reasoning_ends`, plus a per-request field), so it can be watched per model and quant. Complementary to #700 / #966: those handle drift after `</think>` (content state); this handles the call that never left the reasoning state. ## Second commit (defense in depth, template) When the history is rendered, `serve/chat_template.jinja` closes `</think>` before such a call left in an earlier assistant turn. This applies when `content` starts with an open `<think>` and has a line-start `<tool_call>\n<function=`, or when there are no `tool_calls` and `reasoning_content` ends with such a call. The model is then never shown a call inside `<think>` as an example. All 10 `chat_golden.json` cases render byte-identical (now checked by a test). Note: the server reads the template from the pack's `tokenizer/`, so this part reaches users only when packs are rebuilt. It can be dropped from this PR if you prefer. ## Tests `serve/test_reasoning_toolcall.py`, 12 tests: - **The 4 recorded failures** (`serve/fixtures/`): each gives exactly 1 `tool_call`, with the reasoning before it kept. Checked fed 1 character at a time, in random chunks and whole, with `stream_tools` on and off. All 4 give 0 calls on `main`. - **10 mention-only controls**, none of which give a call: - backticks; - a whole call written mid-sentence; - backticks at line start; - a line-start tag with no `<function=`; - another tag after it; - indented; - a `>` quote; - ``` and ~~~ blocks; - a bare tag before `</think>`. They also stay correct when followed by `</think>` and a real call. - **Server-level:** `finish_reason=tool_calls`, the log line, and the `/metrics` counter. - **Template:** `chat_golden.json` matches 10/10, and both history repairs are covered. Full `serve/test_*.py` on this branch: 280 run, 9 skipped. 1 error, `test_responses.OverHttp.test_json_schema_text_format`, fails the same way on plain `main` when `jsonschema` is not installed (the other jsonschema tests skip in that case; this one does not). ## In production This has been running here since 2026-10-04 (Linux, L40S vGPU, Hermes Agent over the OpenAI API): first on IQ3_S / 0.1.38, and since today on UD-IQ4_XS / 0.1.39 with images on. It recovered 2 real cases that previously stopped the agent, with no regressions seen. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Mehr auf der Site
Links zu Install, Modellen, Releases.