Pull requests / #939
#939 serve: a tool call written inside an unclosed thinking span is an implicit think end (#804)
closed · @muriloducatti · 0 comentários · No GitHub
Descrição
## Problem Qwen3-family models sometimes go straight from their reasoning to a tool call without ever emitting `</think>`. `OutputParser` only leaves the `reasoning` state on `</think>`, so the whole call is channeled out as `reasoning_content`: the client sees no `tool_calls`, `finish_reason: "stop"`, and its agent loop stops silently with nothing to run. Reported live in #804 (Qwen3.8-Flash-Next, GSQ-RCO IQ3_S; ~4 of 1050 responses in agent use). ## Fix A `<tool_call>` that ends an unclosed thinking span is an implicit end of thinking, with a strict trigger so quoted markup stays reasoning. In the `reasoning` state, it must be: 1. at the start of a line (right after a newline, or the very first text of the thinking), 2. outside a ``` / ~~~ fenced code block, 3. followed - after only whitespace - by the `<function=` that opens a real call body. The text before it stays `reasoning_content`, and the call goes through the normal `call` state, so it is extracted in streaming (`stream_tools`) and non-streaming paths alike. A candidate whose follower has not arrived is held, so streamed and whole outputs are identical. This follows the proposal in #804 (and vLLM PR #35687's idea, with the stricter trigger). `self.buf` keeps its original trimming - the reasoning budget in `server.py` relies on `not parser.buf` - and the reasoning span is tracked separately in `self.rseen`. Each implicit end is logged: `[strata] implicit end of thinking: N tool call(s) ...`. ## Tests New `ToolCallInsideThinking` in `serve/test_server.py`: - the live shape, fed 1 char at a time, in chunks of 7, and whole, with `stream_tools` on and off: exactly 1 call, the reasoning kept; - `</think>` before the opener is unchanged; - 10 mention-only controls that must stay reasoning (backticked, mid-sentence, indented, markdown quote, ``` and ~~~ fences, no function body, tag after tag, bare tag before `</think>`); - a call the output ends inside is not reported; - server-level: `finish_reason: "tool_calls"` and the arguments. `python -m unittest serve.test_server`: 144 tests OK (Python 3.13; `jsonschema` and the pack tokenizer absent). ## Note The chat template also renders a stranded call from the history in the #804 proposal, but the server loads the template from the pack (`<pack>/tokenizer/chat_template.jinja`), not `serve/`, so that half needs the pack side and is left out here. Fixes #804.
No site
Links install, modelos, releases.