Pull requests / #939

#939 serve: a tool call written inside an unclosed thinking span is an implicit think end (#804)

closed · @muriloducatti · 0 comments · View on GitHub

Server & APIModels & quants

Description

## Problem

Qwen3-family models sometimes go straight from their reasoning to a tool call without ever emitting `</think>`. `OutputParser` only leaves the `reasoning` state on `</think>`, so the whole call is channeled out as `reasoning_content`: the client sees no `tool_calls`, `finish_reason: "stop"`, and its agent loop stops silently with nothing to run. Reported live in #804 (Qwen3.8-Flash-Next, GSQ-RCO IQ3_S; ~4 of 1050 responses in agent use).

## Fix

A `<tool_call>` that ends an unclosed thinking span is an implicit end of thinking, with a strict trigger so quoted markup stays reasoning. In the `reasoning` state, it must be:

1. at the start of a line (right after a newline, or the very first text of the thinking),
2. outside a ``` / ~~~ fenced code block,
3. followed - after only whitespace - by the `<function=` that opens a real call body.

The text before it stays `reasoning_content`, and the call goes through the normal `call` state, so it is extracted in streaming (`stream_tools`) and non-streaming paths alike. A candidate whose follower has not arrived is held, so streamed and whole outputs are identical. This follows the proposal in #804 (and vLLM PR #35687's idea, with the stricter trigger).

`self.buf` keeps its original trimming - the reasoning budget in `server.py` relies on `not parser.buf` - and the reasoning span is tracked separately in `self.rseen`.

Each implicit end is logged: `[strata] implicit end of thinking: N tool call(s) ...`.

## Tests

New `ToolCallInsideThinking` in `serve/test_server.py`:

- the live shape, fed 1 char at a time, in chunks of 7, and whole, with `stream_tools` on and off: exactly 1 call, the reasoning kept;
- `</think>` before the opener is unchanged;
- 10 mention-only controls that must stay reasoning (backticked, mid-sentence, indented, markdown quote, ``` and ~~~ fences, no function body, tag after tag, bare tag before `</think>`);
- a call the output ends inside is not reported;
- server-level: `finish_reason: "tool_calls"` and the arguments.

`python -m unittest serve.test_server`: 144 tests OK (Python 3.13; `jsonschema` and the pack tokenizer absent).

## Note

The chat template also renders a stranded call from the history in the #804 proposal, but the server loads the template from the pack (`<pack>/tokenizer/chat_template.jinja`), not `serve/`, so that half needs the pack side and is left out here.

Fixes #804.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.