Issues / #804

#804 Tool call written inside `<think>` (no `</think>`) is returned as `reasoning_content`: no `tool_calls`, `finish_reason=stop`

Tool call written inside `<think>` (no `</think>`) is returned as `reasoning_con

closed · @talisp · 10 Kommentare · Auf GitHub

Server & APIModels & quantsWindows

Beschreibung

## Summary

Qwen3-family models (seen with **Qwen3.8-Flash-Next, GSQ-RCO IQ3_S**) sometimes emit a complete, well-formed `<tool_call>` **inside the thinking block without closing `</think>` first**. `serve/frontend.py` → `OutputParser` only leaves the `"reasoning"` state on `</think>`, so the whole call ends up in `reasoning_content`. The client (an agent harness) receives `tool_calls: []`, empty `content` and `finish_reason: "stop"`, and the agent loop silently stops: the call is never executed.

- Frequency observed: **4 of ~1050** responses in real agent use (Hermes Agent, OpenAI Chat Completions API, streaming).
- Engine 0.1.38, `serve/server.py`, thinking mode on, official Qwen thinking sampling (temperature 1.0, top_p 0.95, top_k 20, presence 0). Not a sampling issue.

## What the model generated (shape of all 4 cases)

```
...the reasoning text, then the model decides to act.

<tool_call>
<function=execute_code>
<parameter=code>
print("hello")
</parameter>
</function>
</tool_call><|im_end|>
```

No `</think>` anywhere. The call itself is complete and parseable.

## Minimal reproduction (no GPU)

```python
from serve.frontend import OutputParser

text = ("I will run it now.\n\n<tool_call>\n<function=execute_code>\n"
        "<parameter=code>\nprint(1)\n</parameter>\n</function>\n</tool_call>")
p = OutputParser(thinking=True, tools=None, stream_tools=True)
evs = p.feed(text) + p.finish()
print([e.kind for e in evs])
# current main: ['reasoning']: no 'tool_call' event
```

## Root cause

In `OutputParser.feed()`, the `"reasoning"` branch searches only for `THINK_END` (`</think>`). `CALL_START` is only looked for in the `"content"` state. A call generated before `</think>` is never seen as a call.

## Prior art: vLLM

- vLLM **PR #35687** ("[Bugfix] Treat `<tool_call>` as implicit reasoning end in Qwen3 parser", merged 2026-04-24) fixed the same bug. In current vLLM (`vllm/parser/qwen3.py`, the grammar in `qwen3_config()`), the transition `(REASONING, "TOOL_START") -> TOOL_PREAMBLE` emits `REASONING_END` + `TOOL_CALL_START`, so **any** `<tool_call>` ends the reasoning, unconditionally.
- vLLM **PR #59821** ("Preserve quoted Qwen3 tool markup in reasoning") is still **open**. It addresses the false-positive side: a model that only *mentions* `<tool_call>` in its reasoning (quoting the format, explaining it) should not trigger a call.
- vLLM's rule that ignores paired `<tool_call>` tokens applies to **prompt** tokens, not generated output, so it is not relevant here.

## Proposed fix (implemented and tested locally, ready as a PR)

Adopt vLLM's idea (an implicit end of reasoning), but with a **stricter trigger** so quoted markup stays reasoning. In the `"reasoning"` state, a generated `<tool_call>` ends the thinking only when **all three** hold:

1. it is at the **start of a line** (right after `\n`, or the very first text of the thinking);
2. it is **outside a fenced code block** (```` ``` ```` or `~~~`) of the reasoning;
3. it is followed, after only whitespace, by **`<function=`** (the start of a real call body).

The text before it stays `reasoning_content`, and the call goes through the normal `"call"` state, so it is extracted in streaming (`stream_tools`) and non-streaming paths alike. While the text after `<tool_call>` is still only whitespace or a prefix of `<function=`, it is held back (the same approach as for partial tags), so streamed and whole outputs are identical.

**Why stricter than vLLM:** an unconditional rule turns things like "I must emit `` `<tool_call>` `` next", "The format is <tool_call>…" or an indented or `>`-quoted example into a bogus call. All real failures had the exact shape `\n\n<tool_call>\n<function=NAME>`, so the three conditions cost nothing in recall.

**Edge case:** if an implicit call never completes (output truncated), it is returned as reasoning, as before, not as content.

**Observability:** each implicit end is logged (`[strata] implicit end of thinking: N tool call(s) written inside the thinking without </think> were read as calls`) and counted in `GET /metrics` (`totals.implicit_reasoning_ends` plus a per-request field), so the rate can be tracked per model and quant.

## Defense in depth: chat template

In `serve/chat_template.jinja`, when an earlier assistant message has a call left inside the thinking, `</think>` is now closed before the call when the history is rendered. This applies when:

- `content` starts with an open `<think>` and contains a line-start `<tool_call>\n<function=`, or
- there are no `tool_calls` and `reasoning_content` ends with such a call.

This way the model is never shown a call inside `<think>` as an example to imitate. All 10 cases in `chat_golden.json` render byte-identical. (Reference: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates.)

**Note:** the server loads the template from the pack (`<pack>/tokenizer/chat_template.jinja`), not from `serve/`. A template fix must also reach the pack, or `strata_pack` must copy it.

## Tests

New `serve/test_reasoning_toolcall.py` (12 tests):

- **The 4 real failures** (sanitized fixture): each yields exactly 1 `tool_call`, with the reasoning preserved. Checked when fed 1 char at a time, in random chunk sizes, and whole, with `stream_tools` on and off. All 4 yield 0 calls on current `main`.
- **10 synthetic mention-only controls**, none of which yield a call:
  - backticks mid-sentence;
  - a whole call written mid-sentence;
  - backticks at line start;
  - a line-start tag with no `<function=`;
  - another tag after it;
  - indented;
  - a markdown `>` quote;
  - inside ```` ``` ```` and `~~~` code blocks;
  - a bare tag before `</think>`.

  They also stay correct when followed by `</think>` and a real call.
- **Server-level:** `finish_reason=tool_calls`, the log line, and the `/metrics` counter.
- **Template:** `chat_golden.json` matches 10/10, and both history repairs are covered.
- **Full `serve/test_*.py` suite:** 209 pass, 9 skipped (environment: Windows-only, no jsonschema, no pack tokenizer).

## Side finding: agent harnesses and `preserve_thinking`

Hermes Agent strips `reasoning_content` from the history for generic OpenAI-compatible endpoints. It keeps it only for DeepSeek, Kimi and MiMo, or with `model.reasoning_echo: true` in its config (source: `agent/message_sanitization.py`). `/metrics` agrees: after a 9,977-token reply, the next prompt grew by only 64 tokens. With the default `preserve_thinking=true`, past assistant turns therefore render as empty `<think>\n\n</think>` blocks, and prefix-cache reuse stops at the last assistant turn. This is not a bug in Strata, but worth documenting for users of agent harnesses.

## Keywords (for search)

tool_call inside think, missing `</think>`, reasoning_content contains tool_call, empty tool_calls, finish_reason stop instead of tool_calls, agent stops after reasoning, Qwen3 implicit reasoning end, OutputParser, Qwen3.8-Flash-Next, IQ3_S, Hermes Agent, vLLM #35687.


<img width="2533" height="1505" alt="Image" src="https://github.com/user-attachments/assets/07c4b75a-5cd8-44c1-ab27-7d326664788a" />

Mehr auf der Site

Links zu Install, Modellen, Releases.