Pull requests / #970

#970 serve: a tool call written inside the thinking ends it implicitly (#804)

closed · @talisp · 0 评论 · 在 GitHub 查看

Setup & installServer & APIModels & quantsLinux

描述

Fixes #804.

## What

Qwen3.8-Flash-Next sometimes writes a complete tool call **inside the thinking, without `</think>` first** (4 of ~1,050 Hermes Agent replies here). `OutputParser` only left the `"reasoning"` state on `</think>`, so the call ended up in `reasoning_content`: no `tool_calls`, `finish_reason=stop`, and the agent stopped.

Like vLLM's Qwen3 parser (PR vllm-project/vllm#35687, where `<tool_call>` in the reasoning is an implicit end of it), a generated call now ends the thinking. Unlike vLLM, which fires on any `<tool_call>`, it fires only on a real call, so quoted markup stays reasoning. All three must hold:

1. `<tool_call>` at the **start of a line** (right after `\n`, or as the very first text of the thinking);
2. **outside a ``` or ~~~ code block** of the reasoning;
3. followed, after whitespace only, by **`<function=`** (the same rule #700 uses in the content state).

The text before it stays `reasoning_content`, and the call goes through the normal `"call"` state. While the text after the tag is still whitespace or a prefix of `<function=`, it is held back, so streamed and whole outputs are identical (with and without `stream_tools`). An implicit call that never completes is returned as reasoning, as before.

Each one is logged (`[strata] implicit end of thinking: ...`) and counted in `GET /metrics` (`totals.implicit_reasoning_ends`, plus a per-request field), so it can be watched per model and quant.

Complementary to #700 / #966: those handle drift after `</think>` (content state); this handles the call that never left the reasoning state.

## Second commit (defense in depth, template)

When the history is rendered, `serve/chat_template.jinja` closes `</think>` before such a call left in an earlier assistant turn. This applies when `content` starts with an open `<think>` and has a line-start `<tool_call>\n<function=`, or when there are no `tool_calls` and `reasoning_content` ends with such a call. The model is then never shown a call inside `<think>` as an example. All 10 `chat_golden.json` cases render byte-identical (now checked by a test). Note: the server reads the template from the pack's `tokenizer/`, so this part reaches users only when packs are rebuilt. It can be dropped from this PR if you prefer.

## Tests

`serve/test_reasoning_toolcall.py`, 12 tests:

- **The 4 recorded failures** (`serve/fixtures/`): each gives exactly 1 `tool_call`, with the reasoning before it kept. Checked fed 1 character at a time, in random chunks and whole, with `stream_tools` on and off. All 4 give 0 calls on `main`.
- **10 mention-only controls**, none of which give a call:
  - backticks;
  - a whole call written mid-sentence;
  - backticks at line start;
  - a line-start tag with no `<function=`;
  - another tag after it;
  - indented;
  - a `>` quote;
  - ``` and ~~~ blocks;
  - a bare tag before `</think>`.

  They also stay correct when followed by `</think>` and a real call.
- **Server-level:** `finish_reason=tool_calls`, the log line, and the `/metrics` counter.
- **Template:** `chat_golden.json` matches 10/10, and both history repairs are covered.

Full `serve/test_*.py` on this branch: 280 run, 9 skipped. 1 error, `test_responses.OverHttp.test_json_schema_text_format`, fails the same way on plain `main` when `jsonschema` is not installed (the other jsonschema tests skip in that case; this one does not).

## In production

This has been running here since 2026-10-04 (Linux, L40S vGPU, Hermes Agent over the OpenAI API): first on IQ3_S / 0.1.38, and since today on UD-IQ4_XS / 0.1.39 with images on. It recovered 2 real cases that previously stopped the agent, with no regressions seen.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。