Pull requests / #1068

#1068 serve: a tool call from the reasoning only when it ends the turn (#1058)

closed · @sergqwer · 0 Kommentare · Auf GitHub

Server & API

Beschreibung

Fixes #1058. On 0.1.40 the #804 rescue delivers any complete call to a declared tool that the model wrote inside its thinking. That includes a call the model only quoted there, such as an example, a fenced block or a turn cut by `max_tokens` mid-thought. The issue shows a quoted `rm -rf build` returned as a `tool_call` with finish `length`.

### What changes

A call inside the reasoning is held until the turn ends. It is delivered only if all three of these hold:

1. its `<tool_call>` begins a line (the issue's first gate);
2. after it come only whitespace and further such calls;
3. the turn ends by itself (a stop token or a stop string), not by `max_tokens`, a cancel or an error (the issue's second gate).

Otherwise its text goes back to the reasoning unchanged. A call in the answer is unchanged, and so is everything without declared tools.

- **Gate 2** (whitespace only after the call) is the one the issue did not list. A stranding is a call the model wrote and then stopped, so nothing follows it. A quoted example is followed by more thinking or prose, a closing fence, or `</think>` and an answer, and this gate catches those. Without it, the fenced examples still fire, because the opener after "```xml\n" begins a line. So do `mention-trailing-prose` and the prose-between-two-blocks entry of #525's corpus.
- **When the call is sent.** It goes out when the turn ends instead of as soon as it closes. Nothing comes after a delivered call anyway, so a client gets the same call at the same point in the stream, only no earlier than the end of the turn.
- **Observability** (the issue's ask). `GET /metrics` totals get `reasoning_calls_delivered` and `reasoning_calls_kept_as_text`, which counts complete calls to a declared tool that stayed text.

### Measured

Both corpora from the issue were run offline through `OutputParser`. Each was fed whole, one character at a time and in random 1-9 character chunks, and every chunking gave the same result. A natural end is `finish == "stop"`.

| | 0.1.40 | this PR |
|---|---|---|
| #525's 25 specimens (5 acts, 20 negative controls) | 10 negatives deliver a call | all 25 as expected |
| the 12 live strandings (gist from the issue) | `live-02` (finish `length`) and `live-06` (an envelope, `todo_list`) deliver a call | all 12 as expected |
| live strandings that are calls (declared name, natural stop) | 9 of 9 delivered | 9 of 9 delivered |

The recall on the recorded live data is unchanged. Every live stranding opens its call on its own line and ends the turn with it.

### Tests

`serve/test_reasoning_tools.py` now expects these semantics:

- a call that ends the turn (one, two adjacent, and at the start of the reasoning) is delivered and never streamed piece by piece;
- the quoted shapes from #1058 stay text, with the text unchanged:
  - more thought after the call;
  - mid-sentence;
  - prose after it;
  - a fenced block;
  - `</think>` after it;
  - a `max_tokens` cut;
  - prose between two calls;
  - a second call mid-line;
- the counters;
- over HTTP, the same call delivered on a stop and kept as text when `max_tokens` cuts the turn right after it. This test fails if `finish()` ignores how the turn ended.

The serve suite passes (423 tests).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Mehr auf der Site

Links zu Install, Modellen, Releases.