Issues / #1058
#1058 0.1.40: a tool call quoted in the thinking is delivered as a real call (fires on a max_tokens cut and inside code fences)
closed · @47Hunter47 · 5 Kommentare · Auf GitHub
Server & APINVIDIA / CUDAModels & quantsLinux
Beschreibung
**Summary.** 0.1.40 (`9fb1b9f`) fixes #804 and on real traffic it recovers more than the opt-in PR did — on a second harness here it turned 11 of 12 recorded strandings into real calls, where 0.1.39 delivered none. But the rescue's only condition is "the called name is one the request declares", so it also acts on a call the model wrote inside its thinking **as an example**. On a reply that then runs into `max_tokens`, that means the client is handed a real tool call for something the model was only documenting — and on a client that auto-approves, it runs.
The example everyone writes quotes a destructive command, so this is worth a look before it spreads:
```text
finish_reason: "length" tools: [{"name": "fake_tool", ...}]
reasoning_content: "... include this example verbatim (it is only documentation, do not run it):
<tool_call>
<function=fake_tool>
<parameter=cmd>
rm -rf build
</parameter>
</function>
</tool_call>
... Then think about multiples ... 700=7*100, 693=7*99, ..."
tool_calls: [{"function": {"name": "fake_tool", "arguments": "{\"cmd\": \"rm -rf build\"}"}}]
```
That is a real response from a live 0.1.40 server, not a parser unit test.
**Environment.** 0.1.40, tag `1735d64`, engine compiled locally on this machine (`BUILD.json`: `source: local`). Linux, 1× RTX 3090 24 GB, 61 GB RAM, driver 580.178.04. Qwen3.8-Flash-Next GSQ-RCO **IQ3_XXS**, `--max-context 230000`, `--kv int8`, images on (CPU encoder), thinking on, sampling temperature 1.0 / top_p 0.95 / top_k 20, `reasoning_budget_tokens` 16384. Requests over `/v1/chat/completions`, non-streaming and streaming (they agree). Nothing was executed: the client read the response fields only, and the tool named in the request exists nowhere.
**Which conditions it fires under.** Same dictated example, one variable changed per row, two attempts each:
| variant | declared tools | name matches | end of turn | fired |
|---|---|---|---|---|
| V1 | yes | yes | `length` (cut mid-thought) | **2/2** |
| V2 | **none** | – | `length` / `stop` | 0/2 |
| V3 | yes, but a **different** name | no | `length` | 0/2 |
| V4 | yes | yes | long reply, natural end | **1/2** (`finish=tool_calls`, and two calls in one reply) |
| V5 | yes | yes | block inside a ```` ```text ```` fence | **1/2** (`finish=tool_calls`) |
So, in order: tools must be declared (V2 confirms the guard the maintainer described), the name must be one of them (V3 confirms the gate), and once those hold the rescue ignores **where** the call sits — not the start of a line, and not kept out by a code fence (V5) — and it does not care **how the turn ended**: a `max_tokens` cut (V1, the dangerous one) and a natural stop (V4) both fire.
Rate, to be fair about it: with the model asked to copy the exact block, ~50-100% of attempts fire; with a free-form prompt ("give an example of the call format, then keep thinking") it was **1 of 11 attempts** — the model often writes the format loosely enough not to parse. It is stochastic, temperature 1.0.
**The same result offline and deterministically.** The corpus on PR #525 (`serve/fixtures/format_fix_specimens.json` at `4111c85`, 25 entries with stated expectations) run against a 0.1.40 tree: the 5 "act" entries are correct, and **10 of the 20 negative controls now deliver a call** — a fenced example, a mid-sentence quote, prose between two blocks (two calls), a call with no parameters, and `cut-max-tokens`, which is exactly the shape above. No model needed, just `OutputParser`:
```python
import sys; sys.path.insert(0, "<0.1.40 tree>")
from serve.frontend import OutputParser
TOOLS = [{"name": "execute_code", "parameters": {"properties": {"code": {"type": "string"}}}}]
INPUT = open("cut-max-tokens.txt").read() # "I could run\n<tool_call>\n<function=execute_code>…"
p = OutputParser(thinking=True, tools=TOOLS, stream_tools=True)
print([(e.kind, getattr(e.call, "name", None)) for e in p.feed(INPUT) + p.finish("length")])
# 0.1.40 -> [('tool_call', 'execute_code')] 0.1.39 -> only reasoning
```
**Mechanism.** In `feed()`, the reasoning branch enters a new `rcall` state on any `<tool_call>` once `self.schemas` is non-empty (`tool = self.buf.find(CALL_START) if self.schemas else -1` — no line or fence context), holds the block whole, and then:
```python
name = body[:end].strip()[len("<function="):].split(">", 1)[0]
if name in self.schemas:
call = parse_tool_call(body[:end], self.schemas.get(name))
# a declared name -> tool_call; anything else -> reasoning
```
The declared-name condition is doing the work the content channel already does — that part is right, and it is why #970's two bogus `<function=tool_call>` deliveries in #804 do not happen here. What is missing is that the content channel distinguishes a call *the model is making* from text *about the format* by position and by where the turn ended, and this path does not.
**Suggested gates, both cheap and both already in #525** (I am not arguing for the flag, only for the two conditions):
1. **The opener must start a line** (or, at the minimum, be refused inside a fence). All 12 live strandings in the corpus from this machine begin with `\n\n<tool_call>\n<function=NAME>` — opener on its own line — and so do all four in #804, so this costs nothing in recall on the recorded data. It removes the fenced documentation case, the mid-sentence quote, and prose-between-blocks.
2. **Refuse when the turn ended by `length`.** A budget cut is when a thinking span is most often open, and a complete call in that reasoning was written as an example far more often than as an act. This is the case with a destructive command in it.
**One request: make it observable.** `rcall` logs nothing and adds nothing to `GET /metrics` (`totals` still holds only the request/token/draft keys). So an operator cannot tell whether this class is firing on 1 reply in 100 or 1 in 100 000, and cannot tell when it acts on the wrong thing — the parse-level shape that #525 calls a hit is invisible here. A single counter (`totals.tool_calls_from_reasoning` or similar) would make the rate measurable per model and per quant, which is the question every harness will ask before turning this on.
**Corpora, both public.** PR #525's `serve/fixtures/format_fix_specimens.json` (25 entries, expectations included) is the deterministic half. Mine — 12 live strandings from this machine, genericised, in that same schema, 11/12 recovered by 0.1.40: https://gist.github.com/47Hunter47/4abcd3a80701f5634f85b96e4c067cc3 . Happy to re-run either against any head or to add the variants above as corpus entries.
Mehr auf der Site
Links zu Install, Modellen, Releases.