Pull requests / #966
#966 serve: recover tool calls the model writes next to the template's form, and return the rest as content
closed · @signalnine · 0 comentarios · En GitHub
BenchmarksServer & APIDocumentation
Descripción
On a corpus of 1,462 real agent turns, `main` ends 73 of them (5%) with a `server_error`. With this PR it parses 98.3% of the turns as intended, up from 92.3%, and ends none.
This builds on #700. Its commit is the first one here, cherry-picked onto current `main`; if #700 merges first, this rebases to one commit. #700 makes a `<tool_call>` open a call only when a call follows it. This PR widens what counts as a call, for a tool the request declared:
| the model wrote | before | now |
| --- | --- | --- |
| `<tool_call><parameter=NAME>...` (the tool's name as a parameter tag) | `ValueError`, request ends | the call |
| `<tool_call>{"name": NAME, "arguments": {...}}</tool_call>` (also `"parameters"`, or arguments as a JSON string) | `ValueError`, request ends | the call |
| `<function=NAME>...</function>` at the start of a line, no `<tool_call>` (a stray `</tool_call>` after it is dropped) | text to the client, agent stops | the call |
| two calls inside one `<tool_call>` (the second may open as `<parameter=NAME>`) | ONE call with the second's parameters merged into the first | two calls |
Anything else in those forms comes back as content, verbatim: an undeclared name, broken JSON, a JSON object that never closes, `<function=` inside a ``` code block or mid-line. A body that still fails to parse becomes content instead of raising. Every recovery needs a declared tool, so requests without tools behave exactly as before.
Streamed output equals the whole output at any chunk size, with and without `stream_tools`, and an announced call's streamed arguments equal its final ones.
## How I found it
I replayed a corpus of real model output through `OutputParser`: 1,462 agent turns in 176 distinct shapes, captured from Qwen3.8-27B under Claude Code (redacted, from [signalnine/q27](https://github.com/signalnine/q27)'s drift corpus), each labelled with the calls the model meant. `server.py` handles the parser's `ValueError("malformed tool call")` as an engine error, so the client gets `server_error` and the agent's session ends.
| | turns read as intended | shapes | turns that end the request |
| --- | ---: | ---: | ---: |
| main | 92.3% | 113/176 | 73 |
| #700 + this PR | 98.3% | 154/176 | 0 |
The 25 turns still missed are mostly invalid JSON and redaction artifacts. Qwen3.8-Flash-Next drifted less in a live check: 12 SWE-bench Verified instances under Claude Code had 0 malformed calls on `main` and on this branch (117 tool calls on this branch, all parsed; 11 of 12 runs edited the gold patch's file).
## Tests
- New `ToolCallDrift` in `serve/test_server.py`, every case at steps 1 / 7 / whole with `stream_tools` on and off: each recovered form (with corpus shape ids), each form that must stay text, broken and unclosed JSON. The recovery cases fail on `main` + #700 (54 subtest failures) and all pass with the change.
- #210 / #211 / #700 tests pass unchanged.
- `python -m unittest discover -s serve -p 'test_*.py' -t .`: 279 tests, plus one error that `main` has too (`test_responses.OverHttp.test_json_schema_text_format`).
- `docs/DETAILS.md`: a paragraph under Streaming.
One judgement call worth a look: a declared tool's `<function=NAME>` block at the start of a line, outside a code block, counts as a call without the wrapper. A model explaining the format in prose with a real tool's name at the start of a line would run that call. The corpus had 27 turns of the bare form and zero of that prose case.
En el sitio
Enlaces a install, modelos, releases.