Pull requests / #1557

#1557 serve: skip tool-body rescans until a closing tag arrives

open · @InB4DevOps · 0 评论 · 在 GitHub 查看

Setup & installServer & APIWindowsLinux

描述

## Summary

Avoid repeatedly walking an accumulating wrapped tool-call body before a closing `</tool_call>` can possibly be present.

- Search newly appended text plus enough overlap for a closing tag split across feeds.
- Once any closing candidate appears, use the original structural walk for the rest of that call, including literal closing tags inside parameter values.
- Reset the cursor when the call completes.
- Recovery mode and calls inside reasoning retain their existing structural paths.
- Add event-parity and cursor/reset tests.

This is independent of #1542. It contains only the tool-call scan change and tests; no partial-tag optimization, newline-free tracking, profiling instrumentation, native-engine changes or GPU changes.

## Standalone validation

Validated against a same-day untouched checkout of upstream main at `fb58e0dbc8399662c0e47c76578c6e878b14f6cf`. The candidate checkout contained only these two changed files, excluding all other local optimizations.

```text
python -m unittest serve.test_parser_call_scan serve.test_frontend serve.test_reasoning_tools serve.test_server
Ran 295 tests — OK
```

Tests compare emitted events after each fragment, including fixed/random chunk boundaries, literal tags, consecutive calls, incomplete calls, recovery, reasoning tools and server API behavior. Random tool IDs are excluded from semantic comparisons. Cursor tests assert that structural walks are skipped before a closing candidate and reset for subsequent calls. `git diff --check` passed.

**Windows correctness and performance validation: pending.**

## Measurements

Intel Core i7-12700KF, 20 logical CPUs; Linux 7.0.0-38-generic, glibc 2.39, CPython 3.12.3, Jinja2 3.1.6. Default scheduling/power policy, no affinity pinning. Seven fresh-process samples per checkout with alternating baseline/candidate order, two warmups per case, fixed 16-character chunks, profiling disabled. Fixture creation, validation and hashing are outside timing. No model or GPU was needed.

Median elapsed milliseconds:

| Workload | 32 KiB argument: upstream → PR | 128 KiB argument: upstream → PR |
|---|---:|---:|
| Buffered large tool call | 5.37 → 1.32 | 137.31 → 11.50 |
| Streamed large tool call | 9.24 → 5.19 | 153.52 → 26.70 |
| Many short streamed calls (128 / 512 calls) | 3.00 → 2.92 | 11.92 → 11.73 |
| Literal closing tag early in argument | 9.26 → 9.39 | 152.19 → 153.01 |
| Recovery enabled | 13.89 → 14.08 | 280.65 → 280.49 |
| Unbroken-line control (128 / 512 KiB) | 20.45 → 20.35 | 193.63 → 193.73 |

At 128 KiB, buffered parsing used **91.6% less elapsed time** and streamed parsing used **82.6% less**. The streamed candidate range was 26.21–26.99 ms versus 152.40–155.21 ms upstream. These are synthetic CPU parser measurements, **not end-to-end inference-throughput gains**.

The fallback paths show little change, including small slowdowns retained in the table. This change does not fix repeated current-line/tool-body string copying or guarantee linear scaling on all inputs.

All 168 measured outputs matched known content, final arguments and concatenated streamed JSON; hashes matched across both checkouts and every sample.

### Minimal workload reproduction

Run this from each checkout's repository root with Jinja2 installed. It reproduces the 128 KiB streamed workload; change `stream` to `False` for the buffered case. Repeat in fresh processes, alternating checkouts. The snippet prints one post-warmup sample.

```python
import json
import time
from serve.frontend import OutputParser

stream = True
value = "abcdefgh" * (1024 * 16)
text = "<tool_call><function=write><parameter=text>" + value + "</parameter></function></tool_call>"
chunks = [text[i:i + 16] for i in range(0, len(text), 16)]
tools = [{"name": "write", "parameters": {"properties": {"text": {"type": "string"}}}}]
expected = ("", [("write", {"text": value})],
            json.dumps({"text": value}, separators=(",", ":")) if stream else "")

def run():
    parser = OutputParser(thinking=False, tools=tools, stream_tools=stream, recover=False)
    content, calls, args = [], [], []
    def collect(events):
        for event in events:
            if event.kind == "content":
                content.append(event.text)
            elif event.kind == "tool_call":
                calls.append((event.call.name, event.call.arguments))
            elif event.kind == "tool_args":
                args.append(event.text)
    for chunk in chunks:
        collect(parser.feed(chunk))
    collect(parser.finish())
    return "".join(content), calls, "".join(args)

for _ in range(2):
    assert run() == expected
start = time.perf_counter_ns()
result = run()
elapsed_ms = (time.perf_counter_ns() - start) / 1e6
assert result == expected
print(len(value), len(chunks), elapsed_ms)
```

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。