Pull requests / #1542
#1542 serve: avoid exhaustive partial-tag suffix checks
open · @InB4DevOps · 0 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APIWindowsLinux
Beschreibung
## Summary
Speed up `OutputParser._hold()` by searching the bounded suffix for candidate tag-prefix starts instead of checking every possible prefix with `str.endswith` on every token delta.
The earliest matching candidate is the longest proper-prefix suffix. Overlapping candidates, empty/one-character tags, Unicode text, and streamed tag boundaries retain the original behavior.
Adds parity tests against the original algorithm, including exhaustive short strings and randomized chunk boundaries through reasoning, quoted code, tool calls, streamed tools, and recovery.
## Why
Profiling a synthetic 49,664-token output workload identified 496,640 suffix checks in `_hold()`. Ordinary prose generally contains no candidate `<`, so the new path replaces these checks with a bounded C-level string search.
Local Linux measurements on an Intel Core i7-12700KF, CPython 3.12.3, regex 2026.9.29 and Jinja2 3.1.6, with profiling disabled, seven fresh-process samples and two warmups per workload:
| Synthetic workload | Original parser median | New parser median |
|---|---:|---:|
| Detokenization + parser + collection | 88.35 ms | 59.73 ms |
| Parser + collection, predecoded deltas | 71.76 ms | 46.01 ms |
| Collection alone | 6.51 ms | 6.54 ms |
| Detokenization alone | 12.68 ms | 12.75 ms |
The combined CPU output-processing workload used 32.4% less elapsed time. This is a synthetic Python workload, **not an end-to-end inference-throughput claim**. No model or GPU was needed. Windows performance has not yet been measured.
The measurements came from local profiling/benchmark work; this PR contains only the parser optimization and correctness tests.
### Minimal workload reproduction
With `regex` and `jinja2` installed, run this Python snippet from the repository root on the base and PR branches. It recreates the measured combined workload (49,664 tokens, 66,048 UTF-8 output bytes). Run in fresh processes for repeated samples; fixture creation and correctness checks are outside the timer.
```python
import time
from serve.server import Detokenizer
from serve.frontend import OutputParser
from serve.test_detok import synthetic_tokenizer
tok = synthetic_tokenizer(seed=268, n_merges=1500)
code = "def accumulate(values):\n return sum(x * x for x in values)\n# café 日本語 🙂\n"
text = ("The result is correct. café 日本語 🙂\n" + code) * 512
ids = tok.encode(text, parse_special=False)
def run():
detok, parser, chunks = Detokenizer(tok), OutputParser(thinking=False), []
for token in ids:
chunks.extend(e.text for e in parser.feed(detok.push(token)) if e.kind == "content")
chunks.extend(e.text for e in parser.finish() if e.kind == "content")
return "".join(chunks)
for _ in range(2):
assert run() == text
start = time.perf_counter_ns()
result = run()
elapsed_ms = (time.perf_counter_ns() - start) / 1e6
assert result == text
print(len(ids), len(text.encode("utf-8")), elapsed_ms)
```
## Validation
Validated the isolated PR branch, based on current upstream `main`:
```text
python -m unittest serve.test_parser_hold serve.test_frontend serve.test_reasoning_tools serve.test_server
Ran 287 tests — OK
```
`git diff --check` also passed. The implementation uses portable Python string operations; Windows execution is pending.
Mehr auf der Site
Links zu Install, Modellen, Releases.