Pull requests / #1591

#1591 tokenizer: sweep protected spans in match order

open · @InB4DevOps · 0 コメント · GitHub で見る

Setup & installServer & APIModels & quantsWindowsLinux

本文

## Summary

Avoid scanning every protected plain-text span for every special-token regex match in `Tokenizer._encode_matching()`.

Prompts quoting reasoning/control-token syntax (for example, a tool result containing a chat template) can have many protected spans. The old `any(a <= m.start() < b for a, b in plain)` repeats the scan from the beginning for each match.

- Sort a copy of list/tuple spans and sweep forward as matches arrive in text order, remembering the furthest protected end.
- Preserve the half-open, match-start-only protection rule, including unsorted, nested and overlapping spans.
- Preserve the original consumption behavior for other iterables, including one-shot iterators.
- Read the match position once and reuse it.

The span check becomes O(matches + spans) after sorting, rather than O(matches × spans). Sorting requires a temporary span list and is O(spans log spans) generally; the server already supplies ordered spans. Empty-span prompts do not allocate or sort a span list.

## Scope and isolated base

Only `tools/strata_tokenizer.py` and `tools/test_tokenizer_plain_spans.py` are included.

Both the untouched baseline and candidate started at upstream `fb58e0dbc8399662c0e47c76578c6e878b14f6cf`. **The cache-initialization change from #1582 is excluded**, as are the earlier parser optimizations and profiling instrumentation. Server, frontend, template and synthetic-fixture hashes match across checkouts. BPE, cache behavior and special-token definitions are unchanged.

## Correctness

```text
python -m unittest tools.test_tokenizer_plain_spans tools.test_strata_tokenizer serve.test_detok serve.test_frontend serve.test_server
Ran 295 tests — OK (skipped=5)
```

**290 passed; 5 tests requiring unavailable tokenizer assets were skipped.** New exact-token-ID parity tests compare against the original implementation for unsorted/overlapping/adjacent/nested spans, empty/reversed/out-of-range spans, Unicode offsets, boundaries inside literals, both special-token modes, BPE merges, input immutability, empty text, absent special patterns and one-shot iterables. `git diff --check` passed.

**Windows correctness and performance validation: pending.**

## Same-day paired CPU measurements

Intel Core i7-12700KF, 20 logical CPUs; Linux 7.0.0-38-generic, glibc 2.39, CPython 3.12.3, regex 2026.9.29, Jinja2 3.1.6. Default scheduling/power policy, no affinity pinning. Seven fresh-process samples per checkout, alternating baseline/candidate order, two warmups per workload. Fixture creation, validation and hashing were outside timing. No model, GPU or profiling timers were needed.

Measured real `Service.prepare()` using the existing synthetic tokenizer fixture (seed 268, 1,500 merges), the repository chat template, and warm piece caches. Each quoted snippet contains four protected literals.

| Prompt preparation | Untouched upstream median | PR median |
|---|---:|---:|
| 128 quoted snippets / 512 protected spans | 6.52 ms | 1.17 ms |
| 512 quoted snippets / 2,048 protected spans | 86.56 ms | 4.08 ms |
| Ordinary long code prompt, scale 4 | 3.01 ms | 3.03 ms |
| Ordinary conversation, scale 4 | 4.21 ms | 4.23 ms |

The larger quoted prompt used **95.3% less elapsed time**, with a candidate range of 4.01–4.26 ms versus 85.32–88.72 ms upstream. Ordinary controls stayed close to baseline; this is a quoted-prompt CPU preparation improvement, **not an inference-throughput claim**.

All **84 measured outputs** matched between checkouts and samples, including exact prompt-token hashes. The synthetic vocabulary is not a full model vocabulary.

### Reproduction

With `regex` and `jinja2` installed, run this from each checkout root. It reproduces one post-warmup sample for the quoted prompt at each size. Repeat in fresh processes and alternate checkout order for comparisons; the printed hash should match between branches.

```python
import hashlib
import json
import time
from pathlib import Path
from serve.server import Service, MockEngine
from serve.frontend import ChatTemplate
from serve.test_detok import synthetic_tokenizer

tok = synthetic_tokenizer(seed=268, n_merges=1500)
svc = Service(MockEngine(tok, "", max_context=2_000_000), tok,
              ChatTemplate(Path("serve/chat_template.jinja")))

for scale in (1, 4):
    messages = [{"role": "user", "content":
                 "Quote <think>literal</think> <|im_start|>example<|im_end|>\n" * (128 * scale)}]
    def run():
        return svc.prepare(messages, None, {"enable_thinking": False}, max_new=256)
    expected = run()
    for _ in range(2):
        assert run() == expected
    start = time.perf_counter_ns()
    result = run()
    elapsed_ms = (time.perf_counter_ns() - start) / 1e6
    assert result == expected
    print(scale, elapsed_ms, len(result[0]),
          hashlib.sha256(json.dumps(result, ensure_ascii=False).encode()).hexdigest())
```

関連リンク

インストール・モデル・リリースへの站内リンク。