Pull requests / #567
#567 serve: tokenise the prompt from the last shared prefix, not from the start
closed · @gputier · 0 comments · View on GitHub
Setup & installServer & APIModels & quantsWindowsLinux
Description
Based on v0.1.39 since 5 October (it was first written on v0.1.37). The rebase takes blange48's port: on 0.1.39 the plain spans of #537 (a `</think>` quoted in a message is text) also decide which special matches are taken, so the shared prefix also ends where the two prompts' plain spans first differ, and the rest is encoded with its spans shifted by the resume point. A chat client sends the whole conversation every turn, and `Service.prepare` runs BPE over all of it each time (`serve/server.py`, `self.tok.encode(prompt, parse_special=True)`), even when the engine's conversation cache already holds the whole prefix. `/v1/messages/count_tokens` does the same. This PR adds `PromptEncoder` in `tools/strata_tokenizer.py`. It keeps the last few rendered prompts with their ids. For a new prompt it finds the last special-token boundary shared with one of them, reuses the ids up to that boundary, and sends only the text after it through BPE. The ids match a full encode by construction. Encoding splits the text at special-token matches and encodes each stretch between two matches on its own, so nothing crosses a match. Whether a match is taken at position q depends only on `text[q : q + max_special_len]` and on whether q lies inside a plain span, so the cut sits at least `max_special_len - 1` characters before the end of the common prefix. The class docstring has the reasoning. The tokenizer gains `max_special_len` and `encode_marked()` (`encode()` plus the offset and id count after each special-token match). `ByteTokenizer` gets the same two, so the mock-engine tests exercise the encoder. 0.1.39's `Service.encode_prompt()`, which `prepare` and `count_tokens` already use, now calls the encoder. `encode()` itself is untouched. Diff: 5 files, 3 of them tests. ## Measurement A conversation of 100,893 tokens (433 K characters) that grows by one turn, encoding time for the new prompt, 5 repetitions: - full encode: 235 ms median (231 to 241) - `PromptEncoder`: 1.5 ms median (1.05 to 2.07) Container with 2 cores, Python 3.12, llama.cpp's `ggml-vocab-qwen35.gguf` (with the template's added tokens), synthetic text mixing English, code and CJK. This is not the model pack's tokenizer and not the target machine, so the absolute numbers will differ there. ## Evidence that the ids are identical On v0.1.39, in a Linux container (Python 3.13, llama.cpp's qwen35 vocabulary): `serve.test_prompt_encoder` and `serve.test_server` pass 146 tests, `tools/test_strata_tokenizer.py` passes 11 with 1 skipped (the pack tokenizer). The rest of this section was measured on the v0.1.37 version. - `serve/test_prompt_encoder.py` (blange48's, added with the rebase): a fuzz against the full encode with random plain spans, spans that change inside the shared part, and edited histories, plus quoted `<think>` / `</think>` across turns through `Service.encode_prompt`. - `tools/test_strata_tokenizer.py` (new, 11 tests): llama.cpp's qwen35 vectors, `encode_marked` against `encode`, `chat_golden.json`, and conversations compared token for token with `tok.encode`. Those cover several turns with unicode and literal special tokens in the text, tool definitions, calls and responses, images, a changed reasoning effort or system message, an edited earlier turn, a conversation that gets shorter, an empty message, and interleaved conversations, with 1, 2 and 4 kept prompts. Separate tests use byte-level vocabularies whose special tokens overlap each other, which is where a cut without the look-ahead margin breaks. - `serve/test_server.py`, `IncrementalPrompts`: six turns over HTTP against the mock engine, the same conversations through `Service.encode_prompt` with both the byte tokenizer and qwen35, and a tokenizer without `encode_marked`. - Mutations: with the margin set to 0, 5 tests fail. With the cut one id too late, 305 fail. - Fuzz run outside the test suite: 40,000 cases, no difference in ids. 8,000 on qwen35 (seeds 1000 to 1099) and 32,000 on byte vocabularies with overlapping literals. - Full Python suite (`tools/test_*.py`, `serve/test_*.py`) in a Linux container: 454 passed, 9 skipped, 0 failed on this branch, against 440 passed, 9 skipped on v0.1.37. That run uses scratch copies of both trees where `setup.EXE` is forced to `strata.exe`. Without it, 47 tests fail on Linux, the same 47 on a clean v0.1.37 (test_setup_golden and one test_setup_amd test), because they expect the Windows executable name. The change is not in this PR. The skips are the pack tokenizer, `jsonschema` and the Windows job objects. No existing test is modified. ## Limits - No fallback switch. The default path changes, with the same ids, and there is no environment variable to go back to the full encode. I can add one if you want it. - `Service` now imports `strata_tokenizer`, and so `regex`, whenever the tokenizer has `encode_marked`. That includes the byte tokenizer used in tests without a pack. `regex` is already pinned in `requirements.txt`. - The server keeps up to 4 rendered prompts in memory. Measured at 239,670 tokens (572,670 characters): one entry holds 3.6 MB of ids on top of its text, so the cap of 4 entries is about 15 MB of ids. - `/v1/messages/count_tokens` shares the same encoder, so a count request can evict another client's entry. The cost is speed only, the ids stay the same. - Not tested against the real engine or the model pack's vocabulary, only the mock engine and llama.cpp's qwen35 vocabulary.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.