Pull requests / #1045
#1045 serve: the vision markers inside a message's text stay text and no longer take a picture's place (#150 for the whole marker)
closed · @sergqwer · 0 コメント · GitHub で見る
本文
Replaces #554. GitHub closed it on 2026-10-05, when my fork was made private by mistake: that took the fork out of the network for good, so it can no longer open pull requests. This is the same branch and commit (`1273345`), opened from a new fork; the discussion and the measurements are in #554. --- `Service.prepare()` renders the chat template and tokenizes the whole prompt with `parse_special=True`. The template writes `<|vision_start|><|image_pad|><|vision_end|>` for each image item, and the #150 guard turns only an `<|image_pad|>` **not** preceded by `<|vision_start|>` back into text. Message text that contains the whole marker pair becomes the same special ids as a real image. That happens with an agent reading `serve/chat_template.jinja`, these docs, or a tool result quoting them. If such text comes before a real picture: - picture 1's embedding rows go to the text's marker; - every later picture moves up one; - the last real marker becomes text. The counts still match, so nothing reports it. With Claude Code reading a repository and a screenshot in one conversation, this is an ordinary case. **Fix** (`serve/server.py`, `Service.prompt_ids()`): - The vision markers inside the messages' and tools' strings are swapped for a unique placeholder before rendering. - After rendering, each placeholder becomes its marker's literal (`parse_special=False`) tokens. The pieces between placeholders are tokenized as before, since a special token split the text there anyway. - A prompt without such text takes the old path: byte-identical ids, checked on the real tokenizer too. - `count_tokens` uses the same ids. - If text still yields more marker pairs than there are pictures (a marker cut between two text parts), that is now a 400 instead of a silent shift. Test `test_the_whole_marker_in_text_before_images`: each picture's rows follow its own marker, and the text keeps its literal tokens. Without the fix it finds 3 `vision_start` where 2 belong. The #150 test is unchanged. Serve tests: 173 OK (7 skipped). This touches the first lines of `prepare()`, as #529 does; whichever goes in second needs a one-line rebase. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
関連リンク
インストール・モデル・リリースへの站内リンク。