Pull requests / #1045

#1045 serve: the vision markers inside a message's text stay text and no longer take a picture's place (#150 for the whole marker)

closed · @sergqwer · 0 comments · View on GitHub

Server & APIDocumentation

Description

Replaces #554. GitHub closed it on 2026-10-05, when my fork was made private by mistake: that took the fork out of the network for good, so it can no longer open pull requests. This is the same branch and commit (`1273345`), opened from a new fork; the discussion and the measurements are in #554.

---

`Service.prepare()` renders the chat template and tokenizes the whole prompt with `parse_special=True`. The template writes `<|vision_start|><|image_pad|><|vision_end|>` for each image item, and the #150 guard turns only an `<|image_pad|>` **not** preceded by `<|vision_start|>` back into text.

Message text that contains the whole marker pair becomes the same special ids as a real image. That happens with an agent reading `serve/chat_template.jinja`, these docs, or a tool result quoting them. If such text comes before a real picture:
- picture 1's embedding rows go to the text's marker;
- every later picture moves up one;
- the last real marker becomes text.

The counts still match, so nothing reports it. With Claude Code reading a repository and a screenshot in one conversation, this is an ordinary case.

**Fix** (`serve/server.py`, `Service.prompt_ids()`):
- The vision markers inside the messages' and tools' strings are swapped for a unique placeholder before rendering.
- After rendering, each placeholder becomes its marker's literal (`parse_special=False`) tokens. The pieces between placeholders are tokenized as before, since a special token split the text there anyway.
- A prompt without such text takes the old path: byte-identical ids, checked on the real tokenizer too.
- `count_tokens` uses the same ids.
- If text still yields more marker pairs than there are pictures (a marker cut between two text parts), that is now a 400 instead of a silent shift.

Test `test_the_whole_marker_in_text_before_images`: each picture's rows follow its own marker, and the text keeps its literal tokens. Without the fix it finds 3 `vision_start` where 2 belong. The #150 test is unchanged. Serve tests: 173 OK (7 skipped).

This touches the first lines of `prepare()`, as #529 does; whichever goes in second needs a one-line rebase.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.