Issues / #817
#817 Assistant turn can be interrupted when generation emits a reserved ChatML boundary marker
open · @d4mer · 2 comments · View on GitHub
Description
Sorry for the AI generated issue Issue type: Bug Priority: Please triage; it prevents reliable completion of a conversation Summary In a Hermes desktop conversation, the assistant repeatedly emitted the visible spelling of a reserved ChatML message-start marker as part of its response. Afterward, the assistant turn was interrupted or failed to complete normally. The user observed the failure repeatedly and could not complete the exchange until the assistant stopped reproducing the marker. The visible spelling is intentionally replaced below with [ChatML im_start marker] to avoid reproducing the failure in this issue. Environment Hermes desktop app Active model/provider shown by the runtime: gpt-6-luna via openai-codex The GraphRAG model server was not the active assistant model for this conversation. Steps to reproduce In a test harness, request an assistant response that includes the literal spelling of the reserved ChatML im_start marker as quoted text. Capture the raw inference response, including token IDs, decoded text, stream events, and finish reason. Observe whether the assistant turn completes, is cut off, or is interpreted as a new message boundary. Please run this in a test harness rather than reproducing the marker in a live conversation if the client may parse or render it as control syntax. Expected behavior The inference/provider path should not silently truncate or restructure an assistant turn because generated content contains text resembling a protocol marker. Reserved structural tokens should be handled as protocol structure, not reinterpreted from decoded assistant content. Actual behavior In this conversation, after the marker appeared in assistant output: The user reported that the assistant was blocked. Multiple assistant responses were interrupted or failed to finish a complete turn. The problem recurred when the assistant reproduced the marker. Responses that avoided reproducing it completed normally. Why the cause is not yet certain The conversation does not include raw inference or transport traces, so it does not establish whether: the model generated the reserved token ID; the model generated ordinary characters that happen to spell the marker; or a downstream serializer, stream parser, or client treated decoded content as protocol structure. That distinction matters: if the model generated the special token ID, investigate inference-time token handling and stop behavior; if it emitted ordinary text, investigate downstream parsing and serialization. Requested investigation Please check a failing trace for: The raw generated token IDs versus decoded assistant text. The provider stream’s message boundaries and finish_reason. Whether the marker was emitted inside assistant content or treated as a new structural boundary. Stop sequences, tokenizer decoding, and any post-decoding parser that scans text for protocol-marker spellings. Whether escaping the marker before rendering or storing assistant content prevents the interruption. Impact This can prevent a user from completing a conversation and can repeatedly derail a task when the assistant discusses or quotes protocol-level tokens. This is a reportable failure pattern, but the responsible layer still needs to be confirmed from a raw trace.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.