Issues / #537

#537 Literal </think> quoted in reasoning yields empty stop or leaked reasoning (v0.1.37)

closed · @fenrir-labs76 · 6 コメント · GitHub で見る

Server & APINVIDIA / CUDAModels & quantsSecurity

本文

# Literal `</think>` quoted in reasoning yields empty `stop` or leaked reasoning (v0.1.37)

## Summary

Strata itself mishandles a literal `</think>` in ordinary user text and generated
reasoning. A request to **quote that exact string** returns HTTP 200 with
`finish_reason: "stop"` and no answer, well below `max_tokens`. A separate
request asking about its HTML-escaped form can split `reasoning_content` in
the middle of a quote: reasoning then appears in `content` and the actual
reasoning delimiter can leak into the answer. Reproduced with direct
`/v1/chat/completions`, with and without streaming; OpenCode is not required.

## Version / configuration

- Latest stable upstream source: **v0.1.37**, tag commit
  `db4f91a1171d697928b0d2e1f50ef95e25559d4c` (unmodified source).
- Locally built from official Dockerfile with CUDA base pinned to a digest;
  resulting image ID `sha256:c54ffff355c5f7b16807fbcd2c432b872022420f81e43aae5d1919b7056c73a1`.
  Previous unmodified v0.1.33 image also reproduced it.
- Qwen3.8-Flash-Next **IQ3_S** model, model ID
  `qwen3.8-flash-next-iq3_s`; RTX 5090; text-only, context 262144, INT8 KV,
  single serving slot. Live `/props` reports chat-template SHA256
  `12827f24b742ea4e80cdc12dbcf9622227056b9f797252a3149263d4f9aaadce`.
- Explicit `max_tokens: 512`; no request `stop` field, no proxy stop sequence;
  default reasoning on. One request at a time. No patched image or model.

## Minimal direct reproduction

Use Strata's own local OpenAI-compatible endpoint; substitute your local port
if necessary. No OpenCode, auth, extensions, tools or external data required.

```bash
curl -sS --max-time 30 http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3.8-flash-next-iq3_s","stream":false,"max_tokens":512,"messages":[{"role":"user","content":"Quote this exact literal string, then explain why it appears in some chat transcripts: </think>"}]}'

curl -sS -N --max-time 30 http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3.8-flash-next-iq3_s","stream":true,"max_tokens":512,"messages":[{"role":"user","content":"Quote this exact literal string, then explain why it appears in some chat transcripts: </think>"}]}'
```

**Expected:** complete answer quoting `</think>` and explaining that it is a
reasoning delimiter; `reasoning_content` remains separate and
`finish_reason: "stop"` means the answer actually ended. **Actual:** non-stream
response has `message.content: null`, `message.reasoning_content` ends
`... Quote this exact literal string, then explain why it appears in some chat transcripts: `,
and `finish_reason: "stop"`. SSE emits the same reasoning prefix, final chunk
`delta: {}, finish_reason: "stop"`, then `[DONE]`; zero answer deltas. HTTP 200.
Output budget was not exhausted.

Second direct input isolates a misclassified *generated* quoted tag even when
the user does not send a raw delimiter:

```bash
curl -sS --max-time 30 http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3.8-flash-next-iq3_s","stream":false,"max_tokens":512,"messages":[{"role":"user","content":"What does the HTML entity &lt;/think&gt; represent? Please write its literal text and explain."}]}'
```

**Expected:** private reasoning, then a concise answer explaining that the
entities spell `</think>`. **Actual example:** `reasoning_content` ends
`... so literal text is ` while `content` begins
`. It represents closing tag? In HTML, &lt; and &gt; are character references ...`;
later `content` includes another literal `</think>` and a repeated answer.
The SSE path also switches from `delta.reasoning_content` to `delta.content`
at the first quoted tag.

A normal control prompt `What is 1+2+4+8+16? Give sum and formula briefly.`
returns separate reasoning and final answer `31` on both APIs.

Multi-turn control: user `What about 1+2+4+8+16+....+n`; assistant message
`Powers of 2: sum = 2n - 1.\n</think>\nGeometric series, ratio 2: 2n - 1.`;
user `why did you write </think>`. It can start `content` with reasoning like
`". Need understand context...`, return no answer, or end with `length` after
misclassifying the prior assistant message. This is supplementary; single-turn
cases above isolate the defect without prior history.

## Observed frequency / boundary

Three serial repetitions per case per mode against each release, same prompts
and output cap:

| Release | Input / outcome                                 | JSON | SSE |
| ------- | ----------------------------------------------- | ---- | --- |
| 0.1.33  | Ordinary math: correct final answer             | 3/3  | 3/3 |
| 0.1.37  | Ordinary math: correct final answer             | 3/3  | 3/3 |
| 0.1.33  | Literal quote: `stop` with empty final answer   | 3/3  | 3/3 |
| 0.1.37  | Literal quote: `stop` with empty final answer   | 3/3  | 3/3 |
| 0.1.33  | Escaped tag: reasoning starts in answer content | 3/3  | 3/3 |
| 0.1.37  | Escaped tag: reasoning starts in answer content | 3/3  | 3/3 |

The two-turn case varied across repetitions; not claimed deterministic. A
follow-up with `chat_template_kwargs.enable_thinking=false` still truncated
literal-tag quotation on v0.1.33, so merely turning thinking off was not a
complete workaround.

## Source inspection: confirmed vs hypothesis

At the v0.1.37 commit, `serve/frontend.py:508` calls
`self.buf.find(THINK_END)` during reasoning and transitions permanently to
content at that first generated `</think>`. `serve/server.py:1335-1336` renders
the chat template then encodes it with `parse_special=True`;
`serve/server.py:1475` checks configured end-of-turn token IDs. Live tokenizer
encodes literal `</think>` as ID `248069`, even with `parse_special=False`;
the configured end-of-turn IDs are different. These facts explain how a
quoted tag can collide with the parser's boundary. **The exact reason for
the early backend `stop` (including which generated token ended the turn) is
not proven by these HTTP captures**; please do not assume quantization,
OpenCode rendering or an OpenCode stop list is responsible.

Related but distinct: #210 fixed delimiters inside **tool parameters** in
v0.1.28. This report covers reasoning's `</think>` boundary and an early
backend `stop`, still present in the official v0.1.37 build. #530 concerns
`finish_reason: "length"` from exhausting a thinking budget; the empty-answer
quote here returns `stop` after a short response.

関連リンク

インストール・モデル・リリースへの站内リンク。