Issues / #1053
#1053 A reply that ends inside `<think>` (no `</think>`) comes back empty with `finish_reason: "stop"`, and the agent stops — close the thinking and continue once
closed · @davidsheridan77-dot · 3 comentários · No GitHub
BenchmarksSetup & installServer & APIAMD / HIPModels & quants
Descrição
## Summary
With thinking on, Qwen3.8-Flash-Next sometimes writes one sentence of reasoning and then `<|im_end|>`, **without
closing `</think>` and without a tool call**. `OutputParser` is still in the `reasoning` state, so the whole reply is
`reasoning_content`, `content` is empty and `finish_reason` is `stop`. An agent harness sees an empty turn right
after a tool result and stops (ours, OpenClaw, then runs a tools-off "finalization" pass, so the model can only
reply that it cannot do the edit).
This is the same family as #804 and #843, but a different shape: there is **no `<tool_call>` inside the thinking**
(so #970 / the `rcall` state don't catch it), and no earlier empty turns in the history (fresh session).
Strata already has the machinery for this: the #123 thinking-budget wrap-up closes the thinking and continues from
the held prefix. Doing the same once when the model **stops** inside its thinking fixes it here: 20/20 vs ~15% failures.
## Environment
```text
GPU AMD RX 9060 XT 16 GB (gfx1200), PCIe 3.0
CPU/RAM Ryzen 5 5500, 64 GB DDR4
OS Ubuntu 24.04, ROCm (system /opt/rocm), HIP engine built by setup
Strata v0.1.39 (6f32ec0)
Model Qwen3.8-Flash-Next GSQ-RCO Q2_0, --kv int8, --spec 4 --spec-min-p 0.5, MTP, 32K and 64K context
Client OpenClaw agent, OpenAI Chat Completions, streaming, reasoning_effort "low", 16 tools
```
## What the model generated
Task: "One of the files in configs/ has debug turned on. Turn it off, then create changelog.md with one line saying
which file you changed." The agent globs, greps (finds `b.conf: debug = true`), reads `b.conf`. The next reply is
13 tokens. `STRATA_DEBUG=1` (set on the **server** process — in the config's `"env"` block it only reaches the
engine child, so nothing is printed) shows, identical in all 3 failing runs:
```text
[strata] done: 13 tokens in 2 s (50.5 tok/s) (stop, cancel=False), expert cache 90.5% hit ...
[strata] raw: 'I need to investigate further. Let me check the details.<|im_end|>'
```
`OutputParser(thinking=True).feed(...)` on that text gives only a `reasoning` event — no content, no call.
## Frequency (same task, fresh session per run)
| Run | Engine / server | Result |
| --- | --- | --- |
| A | 0.1.39, stock server | 9/10 pass, 1 empty-stop turn |
| B | 0.1.39, stock server, `STRATA_DEBUG=1` | 17/20 pass, 3 empty-stop turns (raw text above) |
| C | 0.1.39 + the patch below | **20/20 pass**; the patch fired 2x, 0 empty turns |
In run C, after the inserted `</think>` the model wrote the correct call both times (tool name and path shortened):
```text
[strata] raw: 'I need to investigate further. Let me check the details.<|im_end|>\n</think>\n\n<tool_call>\n<function=edit_file>\n<parameter=file_path>\n.../configs/b.conf\n</parameter>\n<parameter=old_string>\ndebug = true\n</parameter>\n<parameter=new_string>\ndebug = false\n</parameter>\n</function>\n</tool_call><|im_end|>'
```
(The `<|im_end|>` before `</think>` is only in `raw_ids`; the continuation prompt is `prompt + seg + close`, without
the stop token.)
A full 24-task agent ladder on the patched server: 24/24, decode 45.0 tok/s (44.7 unpatched) — no regressions seen.
`python -m unittest serve.test_server serve.test_responses`: same result as unpatched (1 pre-existing error,
`test_json_schema_text_format`, also without the patch).
## Proposed change
On a stop token while `parser.state == "reasoning"` (and nothing held back), append `"\n</think>\n\n"` and continue
once, the same way the #123 budget wrap-up does. Opt-out: `STRATA_STOP_IN_THINKING=0`. Once per reply, so a model
that stops again after the close ends normally (its text is then `content`).
Against v0.1.39; on current `main` the stop-token branch in `Service.run` is unchanged, so it should port directly
(there it would also want to skip the `rcall` state).
```diff
--- a/serve/server.py
+++ b/serve/server.py
@@ -76,6 +76,9 @@
REASONING_WRAP_UP = "\n\nI have thought about this long enough; time to give my answer.\n</think>\n\n"
+# A reply that stops while still thinking (no </think>, so no answer and no tool call) is closed and continued
+# once, the way #123 closes a thinking budget. STRATA_STOP_IN_THINKING=0 turns it off.
+STOP_IN_THINKING_CLOSE = "\n</think>\n\n"
@@ -2292,10 +2295,12 @@ class Service:
prompt, thought = ids, 0 # thought: the reasoning tokens so far (the budget's count)
+ stop_close = thinking and os.environ.get("STRATA_STOP_IN_THINKING", "1") != "0"
while True:
...
seg, wrap, leaving = [], False, False # this pass's tokens; the budget is reached; closed
+ resume = False # stopped while still thinking
@@ -2309,6 +2314,9 @@ class Service:
if t in self.stop_ids:
finish = "stop"
raw_ids.append(t)
+ if stop_close and parser.state == "reasoning" and not parser.buf \
+ and not detok.pending():
+ resume = True # no </think> yet = no answer, no tool call
break
@@ -2350,6 +2358,22 @@ class Service:
+ if resume and not cancel.is_set():
+ stop_close = False # once per reply
+ extra = self.tok.encode(STOP_IN_THINKING_CLOSE, parse_special=True)
+ if max_new - n - len(extra) >= 1:
+ print("[strata] the reply stopped inside its thinking (no answer, no tool call): "
+ "closing the thinking and letting it answer", flush=True)
+ for t in extra:
+ n += 1
+ raw_ids.append(t)
+ evs = parser.feed(detok.push(t))
+ self._note(n, evs, st, rate)
+ for ev in evs:
+ yield "event", ev
+ finish = "length"
+ prompt = prompt + seg + extra
+ continue
if not wrap or cancel.is_set():
break
```
Open questions for you:
- Default on or opt-in? We run it on; it only acts on a reply that would otherwise be empty.
- Worth a `/metrics` counter like #970's `implicit_reasoning_ends`?
- Only tested with one request at a time (no `--batch` / `parallel`).
Happy to turn this into a PR against `main` with a mock-engine test (stop inside thinking → continued; switch off →
unchanged; normal answer → untouched; stops twice → closed once).
No site
Links install, modelos, releases.