Issues / #843

#843 Tool calls silently dropped in agent sessions after an empty assistant turn (fix: skip empty assistant turns when rendering)

closed · @jramirez-uab · 2 comentários · No GitHub

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsLinux

Descrição

## Summary
With opencode as the client, Strata sometimes returns an **empty assistant turn** (no text, no tool call) when the user asks for an
action. Once one such turn is in the history, the model **imitates it**: later turns close the reasoning and end with
`</think>\n\n<|im_end|>` without emitting the `<tool_call>`. The client receives an empty answer and the agent "stops".
Skipping empty assistant turns from the history in `serve/frontend.py` (`render`) fixes it in our tests: **0-6/20 → 20/20**.

## Environment
- Strata v0.1.39 (engine compiled from source with CUDA 13.2), Linux, 1× NVIDIA L40S (48 GB), 251 GB RAM, AMD EPYC 7402P (AVX2).
- Models: `qwen` IQ3_S (GSQ-RCO) and `swift` IQ3_XXS; 262,144 context, KV int8, images on, default MTP settings.
- Client requests: the real request opencode sends (its system prompt, 11 tools), replayed with `stream: false`, `reasoning_effort: xhigh`,
  at temperature 1.0 (none sent) and 0.3.

## How to reproduce
A multi-turn conversation whose history contains assistant turns with empty `content` and no `tool_calls` (only `reasoning_content`),
e.g. the user asked "haz un ls" three times and each assistant turn came back empty, then a fourth "haz un ls".
With `STRATA_DEBUG=1` the raw model text of the failing answers is:
```
The user is asking me to run `ls`. Let me do that.
</think>

<|im_end|>
```
i.e. the model never writes the tool call; the parser is not at fault.

## Results (20 tries per case)
| Case | T=0.3 | T=1.0 |
|---|---|---|
| fresh request "haz un ls" | 20/20 | 20/20 |
| after a long story, "haz un ls" | 20/20 | 20/20 |
| normal session (earlier turns did call tools) → new request | 20/20 | 20/20 |
| **history with 3 empty assistant turns → "haz un ls"** (Swift IQ3_XXS) | **0/20** | 18/20 |
| same, unpatched, other runs (Qwen IQ3_S / Swift) | 2/5 – 6/20 | 1/20 – 3/5 |
| **same, with the patch below** (Swift IQ3_XXS and Qwen IQ3_S) | **20/20** | **20/20** |

Things we ruled out: the chat template (it is Qwen's original plus Unsloth's fixes), the tool-call parser (the call is never generated),
`--prompt-cache 0` (worse: 0/5), `--no-coupled-draft` (no consistent change), `preserve_thinking: false` (still 6/20).

## Patch
```diff
--- a/serve/frontend.py
+++ b/serve/frontend.py
@@ -49,6 +49,14 @@
 
     def render(self, messages: list[dict], tools: list[dict] | None = None, add_generation_prompt: bool = True,
                **kwargs) -> str:
+        # Skip empty assistant turns (no text, no tool calls) from the history, except the last message. With them in the
+        # history Qwen3.8-Flash-Next / Swift imitate the pattern and end the next turn without the tool call (0-6/20 -> 20/20).
+        # STRATA_KEEP_EMPTY_TURNS=1 keeps the old behaviour.
+        import os
+        if not os.environ.get("STRATA_KEEP_EMPTY_TURNS"):
+            messages = [m for i, m in enumerate(messages)
+                        if i == len(messages) - 1 or not (isinstance(m, dict) and m.get("role") == "assistant"
+                                                          and not _text_of(m.get("content")).strip() and not m.get("tool_calls"))]
         return self.template.render(messages=messages, tools=tools, add_generation_prompt=add_generation_prompt,
                                     **kwargs)
 
```
The last message is never removed. An environment variable (`STRATA_KEEP_EMPTY_TURNS=1`) keeps the old behaviour; feel free to make it
a config option instead.

## Related
- #804 is different: there the model writes the tool call inside `<think>`; here it never writes it.
- Unrelated, but for completeness: with `--gpus 0,1` and `"parallel": 2` we also hit the engine exit reported in #776
  (`captured the batch window over slots 0`) on 2× L40S; one GPU with `parallel` 2, or two GPUs with `parallel` 1, work.

Thanks for Strata: on a single L40S the prompt speed stays flat (~4,300 tok/s) all the way to 198K tokens.

No site

Links install, modelos, releases.