Issues / #710
#710 v0.1.38: cross-request tool loop repeats an ineffective repair 98 times despite a thinking budget
closed · @Traveller23 · 5 comments · View on GitHub
Server & APINVIDIA / CUDAModels & quantsWindowsLinux
Description
On Strata v0.1.38, a text-only review task entered a **cross-request loop of client-executed tool calls**, despite an existing per-request thinking budget. The model repeatedly issued the same ineffective repair command and did not complete delivery before the harness deadline. Opening a standalone issue because #75 and #602 are both closed. Their earlier reports are related history; this report specifically concerns repeated tool calls separated by tool results and subsequent model requests, rather than only repeated thinking text within one response. The root cause has not been isolated. ### Observed behavior During one text-only code/document review, the model drafted a JSONL report, then repeatedly attempted the same ineffective JSON repair: - **98 consecutive identical shell commands**, spanning **2,091.936 seconds (about 34.9 minutes)**. - At an intermediate checkpoint, the draft had 118 lines and the same five lines still failed JSON parsing. The repair command returned exit code 0; successful tool execution did not mean the intended repair had succeeded. - At minute 55, the harness delivered a message requiring the model to consolidate the report and finish within the remaining five minutes. Delivery was confirmed in the original session. - After that message, the model did repair the JSON syntax, but continued rereading source files instead of completing the task. - At minute 60, the harness interrupted the active turn. The native task ended as `incomplete`; total reviewer elapsed time was **3,600.738 seconds**. The syntactically valid draft was retained and was not treated as a completed delivery. This was continued model/tool activity, rather than an engine that stopped emitting tokens. We did not repair the draft ourselves or rerun the task to replace the failed observation. ### Environment and relevant settings - Strata **v0.1.38**, official prebuilt NVIDIA engine; also the latest published release when checked for this report. - Qwen3.8-Flash-Next GSQ-RCO **Q2_0**, text only. - Inference on **Windows native Python/CUDA**; **Codex 0.160.0 on WSL** owns the session and executes shell/file tools. - Hardware: i9-14900K, 64 GB RAM, RTX 4070 Ti SUPER 16 GB. - Codex reasoning effort: `medium`. - Per-request defaults for that effort: `temperature=0`, `top_p=1`, `top_k=20`, `reasoning_budget_tokens=4096`, `max_tokens=16384`. - Speculative decoding setting: `--spec-min-p 0.7` (this is not the sampling `min_p` setting). - A separately trained **expert-cache profile** was loaded, with saving disabled during evaluation. This changes expert placement/profile data, not model weights. It is a relevant condition, not an identified cause; no matched default-profile reproduction of this loop has been established. - Codex Responses requests pass through our local adapter to Strata Chat Completions. Ordinary tool calls are returned to Codex and executed by its harness. This observation has not been isolated to a direct-Strata minimal reproducer, so the adapter/harness remains part of the reproduction conditions. ### Reproduction scope and distinction from other reports This is **one directly observed failed instance**, not a measured failure rate or a deterministic minimal reproducer. The original task was a substantial multi-file review requiring a structured report and validation. Full private session logs and task contents are not attached; the counts above were checked against those logs and preserved drafts. The behavior resembles the repeated verification/tool calls reported in #602. It differs from #606's persistent single-character output on subsequent independent requests, and from #702's repeated sentences/thinking in ordinary chat. We have not established a shared root cause. Since this run used greedy sampling, it also does not establish whether the recommended sampled thinking settings prevent this behavior. ### Questions for mitigation 1. Does Strata have a supported repeated-tool/no-progress guard for **client-owned tool loops across requests**, or a supported configuration we have missed? The v0.1.38 engine-silence watchdog and server-owned MCP `max_rounds` limit appear to address different boundaries. 2. Is there a recommended inference configuration or model/quantization comparison to distinguish model behavior from an inference-path problem? We would prefer a controlled comparison to assuming that token-level repetition penalties solve semantic repetition. 3. Would an optional diagnostic/guard be appropriate, or should this live entirely in the calling harness? A per-request thinking/output cap was already enabled here, but the next request receives a fresh budget. Our current fallback is a task deadline; possible earlier intervention is to detect repeated identical calls with unchanged failure evidence, send one warning, then interrupt if no recovery occurs. A warning produced partial recovery in this instance, but did not produce completed delivery. We are therefore reporting both the recovery and the remaining failure, without claiming that reminders alone prevent recurrence.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.