Issues / #1560

#1560 Field data: the greedy/tool-loop issue measured with a replay harness built from real production loops — and a layer-by-layer view of where each mitigation actually helps

open · @liqiang74 · 0 Kommentare · Auf GitHub

Setup & installModels & quants

Beschreibung

Hi — thank you for #728 and for your Oct 7 diagnosis there ("the cause is greedy default sampling when the client sends none"). That sentence is the actual fix; everything below is just our attempt to measure it properly on our own stack and to share a method other users could reuse. Please read the data as a modest field report, not a claim of anything new.

**Our setup**: Qwen3.8-Flash-Next IQ3_S, Strata v0.1.40.2, single Tesla V100, behind an agent harness (OpenClaw) that talks to the model through tool calls. Our workload is tool calls, so we had also avoided global sampling defaults in the past — which is exactly the trap #728 describes.

### What we think is the one useful contribution: a replay harness from *real* loops

Most loop reports (including #980, #710, and ours earlier) use synthetic induction — a tool that always fails, repeated prompts, etc. Those are easy to dismiss as artificial. So we built the test set out of **loops that actually happened in production**:

- Source 1: our middleware's recovery archive — 17 real quarantined loop cases (12 of them `rule=loop`), each carrying the real task goal, the real repeated command and the real tool result.
- Source 2: our agent session store — 20 real loops found by scanning recent sessions for runs of identical consecutive tool calls (longest streak 37).
- `build-real-corpus.py` extracts "the complete real prefix up to the point the loop started"; `realprefix-replay.py` feeds that prefix back and lets the model continue, with a fixed return per turn (the same semantics that produced the loop).

The value of this is the **two-way anchor**: the old behavior reproduces the *real* production loop (so the sample is genuine, not staged), and the new behavior suppresses it (so the fix is real). A synthetic loop only shows the second half.

### Result (same 16 real cases, loop = max streak ≥ 4 identical calls)

| # | Path | Sampling | Loops reproduced | Mean longest streak |
|---|---|---|---|---|
| a | engine direct | greedy (old) | **11/16 (69%)** | 6.5 |
| b | engine direct | config `sampling {1.0/0.95/20}` | **0/16 (0%)** | 1.1 |
| c | engine + our middleware | greedy | **11/16 (69%)** | 3.2 |
| d | engine + our middleware | config defaults | **0/16 (0%)** | 0.8 |

Cross-checked with a separate synthetic harness (20 runs/arm): config defaults cut loop *onset* ~90% (guard engagement counter 41 → 4), consistent with your #728 conclusion.

### The layer-by-layer part, stated honestly

- **(b vs a)** is the whole story for prevention: the published Qwen sampling defaults in the config remove the onset almost entirely. That is your recommended fix, and it holds on real cases, not just synthetic ones.
- **(c vs a)**, our middleware, **does not reduce the reproduction rate at all**. It only reduces the *severity* (mean longest streak 6.5 → 3.2, with several 12-streak locks cut to 4) and can hard-stop a stuck session to let the harness recover. In other words, a middleware guard is a damage-limiter, not a cure — and it is no substitute for the config defaults. We say this because we previously confused the two ourselves and spent time wondering why "everything was on" and loops still appeared.

### If it is ever useful: the cross-request guard design (not a request to merge)

Your `reasoning_loop_recovery` (c3473262) is per-reply and, as you noted in #710, cannot see a loop that spans requests. If you ever want to cover that family too, here is the minimal thing we tried, in the same opt-in style — offered only as a reference, not as something you should take:

- At the sampling merge point, scan the visible history for a trailing run of identical consecutive tool calls (name + key-sorted arguments).
- On streak ≥ K (we use 4), merge `repetition_penalty=1.1, penalty_last_n=512` **only if the client sent no penalty of its own**; never add temperature/top_p; never rewrite the request bytes (injection happens after tokenization, so caching and pass-through stay intact).
- Off by default; one config block; one-key rollback.
- Measured on our harness: 11/20 → 19/20 converged, Fisher exact p≈0.008, zero false triggers on healthy traffic, 90/90 tool-call JSON still parses. On the real corpus it lowered severity but, as above, not the reproduction rate.

### Limits (so the numbers are not over-read)

16 real cases, one model/tier (IQ3_S), one GPU (V100), one harness — this is *mechanistic* evidence, not a statistical law. The penalty of 1.1 and K=4 are local sweet spots, not universal optima. Shared in the hope the method (real-loop replay + layer attribution) is more useful to others than our particular numbers.

Happy to share the full corpus, the raw JSONL and summary CSVs if anyone wants to reproduce it.

Mehr auf der Site

Links zu Install, Modellen, Releases.