Issues / #1753

#1753 Loop guards miss a periodic loop written as one long word: a 24K-digit string cycling an 11-digit pattern passes both #606 and #728

open · @nomishbhardwaj · 0 comments · View on GitHub

Setup & installServer & APINVIDIA / CUDALinux

Description

Neither loop guard can see a loop that the model writes as one unbroken run of digits; only `reasoning_budget_tokens` ended it. A simple periodicity check catches it early, and in a 70-sample replay it never fired on a reasoning that went on to pass.

**Setup:** Strata 0.1.41 (`fb58e0d`), Qwen3.8-Flash-Next IQ2_XS, RTX 3090 24 GB, Linux. `reasoning_budget_tokens` 24576, `repeat_stop_tokens` 256 (the default). Requests go to `/v1/chat/completions` at temperature 0: our eval harness pins greedy decoding on purpose. Your diagnosis in #728 says greedy makes loops likelier, but the blind spot below does not depend on sampling.

**Prompt:** `How many times does the digit 8 appear in the decimal expansion of 9^104? End your reply with exactly: "FINAL: <answer>" and nothing after it.` (Truth: 13.)

**What happened:** in two captured runs, the reasoning wrote one unbroken string of digits: 24,255 characters (95% of the reasoning) and 32,455 characters (96%). Both repeat the 11-digit cycle `58747693387` until the thinking budget closes them, at 24,599 and 32,791 tokens, and both answer 1. A third, earlier run ended the same way (24,599 tokens, answer 1); its reasoning was not saved.

**Why neither guard fires:**
- #606 ends a reply after 256 identical tokens in a row. This tokenizer has no multi-digit tokens, so the cycle is 11 tokens long, and the longest run of one token inside it is 2.
- #728's `reasoning_repeat_coverage` splits text with `\w+|[^\w\s]`, so the whole digit string is ONE word. No 12-word passage can repeat inside it, and the reasoning (about 250 words besides the string) never reaches the 2,000-word window either. Replayed through an exact copy of the function, the coverage is 0.00 at every check in both runs.

**A candidate check (replayed offline; we have not patched the server):** at each look, test whether the last 2,048 characters of the reasoning repeat with some period p ≤ 64 (`w[p:] == w[:-p]`). We replayed it over all 70 reasonings captured in a budget A/B (35 long-arithmetic prompts × budgets 24576 and 32768):

| | current detector | periodic check |
|---|---|---|
| the two digit-string runs | never fires (coverage 0.00) | fires at 4,096 characters (period 11) |
| the two pure loops of one other prompt | fires | fires (periods 8 and 64) |
| total fires | 7 of 70 | 4 of 70 |
| fires on a reasoning that went on to pass | 2: one finished naturally at 14,204 tokens after peaking at 0.46; one was a budget close that landed on the right answer | 0 |

The periodic check complements the passage detector rather than replacing it: it misses three passage loops that are not strictly periodic, which the current detector catches.

In the engine, the same idea could work on token ids as an extension of #606 from period 1 to short periods. For each p ≤ 64, keep a count of consecutive positions where `t[n] == t[n-p]`, and act when any count passes a threshold such as 1,024. That costs O(64) per token and needs no decoding.

A smaller change that keeps the current detector did worse. Cutting long words into 16-character pieces (`\w{1,16}`) caught the 32K-digit run but not the 24K one, because that reasoning still had fewer than 2,000 "words".

Happy to share the 70 captured reasonings (about 1 MB of JSONL) and the replay script.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.