Pull requests / #869

#869 serve: opt-in recovery from repeated reasoning passages

closed · draft · @huangserva · 0 comments · View on GitHub

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsLinux

Description

A high-effort IQ3_S request can repeat whole reasoning passages until `max_tokens` is exhausted and return no final answer. The 0.1.39 guard from #606 ends runs of one identical token; it does not detect a repeated sentence or a repeated code-verification pass. This draft adds an opt-in, one-shot recovery policy for those passages. It does not claim to fix numerical corruption or replace #838.

Set `"reasoning_loop_recovery": true` in the model config (default: false). At clean parser and UTF-8 boundaries every 512 output tokens, the server checks whether at least 25% of the most recent 2,000 word/punctuation units belongs to 12-unit passages seen at least three times in the reasoning history. It stops and drains the current generation and resumes with every generated token ID preserved. Only the native high-effort hint in the initial system message changes to the template's low-effort hint. No final answer, synthetic reasoning, or `</think>` is inserted; both passes share the original completion limit.

On recovery only, temperature is raised to at least 1.0 and presence penalty to at least 1.5, with top-p 0.95/top-k 20 and the same seed. This deliberately overrides an explicit low-entropy client sampler and is documented as opt-in. An already-matching sampler is passed through unchanged. History and request traces record recovery count and repetition coverage. The 0.1.39 single-token guard stays enabled.

Validation:

- CPU regressions cover the single-token guard's passage-loop gap, unchanged user task and generated IDs, unchanged matching sampler, old explicit sampler recovery without request mutation, shared output limit, one-shot behavior, cancellation, repeated final text, missing native hint, and batch status cleanup.
- `python -m unittest discover -s serve -p 'test_*.py'` on base `6f32ec070f23ced9f50e704d854d775da52591ab`: 281 tests run, 274 passed, 7 skipped (no GPU/model fixtures). The optional jsonschema package was installed for schema-validation tests. Includes 13 new recovery tests.
- v0.1.39 GPU control, built for Linux CUDA sm_89 with pinned llama.cpp `3cf03257f219afbe7334045ff7c6a06ac68c627d`: the captured original IQ3_S failure prefix produced **1,487 consecutive repeated passages**, reached the **16,384-token** cap, and never closed thinking or delivered a final answer. This is a multi-token loop, so the server's 256-identical-token guard does not match it. **Recovery enabled on that same loaded engine, input, seed and cap** detected coverage 1.0 at 2,048 generated tokens, recovered once, and finished naturally with `stop` at **9,629 total completion tokens**, with a nonempty final answer (SHA-256 `5a35536b34032fce0094f6453067efb6952f2ae15a6620cbc59f863c5f188cad`). The repeated prefix is retained, including 187 phrase repetitions generated before recovery. This is one diagnostic replay, not a new full-quality benchmark on 0.1.39.
  - `STRATA_NO_LARGEPAGES=1` was used in both arms, after the default large-page startup was manually aborted before READY due to slow allocation. This switch changes page backing, not model arithmetic. No engine source was modified.
- Prior 0.1.38 hardware verification: an actual captured failure prefix with the same loaded engine, seed and 16,384-token cap produced 1,487 consecutive repeated passages without recovery and a native `stop`/nonempty final answer at 6,655 completion tokens with one recovery. A frozen 31-request cohort then returned 31 complete final answers. [Protocol, limitations and public answers](https://github.com/huangserva/model-evaluation/blob/strata-iq3s-loop-recovery-20261004/strata-iq3s-2026-10-04/fix-verification.md).

Limits: this is a heuristic policy and may trigger on useful repeated code/checks. It supports the exact native high-effort hint in the first system message, skips images, and can require re-prefilling the modified prompt. Other templates/effort positions are not validated. A resumed model can still loop or reach the shared output limit. The 31-request cohort was run on 0.1.38; the 0.1.39 GPU check above is one diagnostic pair, not evidence that every 0.1.39 request completes. The draft is open for feedback on the opt-in API, threshold, and sampler policy before broader use.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.