Issues / #606

#606 After a 36,689-token request at 155k, every later request answers one repeated token until the engine reloads

open · @66419118nnn · 7 comments · View on GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsWindowsLinux

Description

### Summary

One request generated 36,689 tokens at a 155,763-token prompt. After it, **the server stayed up, `/health` stayed `loaded: true`, decode stayed at 157–183 tok/s — and every subsequent request answered with a single character repeated**, including a 20-token question. Sampling parameters make no difference at all: six requests with `temperature` 0 / absent / 1.0, `repetition_penalty` 1.05 / 1.3 and `penalty_last_n` 64 / 4096 produced **byte-identical** output (`!!!!…`, 1,500 of them). Only stopping the service and letting the engine reload its experts fixed it.

This looks like in-process state that survives the end of a request (here: a request the client aborted), keeps feeding a constant token to *later* requests, and is invisible to every health check we have.

### Environment

- engine **0.1.35** (prebuilt Linux / CUDA 13), `serve/server.py` from the same tag
- RTX 5090 D 32 GB, driver 595.91.07; Ryzen 9 9950X3D; 64 GB kit (OS sees 59.2 GiB)
- **resident low-RAM mode**, `--resident-experts`: 39.82 GiB of experts page-locked, `0 blob reads from the file`
- model: Qwen3.8-Flash-Next **IQ3_S** (original, GSQ-RCO), `--max-context 224000 --kv int8 --prefill auto:32768 --spec 3 --spec-min-p 0.5 --mtp …/rt --expert-cache auto --vision --vram-reserve-mib 700 --api-key …`
- run config `sampling` block: `temperature 0.6, top_p 0.9, top_k 30, min_p 0.05, repetition_penalty 1.05`
- one slot in play, `--conversation-cache-mib 0` (off)

### How it unfolded (serve log, the same process throughout)

```
prompt 153,171 … 876  generated in 6,095 ms (143.7 tok/s), drafts accepted 596 of 688
prompt 154,112 … 958  generated in 8,309 ms (115.3),         drafts accepted 445 of 643   hit rate 92.9%
prompt 155,121 … 594  generated in 3,961 ms (150.0),         drafts accepted 429 of 472   hit rate 88.6%
prompt 155,763 = 155,715 reused + 48 read, 36,689 generated in 232,175 ms (158.0),
                       drafts accepted 18,441 of 18,551                                   hit rate 99.5%
        suffix drafts: 18,069 windows, 18,296 of 18,316 drafts accepted
prompt 156,132 … 780  generated in 4,957 ms (157.3),         drafts accepted 388 of 388   hit rate 99.6%
prompt 156,019 … 348  generated in 2,210 ms (157.4),         drafts accepted 172 of 172   hit rate 100.0%
        resident RAM: 39.82 GiB in RAM, exchanged with the VRAM tier: +3,039 → +85 → +8 per request
```

The client aborted that 232-second request (`finish` recorded as a disconnect). From then on everything collapsed, at any prompt size:

| prompt | request | output |
| --- | --- | --- |
| 23 tokens | plain question, sampling default (0.6) | 1,500 × `!` |
| 23 tokens | `temperature: 0` | 1,500 × `!` — **identical text** |
| 23 tokens | `temperature: 1.0` | 1,500 × `!` — **identical text** |
| 23 tokens | `repetition_penalty: 1.3, penalty_last_n: 4096` | 1,500 × `!` — **identical text** |
| 8,192 / 65,536 tokens | same task | 1,500 × `!` |

After `systemctl --user restart` (engine reloads 40 GiB of experts): the same 23-token, 8k and 64k prompts all answer normally, drafts accepted back to 57–84%, longest same-character run 1–2.

### The tell I would like your read on

Going from `temperature: 0` to `temperature: 1.0` cannot produce byte-identical text unless the sampler is no longer what picks the token. So whatever is stuck sits upstream of sampling — yet it is not per-request, because a brand-new request with a 20-token prompt inherits it. Candidates I can see from the log: the resident expert mapping in `src/core/expert_source.cpp`, and the prompt path's borrowed cache slots ("the prompt path borrows 7,233 CUDA0 cache slots"), with that request sitting at 6 checkpoints.

### What I could NOT reproduce (each ended with a correct answer)

- 150k-token prompt + `max_tokens 40000` on a long list task → wrote ~7k characters and stopped at `end_turn`
- aborting **mid-decode** at 150k twice (~700 KB of SSE streamed, then the socket closed)
- aborting **during prefill** at 64k three times

So the trigger looks content- or history-dependent and I cannot hand you a small reproducer. What I can hand you is a cheap observable: **longest run of one repeated character ≥ 20 in the answer**. We now run a canary every 5 minutes on that rule and restart when it fires.

### Two asks

1. **Which in-process state can outlive a request (or a cancelled one) and pin the output to a constant token?** #481 (a silent engine) is the neighbouring failure, but nothing went silent here — the engine kept answering at full speed with the same character, so the silence watchdog would never fire.
2. **A degeneration guard inside the engine?** Re #588: I read it, and the `--pcie-frac` denominator artifact does not cover this — in resident mode there were `0 blob reads from the file`, and the reported hit rate rose to ~100% while decode stayed at 157–183 tok/s rather than falling. My reading is that `hit rate ≈ 100% + drafts accepted ≈ 100% + one character repeating` is a *symptom* triple: a constant token stream makes every expert lookup and every suffix-lookup draft trivially right. If you agree, that triple would make an inexpensive in-engine breaker — with `--suffix-draft` on, a collapse is fast (158 tok/s) and, because a client that omits `max_tokens` gets `max_new = room` (66,967 here), it can run for four minutes before anyone stops it.

### Side note, not the bug

`serve/server.py` gives a request that omits `max_tokens` the whole remaining context as its runway. `POST /settings` can cap it, but that also changes the default for every client that leaves the field out, so an engine-side "same token N times in a row → stop" would be the better knob.

Happy to test candidates on this box (5090 D 32 GB / 64 GB RAM / IQ3_S resident / 224k ctx), and to re-run this on a newer engine — we are on 0.1.35 and I see 0.1.36–0.1.38 landed since, including the #463 greedy-determinism change.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.