Issues / #879
#879 All-logits-NaN degeneration (one repeated token forever) still fires on 0.1.39b / current main — first poison traced to the MoE output rows: garbage bits, varying layer
open · @66419118nnn · 5 コメント · GitHub で見る
Server & APINVIDIA / CUDAModels & quants
本文
Follow-up to #606 with a working statistical reproducer and a full instrumentation suite. **TL;DR: the degeneration survives 0.1.39 (both q8_1 clamps + `repeat_stop_tokens`); it also survives disabling the unclamped fused-SwiGLU quantizer that rwkeyes identified below. Layer-trace instrumentation on current main (`6f32ec0`) shows the first non-finite value of a window being produced in the MoE path — of the 13 fires with per-layer traces, the first poisoned signal is the routed-expert rows in 11, the attention block's output of layer 28 in one (with that layer's input, GDN state and PLE all clean), and the layer's combined contribution only (a 2-signal early build) in the last — while the layer input, GDN/QSA state, PLE history, KV, CPU-pool output and the expert-cache slot bytes are clean in every fire. In all 6 fires that recorded raw bits, the GPU-side value is the same non-canonical NaN payload 0x7FFFFFFF; the miss side recorded garbage payloads twice (0x7FD6E000, 0x7FD3E000) and a canonical arithmetic NaN once (0xFFC00000).** This is a data-path defect, not an overflow, so the clamp family of fixes cannot close it.
## Environment
- Engine: current `main` `6f32ec0` (0.1.39b), built locally (CUDA 13.0, sm_120) — and, for the earlier half of the fires, 0.1.35 with an equivalent local patch set. The signature is identical across both (300 commits apart, including the #606 fix).
- GPU: RTX 5090 D 32 GB (compute capability 12.0), memory clock at stock (14 001 MHz) since 2026-10-03 — every fire in the 21-fire set postdates that; ECC is not available on this board.
- CPU: Ryzen 9 9950X3D, 64 GB RAM; 15 expert-pool workers.
- Model: Qwen3.8-Flash-Next, Unsloth UD-Q4_K_XL (GGUF-in-place, native pack), `--expert-cache auto` (6 493–6 714 of 24 576 experts resident in VRAM across boots, ≈ 26–27%), `--resident-budget-gib 53` (clamped to 50.5–51.6 GiB page-locked complement, whatever the boot leaves free), `--spec 3`, suffix drafts on, `--pcie-frac 0.55`, `--kv int8`, `--prefill auto:32768`, ctx 224 000.
- Tier-independence: the same fires were produced on UD-IQ4_XS (4 fires) and, per the Q8 report in #606's thread, on someone else's Q8 — this is not a quantization-tier artifact.
## Reproducer (statistical: 1 fire per 1–18 replays; every completed 30+-try round has fired at least once)
A real agentic request whose window failed (85 688-token prompt, 47 messages, 26 tools, `reasoning_effort xhigh`), captured verbatim at guard failure by a serve-layer hook in our local patch, then replayed with fresh seeds:
```
STRATA_API_KEY=… python3 replay-degen.py 20261004-211527-request.json \
--tries 40 --max-tokens 3000 --model <tier> --strip-images
```
(`replay-degen.py` is in the gist; the capture contains my private conversation, so I will share it directly on request rather than attaching it. It reproduces on this 5090 D with a partial-residency tier; not yet verified on other cards.)
Every observed ignition in long-context traffic lands at positions 77 261–88 443 — in every fire whose prompt length is known, within ~3 000 decoded positions after prefill — plus one short-context fire at position 311, three minutes after a cold start. On four 0.1.39b rounds today the fires landed at try 6, 18, 1 and 9 of 40; earlier 0.1.35 rounds hit ~1 fire in 15–20 tries. Deterministic per-seed reproduction does **not** survive an engine restart — the trigger needs process-level state (cache residency/swap history), which is why rounds must be Monte-Carlo.
## What the instrumentation is (all of it is in `local-patches-20261005-0139b.diff`, ~440 lines on top of `6f32ec0`)
1. **A logits finiteness guard** (`logits_nonfinite` in `sampler.cu`, enqueued into the captured verify-window graph, flag read after the sync the window already performs — measured zero cost): a window whose logits are non-finite fails instead of being sampled, so the symptom becomes a clean request error instead of 9 613 consecutive `!`.
2. **Seven per-layer poison flags baked into the window graph** (`lay_flags_` in `verify.cpp`): layer input (folded residual), attention block, shared expert, routed total (post-fold), combined output, GPU cache-hit rows (`hit_out_`), CPU miss rows (`parts_` pre-fold).
3. **First-non-finite recorders** (row, col, raw bits) for the miss side and the hit side.
4. **A planned-row audit** (`audit_hit_rows`): scans only the rows in this layer's GPU plan (`p_dst[0, counts[1])`), so a stale row the engine never reads cannot masquerade as poison.
5. **A kernel-stage audit** (`STRATA_KERNEL_AUDIT=1`, in `native_expert_grouped`): scans the gu (gate/up) outputs, the SwiGLU output, and the down-projection output.
6. (0.1.35 build additionally: slot-vs-file memcmp of the poisoned expert's cache bytes, and a host-side scan of the CPU pool's rows at publish time.)
## Evidence
**Fire A (0.1.39b + guard only, try 6):** position 88 443 — fires with the owner's full 0.1.39 fix in place.
**Fire B (0.1.39b + `STRATA_FUSED_SWIGLU_Q81=0`, try 18):** position 87 353 — the third unclamped quantizer is not the cause either.
**Fire C (0.1.39b + layer trace, try 1):** position 87 611 — *(log lines verbatim; the ← annotations are mine)* —
```
layer trace: first poisoned input/attn/shared at layer 14 (GDN) ← downstream
layer trace: first poisoned routed total / combined / GPU cache-hit at layer 13 (GDN) ← FIRST
layer trace: first poisoned CPU miss rows at layer -1 ← clean the whole window
first non-finite HIT value: row 26 (token 2, expert slot 6), col 2092, bits 0x7fffffff
mapped miss rows: 0 of 76800 non-finite
```
**Fire D (0.1.39b + everything, try 9):** position 85 965 —
```
layer trace: first poisoned routed total / combined / CPU miss rows at layer 37 (GDN) ← miss side FIRST
layer trace: first poisoned GPU cache-hit / input / attn / shared at layer 38 (GDN) ← downstream
first non-finite MISS value: row 17 (token 1, slot 7), col 2144, bits 0xffc00000 ← canonical -qNaN
first non-finite HIT value: row 15 (token 1, slot 5), col 736, bits 0x7fffffff ← garbage payload
kernel-stage audit: first non-finite in gate (gu kernel) of layer 47, flat 6848, bits 0x7fffffff
planned-GPU-row audit: layer 38, row 15 (token 1, slot 5), col 2176, bits 0x7fffffff, expert 5 ← in the plan
```
**From the 0.1.35 instrumentation (13 fires with dumps, same shapes):**
- The poisoned GPU row is **inside the plan's dst table** (the engine consumes it) in both fires where the planned-row audit ran — the stale-row objection is ruled out.
- The poisoned expert's cache slot is **byte-identical to the pack file** (slot 4 203, layer 21, expert 1) — the blob arrived fine; the poison is made after it lands.
- The CPU pool's rows at publish time are clean in every hit-side fire.
- GDN state (36 layers), PLE history (92 160 values), layer inputs up to the first poisoned layer: all finite.
**Across 21 fires (0.1.35 + 0.1.39b):** the first poisoned layer varies — {3, 13, 15, 16, 18, 20, 22, 26, 28, 37} — with GDN and QSA layers both appearing, and both sides of the hit/miss split having fired first (the routed-expert rows are the first poisoned signal in 11 of the 13 fires with layer traces; one, position 85 683, has the attention block's output of layer 28 first, with that layer's input, the GDN state and PLE all clean — so the defect is not exclusive to the expert GEMMs). What is invariant: the GPU-side first non-finite value is the same garbage payload 0x7FFFFFFF in all 6 fires that recorded bits, and everything else — weights, cache slots, pool output, persistent state — is clean at the moment of the failure dump.
## What is ruled out, and what is left
- It is **not** the q8_1/q8_0/KV fp16 overflow family: our build had those clamped before 0.1.39 existed, fires continued; 0.1.39b (official clamps) fires; the fused-SwiGLU site disabled still fires.
- It is **not** precision-tier, cache-slot corruption, stale-row instrumentation, GDN/PLE/KV state, the AVX-512/AVX-2 multi-token CPU kernels (forcing the ggml single-token dot path fires on try 1), or sampling.
- The first non-finite value appears **in the routed-expert pipeline's output rows in 11 of the 13 traced fires — and in the attention block's output of layer 28 in the other fully-instrumented one — at a varying layer, with raw garbage bits on the GPU-computed side** — i.e. bytes that are not the result of any float arithmetic in that path. Where exactly the garbage bytes enter (the gu/down GEMM's output write coverage vs. the plan's entries, scratch-buffer reuse across graph replays, the `ent_dst`-indexed write, or — for the miss-side fires — the host-publish → mapped-memory read-back of the pool's rows) is what the instrumentation localizes to; I have not been able to close that last inch from outside the kernels.
## Suggested places to look (in `native_expert_grouped` and the window graph)
1. **Output-row coverage vs. plan entries in `launch_gu` / `launch_down`**: the GEMMs write `gate`/`up`/`out` rows for the plan's entries; if any launch's row range or `ent_dst` base assumes `counts[1] == cap` (or strides by `cap` while the plan is shorter), it writes or reads rows the plan does not own — which is exactly where a stale-scratch 0x7FFFFFFF can be laundered into a consumed row.
2. **`hit_scratch_` / `nat_xq_` reuse across CUDA-graph replays** with varying window T (spec 3 + suffix drafts means T cycles 1–5 and the graphs are instantiated per T): the stage audit caught garbage in the gu output of a *late* layer, i.e. bytes surviving in scratch from earlier in the window or from a previous window.
3. **The miss-side path** (`expert_pool_dispatch_multi` → host publish → `copy_rows_from_mapped`): the host's rows are finite at publish and at failure time, yet the pre-fold miss row held 0xFFC00000 — worth checking the flag-ring (`m_flag_`/`wait_flag_ge`) sequencing under graph replay for the group-skip (`skip_ + grp`) cases, and the pool's memset of GPU-kind rows (`expert_source.cpp`, the `kind[i] >= 0` branch).
4. The `0x7FFFFFFF` payload specifically: four 0xFF bytes. In a q8_1 block that is the pattern of `ds = {inf-ish, inf-ish}` written as raw bytes, or a half2 sentinel — a hexdump of the 16 bytes around the recorded (row, col) in `gate`/`hit_out` would tell which.
## Mitigation we run in production (offered as-is)
The verify-window finiteness guard from the attached diff: a non-finite window fails cleanly (one bad request instead of minutes of `!` spam), the engine exits, and the failure dump lands in the engine log with position/row/bits. Together with `repeat_stop_tokens` it bounds the damage; it does not, of course, cure the root cause.
Everything referenced above is in this gist: **https://gist.github.com/66419118nnn/7c9399d982d98229c61fcefaaa0b9215** — the full experiment report (Chinese), `fires.json` with all 21 fires and their verbatim dumps, `replay-degen.py`, both patch sets (0.1.35 and `6f32ec0`), and the four round logs.
関連リンク
インストール・モデル・リリースへの站内リンク。