Pull requests / #1448

#1448 #606 follow-up: clamp the three remaining unclamped q8_1 scale/sum emit sites

open · @MatthewHines · 0 评论 · 在 GitHub 查看

Server & APINVIDIA / CUDAModels & quantsDocumentation

描述

## Summary

Three kernels still emit the `(d, sum)` pair of a q8_1 activation block with a raw `make_half2(d, sum)`, bypassing the finite helper the rest of the tree has used since `19de7f4` / `f6601bf` (`include/strata/kernels/q8_1_finite.hpp`):

| site | path |
|---|---|
| `fused_gr.cu` — `gr_q8_tail` | fused gate-reduce tail (QFUSE) |
| `verify_kernels.cu` — `gdn_q8_1_store` | batched GDN verify store (QFUSE) |
| `iq_kernels.cu` — `s26_swiglu_q8_1_kernel` | S26 swiglu quantize |

A block whose `amax` overflows (or that reads a non-finite input) emits `d = inf/NaN`; fp16 turns the pair into NaN (or the pathological `(0, 0)`). Nothing downstream tests for finiteness — the dot kernels assume the producer upheld the invariant, because every other producer shares `q8_1_ds()`. A NaN scale poisons the fp16 dot of the quantized activation and every sum downstream of it through the layer stack; a fully non-finite logit row ends at the sampler's all-non-finite fallback — vocabulary id 0, the `!` face of #606.

This is the same defect class already fixed at the other emit sites; these three were missed. The fix is the same one-liner pattern: route the pair through `q8_1_ds()`, which clamps both halves into the fp16-finite range (finite `d`, non-negative finite `sum` — the sum is a mean of `|x|`) and produces for the GPU paths the same bits the CPU pool and the already-clamped GPU sites produce for the same block. Content-blind producer-side clamp; nothing model-text-aware anywhere in it.

## Verification

- `nvcc -std=c++20 -arch=sm_120 -c` clean on all three changed translation units (CUDA 13.4.2).
- Field evidence from a Qwen3-Next-class hybrid deployment (GDN/PLE + sparse experts, conversation cache + speculative decoding): the id-0 (`!`) degeneration face of #606 appears on builds where the fused/S26 paths are enabled and is absent on boots that use the clamped path exclusively — consistent with the site inventory. The wider carried-state failure (restart-cures-it, spans sessions) is analyzed in `docs/DEGENERATION_RCA.md`; it involves two further purity violations of the state cache that this PR deliberately does **not** change code for.

## Scope

3 clamp lines + 2 includes + 2 docs. No behavior, tuning, or policy changes. Deployment-side mitigations we use in the interim (freezing the expert migrator with the existing flag, disabling the suffix-draft source, serve-layer repetition detectors) are listed as prose in the RCA doc so maintainers can route them properly — no code for them here.

Docs added:
- `docs/Q8_1_FINITE_DS.md` — the invariant, the missed sites, diagnostic criteria for this defect class.
- `docs/DEGENERATION_RCA.md` — the full root-cause analysis: the restart-cure and cross-session-persistence localization levers, the three purity violations of `state = f(prefix)`, the falsification tests that pinned each one, a field diagnostic protocol, and recommended engine-side remedies (tier-aware checkpoint identity for the residency split; reproducible policy inputs for the drafting side).

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。