Pull requests / #838

#838 #606: the fused SwiGLU q8_1 quantizers keep the block's scale and sum finite too

closed · @sergqwer · 0 comentários · No GitHub

NVIDIA / CUDAModels & quants

Descrição

Completes the #606 clamp. 0.1.39 made two q8_1 activation quantizers keep the block's fp16 scale and sum finite: `q8_1_store` in iq_kernels.cu and `native_quantize_q8_1_kernel` in native_mmvq.cu. Two fused SwiGLU + q8_1 kernels still store the block raw:

| kernel | where it runs |
|---|---|
| `native_swiglu_quantize_q8_1_kernel` (native_mmvq.cu:166-169) | the shared expert, decode (shared_expert.cu:203) and the multi-token pass (:310); on by default (`STRATA_FUSED_SWIGLU_Q81` unset) |
| `native_gu_fused_kernel` (iq_kernels.cu:2416-2420) | the routed experts' fused gate/up on gfx906 builds with `STRATA_EXP_MODE=8` (IQ types 18/21/22/23) |

Both now go through `q8_1_finite` / `q8_1_quant` / `q8_1_ds`, the same three lines as their siblings. Every block that was finite before is stored bit for bit as before, so normal output does not change.

rwkeyes found the first one in #606 but has no NVIDIA card to test it. The second one turned up in a sweep for `amax / 127.0f` after it; nothing else in `src` stores a q8_1 block without the clamp now.

### Test

`iq_multi_parity`'s #606 check now runs `native_swiglu_quantize_q8_1` as a third path. It uses gate = 32 and up = x / 32: `silu(32)` is exactly 32 in float, so the fused kernel sees the same activations x as the other two paths and is held to the same expectations.

RTX 5090, CUDA 13, sm_120:

```
# this branch
q8_1 finite (quantize_q8_1_rows): 237 finite blocks as before, 3 overflowed blocks clamped finite
q8_1 finite (native_quantize_q8_1): 237 finite blocks as before, 3 overflowed blocks clamped finite
q8_1 finite (native_swiglu_quantize_q8_1): 237 finite blocks as before, 3 overflowed blocks clamped finite
iq_multi_parity: 0 failures

# the same test with native_mmvq.cu from v0.1.39
q8_1 finite (native_swiglu_quantize_q8_1): 237 finite blocks as before, 0 overflowed blocks clamped finite  FAIL
iq_multi_parity: 4 failures
```

The gfx906 kernel is not covered: I have no gfx906 card to run it on. Its change is the same three lines as `q8_1_store` in the same file.

### Data point

On a fork of 0.1.39 (RTX 5090), a Qwen3.8-Flash-Next served conversation also turned into '!' (token 0) mid-thinking at about 99K positions, and later requests carrying that history kept producing it. Replaying the conversation up to that point did not reproduce it. So I cannot tie that run to this kernel; it is only the same symptom #606 describes, on a build that still had this store unclamped.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

No site

Links install, modelos, releases.