Issues / #1139

#1139 GDN batch path skips activation quantization when STRATA_QFUSE=1

closed · @pepuscz · 5 评论 · 在 GitHub 查看

BenchmarksModels & quants

描述

Static inspection of the public upstream source shows that the batched GDN path can skip writing its Q8 activation buffer when `STRATA_QFUSE=1`.

Affected source: [src/core/verify.cpp at 1735d6471df29b42c26170efaac1f1446a58640f, lines 872–887](https://github.com/Niko1221/Strata/blob/1735d6471df29b42c26170efaac1f1446a58640f/src/core/verify.cpp#L872-L887). The same control flow is present at [current main 82f46a8c8f475f001ad76d92f58f4a4f8ffb0253](https://github.com/Niko1221/Strata/blob/82f46a8c8f475f001ad76d92f58f4a4f8ffb0253/src/core/verify.cpp#L873-L888).

1. With `batch_rec_`, each slot calls `gdn_step_norm_multi` without the optional fused Q8 output pointer; it defaults to null.
2. The solo branch supplies `xq_` when `g_qfuse()` is true.
3. After both branches, `native_quantize_q8_1` runs only when `!g_qfuse()`.
4. The output `native_mmvq` then consumes `xq_`.

For `batch_rec_ && g_qfuse()`, neither the recurrence call nor the explicit quantizer writes this projection's activation into `xq_`. It can therefore consume an earlier activation. This appears to be a correctness defect, rather than permissible numerical rounding drift.

A possible correction is to run the explicit quantizer when `batch_rec_ || !g_qfuse()`, or pass a correctly offset fused output buffer for each slot. Neither proposed source fix has been tested here. A focused regression should exercise QFUSE-on batched GDN and check that the output projection receives the quantized current GDN result.

The configuration workaround is `STRATA_QFUSE=0`. This report is limited to public-source analysis; it includes no private benchmark logs or deployment configuration.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。