Issues / #1139
#1139 GDN batch path skips activation quantization when STRATA_QFUSE=1
closed · @pepuscz · 5 comentários · No GitHub
Descrição
Static inspection of the public upstream source shows that the batched GDN path can skip writing its Q8 activation buffer when `STRATA_QFUSE=1`. Affected source: [src/core/verify.cpp at 1735d6471df29b42c26170efaac1f1446a58640f, lines 872–887](https://github.com/Niko1221/Strata/blob/1735d6471df29b42c26170efaac1f1446a58640f/src/core/verify.cpp#L872-L887). The same control flow is present at [current main 82f46a8c8f475f001ad76d92f58f4a4f8ffb0253](https://github.com/Niko1221/Strata/blob/82f46a8c8f475f001ad76d92f58f4a4f8ffb0253/src/core/verify.cpp#L873-L888). 1. With `batch_rec_`, each slot calls `gdn_step_norm_multi` without the optional fused Q8 output pointer; it defaults to null. 2. The solo branch supplies `xq_` when `g_qfuse()` is true. 3. After both branches, `native_quantize_q8_1` runs only when `!g_qfuse()`. 4. The output `native_mmvq` then consumes `xq_`. For `batch_rec_ && g_qfuse()`, neither the recurrence call nor the explicit quantizer writes this projection's activation into `xq_`. It can therefore consume an earlier activation. This appears to be a correctness defect, rather than permissible numerical rounding drift. A possible correction is to run the explicit quantizer when `batch_rec_ || !g_qfuse()`, or pass a correctly offset fused output buffer for each slot. Neither proposed source fix has been tested here. A focused regression should exercise QFUSE-on batched GDN and check that the output projection receives the quantized current GDN result. The configuration workaround is `STRATA_QFUSE=0`. This report is limited to public-source analysis; it includes no private benchmark logs or deployment configuration.
No site
Links install, modelos, releases.