Pull requests / #318

#318 iq_pack: store an F32 ple_conv1d as F16

closed · @cripto-bot · 0 comments · View on GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quants

Description

Reaching the same root cause as #303, independently, on a second sm_70 rig (Tesla V100) with the same artifact family (Unsloth Qwen3.8-Flash-Next-UD-IQ3_XXS).

`blk.1.ple_conv1d.weight` is F32 in that GGUF, but the PLE conv kernel consumes the tensor as F16 bits (`ss.ple.w.conv1d_f16 = (const uint16_t*) wc->data`). `iq_pack` passed the F32 source through as index kind 2 ("copy 4 bytes per element"), so the kernel read F32 bytes as packed F16, the conv output came out ~500x too large and every generated token was noise. The official GSQ-RCO artifacts ship this tensor as F16 (kind 5) and are unaffected - which is why the bundled tests (synthetic F16 weights) cannot see it.

This PR fixes the pack side: when the source is F32, convert to F16 (same nearest-even rounding as the existing BF16 path) and write kind 5 (81,920 bytes instead of 163,840). It does not help packs already built from an F32 source - the load-time re-round suggested in #303 covers those; the two are complementary.

Validated end to end on the V100 with the fixed pack: "The capital of France is" -> " Paris. The capital of Germany is Berlin..." (was "zocz6wedawyzadsowjhcwways w" before), 15-30 tok/s decode with MTP spec 4 (60-75% draft acceptance), ~40 tok/s prefill, and the repo's parity checks pass with this artifact's data.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.