Pull requests / #318
#318 iq_pack: store an F32 ple_conv1d as F16
closed · @cripto-bot · 0 评论 · 在 GitHub 查看
BenchmarksAMD / HIPNVIDIA / CUDAModels & quants
描述
Reaching the same root cause as #303, independently, on a second sm_70 rig (Tesla V100) with the same artifact family (Unsloth Qwen3.8-Flash-Next-UD-IQ3_XXS).
`blk.1.ple_conv1d.weight` is F32 in that GGUF, but the PLE conv kernel consumes the tensor as F16 bits (`ss.ple.w.conv1d_f16 = (const uint16_t*) wc->data`). `iq_pack` passed the F32 source through as index kind 2 ("copy 4 bytes per element"), so the kernel read F32 bytes as packed F16, the conv output came out ~500x too large and every generated token was noise. The official GSQ-RCO artifacts ship this tensor as F16 (kind 5) and are unaffected - which is why the bundled tests (synthetic F16 weights) cannot see it.
This PR fixes the pack side: when the source is F32, convert to F16 (same nearest-even rounding as the existing BF16 path) and write kind 5 (81,920 bytes instead of 163,840). It does not help packs already built from an F32 source - the load-time re-round suggested in #303 covers those; the two are complementary.
Validated end to end on the V100 with the fixed pack: "The capital of France is" -> " Paris. The capital of Germany is Berlin..." (was "zocz6wedawyzadsowjhcwways w" before), 15-30 tok/s decode with MTP spec 4 (60-75% draft acceptance), ~40 tok/s prefill, and the repo's parity checks pass with this artifact's data.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。