Pull requests / #290
#290 Token embedding in BF16 from the checkpoint (--embd-gguf)
closed · @sergqwer · 0 评论 · 在 GitHub 查看
AMD / HIPNVIDIA / CUDAModels & quantsWindows
描述
## Summary The native path reads the token embedding from the model GGUF, where it is quantized (IQ4_XS in ISTA-DASLab's IQ2_XS). The checkpoint ships it in BF16. - `tools/embd_bf16_pack.py` copies `embed_tokens` into a one-tensor GGUF, bytes unchanged. Only the safetensors shard that holds it is needed: 3 GB for `Qwen/Qwen3.8-Flash-Next` (`model-00130-of-00131.safetensors`). - `--embd-gguf PATH` makes the native embedding read `token_embd` from that file. `iq_embed_rows` and `iq_dequant_f32` decode BF16 (`embed_type_supported`). - The table stays in mapped host memory, as before: 1.2 GiB instead of 0.3, and no VRAM. ## Measured Measured on current main (0.1.29, d6708a4) with ISTA-DASLab's IQ2_XS, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB DDR5, Windows 11. KL is first-token KL divergence (`STRATA_DUMP_FIRST_LOGITS`, #276); two runs of main itself differ by 0 on 5 of 6 short prompts and by 0.00013 on the sixth. First-token KL of the GGUF's IQ4_XS embedding against the BF16 one: 0.0015 (1K), 0.0089 (2K; the top token changes), 0.0010 (4K), 0.037 (32K); mean 0.012, ~100x the run-to-run noise. On the NVFP4 fork, whose GGUF embedding is Q8_0, the same comparison gave 0.0025. Decode and prompt speed are unchanged (fork); the start takes ~0.2 s more. ## Switch Off unless `--embd-gguf` is given. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。