Pull requests / #290

#290 Token embedding in BF16 from the checkpoint (--embd-gguf)

closed · @sergqwer · 0 Kommentare · Auf GitHub

AMD / HIPNVIDIA / CUDAModels & quantsWindows

Beschreibung

## Summary

The native path reads the token embedding from the model GGUF, where it is quantized (IQ4_XS in ISTA-DASLab's
IQ2_XS). The checkpoint ships it in BF16.

- `tools/embd_bf16_pack.py` copies `embed_tokens` into a one-tensor GGUF, bytes unchanged. Only the safetensors shard
  that holds it is needed: 3 GB for `Qwen/Qwen3.8-Flash-Next` (`model-00130-of-00131.safetensors`).
- `--embd-gguf PATH` makes the native embedding read `token_embd` from that file. `iq_embed_rows` and
  `iq_dequant_f32` decode BF16 (`embed_type_supported`).
- The table stays in mapped host memory, as before: 1.2 GiB instead of 0.3, and no VRAM.

## Measured

Measured on current main (0.1.29, d6708a4) with ISTA-DASLab's IQ2_XS, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB DDR5, Windows 11. KL is first-token KL divergence (`STRATA_DUMP_FIRST_LOGITS`, #276); two runs of main itself differ by 0 on 5 of 6 short prompts and by 0.00013 on the sixth.

First-token KL of the GGUF's IQ4_XS embedding against the BF16 one: 0.0015 (1K), 0.0089 (2K; the top token changes), 0.0010 (4K), 0.037 (32K); mean 0.012, ~100x the run-to-run noise. On the NVFP4 fork, whose GGUF embedding is Q8_0, the same comparison gave 0.0025. Decode and prompt speed are unchanged (fork); the start takes ~0.2 s more.

## Switch

Off unless `--embd-gguf` is given.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Mehr auf der Site

Links zu Install, Modellen, Releases.