Pull requests / #255
#255 Q8 0 support
closed · @gopinath87607 · 0 comentários · No GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
Descrição
Q8_0: the checkpoint quant the engine did not know
Stacked on #216 (the session carve) - merge that first. Until it merges, this diff shows both changes;
afterwards GitHub shrinks it to the Q8_0 commit alone.
`Qwen3.8-Flash-Next-Q8_0` is the same `qwen4exp` topology as the IQ3_S build, but every expert, embedding and
PLE tensor arrives Q8_0 - a type three hard gates refused. All of this is code: no data ships with the PR
(`.gitignore` keeps `/packs/` and `/mtp/` out), so other users pack their own GGUF locally, the same flow
IQ3_S/OrcaRouter users already follow (docs/ORCA.md).
* GPU EXPERT KERNELS. `Fmt<8>` in iq_kernels.cu (qk 32, 4 ints per block, `vec_dot_q8_0_q8_1`), a `dq_q8_0`
dequantizer, and `case 8` in the grouped gate/up/down switches: `native_expert_grouped` runs Q8_0 experts,
the embedding table is Q8_0, and parity against ggml-cpu's own `vec_dot` on identical weights holds to 1e-7.
* THE PLE READER LOSES ITS 90-BYTE ROW. `ROW_BYTES` becomes a runtime `row_bytes` (90 B at IQ4_NL, 170 B at
Q8_0) everywhere it was compile-time - row cache, page bounds, `issue()` - so the per-layer token table
reads either quant.
* CONV1D, THE ONE REAL BUG. The Q8_0 checkpoint ships `blk.1.ple_conv1d.weight` as F32 (163,840 B) but the
engine's PLE kernel reads F16, so half of every f32 word was decoded as a weight: mean |w| 0.43 against the
true 0.0062, layer-1 routing flipped, and the output was PLAUSIBLE GARBAGE, not a crash. iq_pack.py now
writes that tensor as pack index kind 3 (the loader narrows F32 to F16 round-to-nearest-even; proven
bit-exact against the IQ3_S pack's own tensor), narrowing happens whether or not `--compat-bf16` is given,
and generate.cpp refuses to boot a pack whose conv1d is still wide instead of decoding it as garbage.
* SHARDS BY ROLE. In this split the gate/up/down of layers 17/35 sit in different shards than their
neighbours; the pack index carries a per-role shard column and the engine resolves each role's file, so
`--native` names any shard of the family and the head is auto-located by `output.weight`.
* THE ARENA ASKS THP FOR 2 MB PAGES. `reserve()` falls back from `MAP_HUGETLB` to plain 4 KB pages and stayed
there; `MADV_HUGEPAGE` costs nothing and needs no hugetlb pool or privilege, and 116.3 of the arena's
119.5 GiB lands on transparent 2 MB pages. Measured neutral on the reference rig - the expert streams are
multi-MB sequential reads, so TLB reach was never the wall - kept for the documented design intent
(include/strata/core/pinned.hpp), A/B with `STRATA_NO_THP=1` like the Windows large-pages switch.
* SERVER. The start-up "reading N GB" narration now takes the true arena size from the pack's
native_experts.txt header (128 GB; the named shard holds 49) with the shard-family sum as fallback.
HOW OTHERS GET IT - nothing is uploaded, every byte of model data is produced on their machine:
```sh
git fetch && git checkout <branch> # the code IS the Q8_0 support
python3 tools/iq_pack.py --gguf Qwen3.8-Flash-Next-Q8_0-00001-of-00006.gguf \
--out packs/q8_0 --compat-bf16 # local conversion from their own GGUF
python3 tools/mtp_fetch.py fetch --out mtp-src # the MTP drafter, once per model family
python3 tools/mtp_pack.py --src mtp-src --out mtp/mtp-q2_0.gguf
python3 tools/mtp_rt.py --gguf mtp/mtp-q2_0.gguf --out mtp/rt
```
Three things worth saying out loud: `--compat-bf16` is mandatory for the plain Q8_0 checkpoint (the routers
and projections ship Q8_0 but the engine reads BF16; the packer refuses without the flag, by design); the
conv1d fix lives in the packer, so every rebuilt pack is correct and every OLD broken pack is refused at
boot; and the MTP drafter is per model family, not per quant - build it once, share it between IQ3_S and
Q8_0.
* VERIFIED AGAINST AN ORACLE, not against another quant. llama.cpp's CPU build (which knows `qwen4exp`) fed
the exact token ids the engine processed, greedy: the fixed build's first tokens match it exactly and the
two then split on a genuine near-tie (Strata's GPU experts quantize activations to Q8_1; llama.cpp CPU uses
Q8_0 for Q8_0 weights) - the "divergence from IQ3_S" that started this was IQ3_S being a different
quantization, not a bug.
* COST, STATED PLAINLY. Experts stream 2.51 GB per token against IQ3_S's 0.61, so decode runs ~21-23 tok/s
against IQ3_S's 38.7 on the reference 4-card rig, with ~51% expert-cache hits (3,637 of 24,576 pairs
resident at 5.22 MB an expert). Prefill is chunk-bound: 115 tok/s at the 2048-token chunk the first config
shipped, 202 tok/s at 4096 (+75%) - the engine clamps an explicit chunk to what every cache participant can
lend. The 119.5 GiB expert arena needs `numactl --interleave=0,1` on dual-socket boards. Quality is the
point of Q8_0; the trade is deliberate.
🤖 Generated with [Claude Code](https://claude.com/claude-code)No site
Links install, modelos, releases.