Issues / #803
#803 Perplexity ~7-10% above llama.cpp on the same GGUF (Qwen3.8-Flash-Next), also at 2K context, every version tested
closed · @mr20399 · 2 Kommentare · Auf GitHub
Setup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindows
Beschreibung
## Summary Same weights, same token ids, same scoring positions: Strata's teacher-forced perplexity is consistently ~7-10% higher than llama.cpp's. The gap is the same on 0.1.32, 0.1.39 and the sergqwer/strata-nvfp4 fork (0.1.38-nvfp4.2), with every expert pack we tried, and it is already present on 2048-token chunks (where QSA is exactly dense), so it looks like a per-token numerics difference rather than QSA selection. Short exact-answer tasks still come out right; the gap shows in perplexity. ## Setup - RTX 5090 32 GB (driver 616.92), Intel i9-14900K (no AVX-512, AVX2 path), 96 GB DDR5-5600, Windows 11 - Strata release binaries: upstream v0.1.32 and v0.1.39 (`strata-windows-x64.zip`), and sergqwer/strata-nvfp4 v0.1.38-nvfp4.2 - Model: `agentionai/Qwen3.8-Flash-Next-AP-GGUF`, `AP-Q4_K_XL/Qwen3.8-Flash-Next-AP-Q4_K_XL.gguf` (101,142,769,536 bytes; experts Q4_K gate/up + IQ4_NL down, Q6_K dense, IQ4_NL PLE). Packed with `tools/iq_pack.py --compat-bf16`; run with `--embd-gguf` (a BF16 token embedding, because the native embedding has no Q6_K GPU dequantizer). - llama.cpp references: unslothai/llama.cpp `mtp/qwen4exp-nextn` a9e9c3c (b10909-mix) and ggml-org master f1cee99 (includes #29751), `llama-perplexity -c 8192 -b 2048 -fa on -ctk q8_0 -ctv q8_0 -ncmoe 39 -ot per_layer_token_embd=CPU` ## Method - Text: 615 KB of English markdown + Python. Token ids come from `llama-tokenize` on the same file, so both engines score identical ids. - llama.cpp: `llama-perplexity` over 8192-token chunks (second half of each chunk scored). - Strata: `--serve` with `--short-read 8192 --prompt-cache 0`, raw `GEN 1 <ids>` per chunk, `STRATA_LOGPOS` column 3 (full-vocabulary log-probability of the actual next token), the same positions (4096..8190 of each chunk). - Engine flags: `--pack <pack> --native <gguf> --resident-budget-gib 56 --embd-gguf <bf16 embd> --expert-profile data/expert-profile.bin --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <draft> --kv int8 --max-context 16384`. ## Results (chunk 1 of 8192 unless noted; lower is better) | Engine | Weights | Perplexity | |---|---|---| | llama.cpp b10909 | AP-Q4_K_XL | 4.0015 | | llama.cpp master f1cee99 | AP-Q4_K_XL | 4.0157 | | Strata 0.1.32 | AP-Q4_K_XL | 4.3871 | | Strata 0.1.39 | AP-Q4_K_XL | 4.3871 | | strata-nvfp4 0.1.38-nvfp4.2 | AP-Q4_K_XL | 4.4024 | | strata-nvfp4 0.1.38-nvfp4.2 | nvidia/Qwen3.8-Flash-Next-NVFP4 (ModelOpt, exact BF16 HC) | 4.3988 | | strata-nvfp4 | GPTQ NVFP4 pack from the BF16 checkpoint (1/3 of the expert error) | 4.3800 | Over 6 chunks: llama.cpp 5.0175 (b10909) / 5.0146 (master), Strata (ModelOpt NVFP4 pack) 5.5524. 2048-token chunks (4 chunks, second half scored; QSA exactly dense): | | after chunk 1 | 2 | 3 | 4 | |---|---|---|---|---| | llama.cpp b10909, AP | 7.3083 | 5.5979 | 6.3807 | 5.5708 | | Strata (fork), AP | 7.6560 | 5.7503 | 6.6659 | 5.9605 (+7.0%) | ## A/B switches that did NOT move it (ModelOpt NVFP4 pack, chunk 1, baseline 4.3988; runs are deterministic) `--kv fp16` 4.3850 · batched prompt path for the first 4097 tokens, then windows: 4.3920 · `STRATA_ROPE_TABLE=0` 4.3815 · `--adapt-every 100000` 4.3988 · `--pcie-frac 1.0` (no CPU experts) 4.3975 · `--pcie-frac 0` (all misses on the CPU) 4.3966 · `STRATA_GR_V3=0` 4.3966 · `--gr-fp32-activations` 4.3988 (no-op with `--native`) · upstream defaults restored in the fork (`STRATA_KV_ROT=0 STRATA_ROPE_TABLE=0 STRATA_GR_V3=0 STRATA_PREFILL_BF16X2=0 STRATA_KV_GROW=0`) 4.3963 · `--no-fused-gr`: the engine exits before READY. ## Notes - The expert weights barely matter inside Strata: a GPTQ pack with 1/3 of ModelOpt's expert error moved chunk 1 from 4.3988 to 4.3800, and a pack with 14% of it (Q8_0 down everywhere) gave 4.4007. The gap seems to sit outside the experts. - docs/UNSLOTH_Q4.md reports parity with llama.cpp 3cf0325 (CPU build) on UD-Q4_K_XL (4.27 vs 4.29 after a 16K prompt). We cannot reproduce that kind of agreement on AP-Q4_K_XL with GPU llama.cpp builds. The AVX2 CPU path is not the cause (pcie-frac 0 / 1 give the same result). - Happy to run any diagnostic build or dump you suggest.
Mehr auf der Site
Links zu Install, Modellen, Releases.