Issues / #1235
#1235 Sapphire Rapids (Xeon w7-2475X): +8% decode, +3% prefill and byte-identical greedy output. STRATA_IQ_MT_MIN=1 is faster here, and 8 stager threads beat 32
open · @enkynakamura · 0 commentaires · Sur GitHub
BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentationWindows
Description
Hi @Niko1221 , I ran a measured tuning pass on my workstation and three results go against current defaults or docs. Numbers only, all from the CLI harness. Happy to run more arms if useful. ### Setup - CPU: Xeon w7-2475X (Sapphire Rapids, 20 cores, AVX-512), 4 × 32 GB DDR5-4800 ECC, quad channel, JEDEC - GPU: RTX 5080 16 GB. Windows, CUDA 13.3, local build (`-DCMAKE_CUDA_ARCHITECTURES=120`, Release) - Model: Qwen3.8-Flash-Next UD-IQ4_XS, GGUF read in place (`--native`), `--expert-cache 2390` (2672-2675 slots in every run), `--pcie-frac 0.33`, `--kv int8`, yarn ×3 - Measured on v0.1.39; reproducibility and the tuned vs default comparison rechecked on v0.1.40.1 (82f46a8) - Method: greedy, 512 new tokens, arms interleaved, 3 runs per arm. Values are median [min-max]. Output identity is a hash of the generated ids. ### 1. `STRATA_IQ_MT_MIN=1` is faster on this CPU, not slower `native_expert.cpp:79-96` and DETAILS.md give -1..-3% decode for MT_MIN=1 (Ryzen 7600, IQ3_S). On the w7-2475X with UD-IQ4_XS it is the opposite. Clean comparison, both arms with `--suffix-draft 0` (output identical within each arm), 1424-token prompt, `--spec 4`: | arm | decode tok/s | gate/up ms/round | pool GB/s | distinct CPU experts/layer | |---|---|---|---|---| | `--suffix-draft 0` | 71.59 [70.03-71.60] | 9.765 [9.739-10.314] | 60.5 [57.0-60.6] | 7.49 | | `--suffix-draft 0` + `STRATA_IQ_MT_MIN=1` | 75.08 [75.07-75.21] | 8.089 [8.061-8.178] | 69.1 [68.5-69.6] | 7.77 | +4.9% decode, -17.2% gate/up time, with slightly more distinct experts per layer in the MT_MIN=1 arm. The 0.1.40 `STRATA_IQ3S_MT1` path does not cover this case (it requires `!cpu_avx512_ok()` and IQ3_S). ### 2. `STRATA_STAGER_THREADS`: 8 beats the default 32 for a UD pack with its experts in RAM `prefill.cpp:861` picks 32 threads whenever any expert is `transient()`. For this pack everything sits in RAM / file cache (the decode log reports `files 0 blobs, 0.0 MB read`), and 32 is slower. 352K-token prompt, tuned config otherwise identical: | STRATA_STAGER_THREADS | prefill tok/s | decode tok/s | |---|---|---| | 8 | 2284.4 [2283.0-2310.1] | 54.94 | | 16 | 2273.4 [2256.2-2277.0] | 55.06 | | 32 (default here) | 2156.3 [2152.3-2157.8] | 55.00 | +5.9% prefill for 8 vs 32. Outputs identical in all 9 runs. The 32-thread rule was measured on a pack that really reads the SSD; maybe the choice could depend on whether the transient experts are actually cached. This may be the same mechanism as #1056 (Stager threads yield-spin; PR #1101 is still open). I have not tested #1101 here. ### 3. Reproducible greedy output without the speed cost Same 1424-token prompt, `--spec 4`, `STRATA_IQ_GATHER=1`, `STRATA_IQ_PREFETCH=4096` in all arms: | arm | decode tok/s | distinct output hashes (3 runs) | |---|---|---| | defaults | 70.48 [68.48-71.91] | 3 | | `STRATA_IQ_MT_MIN=1` | 75.07 [75.01-75.42] | 2 | | DETAILS.md recipe: MT_MIN=1 + `--adapt-swaps 0 --pcie-frac 0` | 68.05 [67.90-68.33] | 1 | | `--suffix-draft 0` | 71.59 [70.03-71.60] | 1 | | `--suffix-draft 0` + MT_MIN=1 | 75.08 [75.07-75.21] | 1 | MT_MIN=1 alone did not make the output reproducible; turning the suffix drafter off did, while keeping `--pcie-frac 0.33` and the adaptive tier on. Most of the recipe's cost seems to come from `--pcie-frac 0`: that arm computes 17.6 distinct experts per layer on the CPU against 7.5-7.8 in the others. My guess at the cause (not proven): the lookup vs MTP choice in `DraftPolicy` is driven by measured wall-clock round time (`generate.cpp:10831-10834` → `draft_policy.cpp:84-87`), so timing noise changes the windows and from there the expert grouping and the PCIe share. On v0.1.40.1, 352K-token prompt: the tuned config (`--suffix-draft 0`, MT_MIN=1) gave 1 hash in 3 runs; the default drafting gave 2 hashes in 3 runs. In #410 (0.1.32, Zen 3, IQ2_XS), fresh CLI runs with MT_MIN=1 were stable. Here fresh CLI runs with MT_MIN=1 still varied while the suffix drafter was on. In the #410 trace, the first difference is a window that held 2 tokens in one request and 4 in the other. These are fresh CLI runs. I have not tested repeated requests through the server (#410). ### Minor - `STRATA_IQ_GATHER=1` on Sapphire Rapids: pool GB/s 55.9 [55.7-56.0] → 59.2 [59.1-59.5] (+5.9%). It is bit-exact: same hash with and without it under `--suffix-draft 0`. The comment in `iq_avx512.cpp:31-32` only reports the Zen 4 slowdown; on Intel it may be worth enabling by default. - The `--expert-cache N` help text (`generate.cpp:803`) says "keep N expert blobs resident". With a native pack and a profile, N × max_blob is a byte budget (`generate.cpp:3921`), so 2390 gave 2672-2675 slots here. ### End result on v0.1.40.1 (352K-token prompt) Both arms use the same args and differ only here. Defaults: `--spec 4`, no env. Tuned: `--spec 5 --suffix-draft 0`, plus `STRATA_IQ_MT_MIN=1 STRATA_IQ_GATHER=1 STRATA_IQ_PREFETCH=4096 STRATA_STAGER_THREADS=8`. | | decode tok/s | prefill tok/s | |---|---|---| | defaults | 51.20 [51.15-53.39] | 2191.7 | | tuned | 55.38 [53.29-55.56] | 2257.8 | Decode +8.2%, prefill +3.0%. The tuned output was identical in 3 runs; the defaults produced 2 different outputs.
Sur le site
Liens install, modèles, releases.