Issues / #1235

#1235 Sapphire Rapids (Xeon w7-2475X): +8% decode, +3% prefill and byte-identical greedy output. STRATA_IQ_MT_MIN=1 is faster here, and 8 stager threads beat 32

open · @enkynakamura · 0 Kommentare · Auf GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentationWindows

Beschreibung

Hi @Niko1221 , I ran a measured tuning pass on my workstation and three results go against current defaults or docs. Numbers only, all from the CLI harness. Happy to run more arms if useful.
 
### Setup
- CPU: Xeon w7-2475X (Sapphire Rapids, 20 cores, AVX-512), 4 × 32 GB DDR5-4800 ECC, quad channel, JEDEC
- GPU: RTX 5080 16 GB. Windows, CUDA 13.3, local build (`-DCMAKE_CUDA_ARCHITECTURES=120`, Release)
- Model: Qwen3.8-Flash-Next UD-IQ4_XS, GGUF read in place (`--native`), `--expert-cache 2390` (2672-2675 slots in every run), `--pcie-frac 0.33`, `--kv int8`, yarn ×3
- Measured on v0.1.39; reproducibility and the tuned vs default comparison rechecked on v0.1.40.1 (82f46a8)
- Method: greedy, 512 new tokens, arms interleaved, 3 runs per arm. Values are median [min-max]. Output identity is a hash of the generated ids.
### 1. `STRATA_IQ_MT_MIN=1` is faster on this CPU, not slower
`native_expert.cpp:79-96` and DETAILS.md give -1..-3% decode for MT_MIN=1 (Ryzen 7600, IQ3_S). On the w7-2475X with UD-IQ4_XS it is the opposite. Clean comparison, both arms with `--suffix-draft 0` (output identical within each arm), 1424-token prompt, `--spec 4`:
 
| arm | decode tok/s | gate/up ms/round | pool GB/s | distinct CPU experts/layer |
|---|---|---|---|---|
| `--suffix-draft 0` | 71.59 [70.03-71.60] | 9.765 [9.739-10.314] | 60.5 [57.0-60.6] | 7.49 |
| `--suffix-draft 0` + `STRATA_IQ_MT_MIN=1` | 75.08 [75.07-75.21] | 8.089 [8.061-8.178] | 69.1 [68.5-69.6] | 7.77 |
 
+4.9% decode, -17.2% gate/up time, with slightly more distinct experts per layer in the MT_MIN=1 arm. The 0.1.40 `STRATA_IQ3S_MT1` path does not cover this case (it requires `!cpu_avx512_ok()` and IQ3_S).
 
### 2. `STRATA_STAGER_THREADS`: 8 beats the default 32 for a UD pack with its experts in RAM
`prefill.cpp:861` picks 32 threads whenever any expert is `transient()`. For this pack everything sits in RAM / file cache (the decode log reports `files 0 blobs, 0.0 MB read`), and 32 is slower. 352K-token prompt, tuned config otherwise identical:
 
| STRATA_STAGER_THREADS | prefill tok/s | decode tok/s |
|---|---|---|
| 8 | 2284.4 [2283.0-2310.1] | 54.94 |
| 16 | 2273.4 [2256.2-2277.0] | 55.06 |
| 32 (default here) | 2156.3 [2152.3-2157.8] | 55.00 |
 
+5.9% prefill for 8 vs 32. Outputs identical in all 9 runs. The 32-thread rule was measured on a pack that really reads the SSD; maybe the choice could depend on whether the transient experts are actually cached.
 
This may be the same mechanism as #1056 (Stager threads yield-spin; PR #1101 is still open). I have not tested #1101 here.
 
### 3. Reproducible greedy output without the speed cost
Same 1424-token prompt, `--spec 4`, `STRATA_IQ_GATHER=1`, `STRATA_IQ_PREFETCH=4096` in all arms:
 
| arm | decode tok/s | distinct output hashes (3 runs) |
|---|---|---|
| defaults | 70.48 [68.48-71.91] | 3 |
| `STRATA_IQ_MT_MIN=1` | 75.07 [75.01-75.42] | 2 |
| DETAILS.md recipe: MT_MIN=1 + `--adapt-swaps 0 --pcie-frac 0` | 68.05 [67.90-68.33] | 1 |
| `--suffix-draft 0` | 71.59 [70.03-71.60] | 1 |
| `--suffix-draft 0` + MT_MIN=1 | 75.08 [75.07-75.21] | 1 |
 
MT_MIN=1 alone did not make the output reproducible; turning the suffix drafter off did, while keeping `--pcie-frac 0.33` and the adaptive tier on. Most of the recipe's cost seems to come from `--pcie-frac 0`: that arm computes 17.6 distinct experts per layer on the CPU against 7.5-7.8 in the others.
 
My guess at the cause (not proven): the lookup vs MTP choice in `DraftPolicy` is driven by measured wall-clock round time (`generate.cpp:10831-10834` → `draft_policy.cpp:84-87`), so timing noise changes the windows and from there the expert grouping and the PCIe share.
 
On v0.1.40.1, 352K-token prompt: the tuned config (`--suffix-draft 0`, MT_MIN=1) gave 1 hash in 3 runs; the default drafting gave 2 hashes in 3 runs.
 
In #410 (0.1.32, Zen 3, IQ2_XS), fresh CLI runs with MT_MIN=1 were stable. Here fresh CLI runs with MT_MIN=1 still varied while the suffix drafter was on. In the #410 trace, the first difference is a window that held 2 tokens in one request and 4 in the other.
 
These are fresh CLI runs. I have not tested repeated requests through the server (#410).
 
### Minor
- `STRATA_IQ_GATHER=1` on Sapphire Rapids: pool GB/s 55.9 [55.7-56.0] → 59.2 [59.1-59.5] (+5.9%). It is bit-exact: same hash with and without it under `--suffix-draft 0`. The comment in `iq_avx512.cpp:31-32` only reports the Zen 4 slowdown; on Intel it may be worth enabling by default.
- The `--expert-cache N` help text (`generate.cpp:803`) says "keep N expert blobs resident". With a native pack and a profile, N × max_blob is a byte budget (`generate.cpp:3921`), so 2390 gave 2672-2675 slots here.
### End result on v0.1.40.1 (352K-token prompt)
Both arms use the same args and differ only here. Defaults: `--spec 4`, no env. Tuned: `--spec 5 --suffix-draft 0`, plus `STRATA_IQ_MT_MIN=1 STRATA_IQ_GATHER=1 STRATA_IQ_PREFETCH=4096 STRATA_STAGER_THREADS=8`.
 
| | decode tok/s | prefill tok/s |
|---|---|---|
| defaults | 51.20 [51.15-53.39] | 2191.7 |
| tuned | 55.38 [53.29-55.56] | 2257.8 |
 
Decode +8.2%, prefill +3.0%. The tuned output was identical in 3 runs; the defaults produced 2 different outputs.

Mehr auf der Site

Links zu Install, Modellen, Releases.