Pull requests / #934
#934 Per-layer expert cache: each layer's slots are that layer's blob
closed · @Chromadera · 0 commentaires · Sur GitHub
BenchmarksServer & APIAMD / HIPModels & quantsDocumentation
Description
## Credit Measured and prepared by Chromadera. The arrangement comes from the report that this model wrote on its own expert pack: **Qwen3.8-Flash-Next GSQ-RCO IQ3_XXS**, the weights served as `qwen3.8-flash-next-gsq-iq3-xxs` (`Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf` and shard 2, pack `strata-gsq-iq3-xxs`). ## Schematic  `--expert-cache N` with `--expert-cache-per-layer` is still a budget of N slots of the largest blob. On a native pack the engine now spends that budget on each layer's real blob. Every layer gets the same number of slots, because a per-layer range is `layer * quota + n`. On Qwen3.8-Flash-Next GSQ-RCO IQ3_XXS, 48 layers, `--expert-cache 7008`: | | before | after | |---|---|---| | slots | 7,008 (146 per layer) | 9,312 (194 per layer) | | VRAM | 15.20 GiB | 15.14 GiB | | experts page-locked in RAM | 38.00 GiB | 37.10 GiB | Startup prints: `per-layer slots use each layer's blob: 7008 uniform slots (15.20 GiB) -> 9312 slots, 194 per layer (15.14 GiB)` ## Why this does not hit the #369 problem #369 kept per-layer mode on the largest blob because the sized-slot list was built in profile order. A size chosen for one layer could land in another layer's slot. This list is built in layer order. Slot `layer * quota + n` is that layer's blob, and every expert in a layer is one size. `open_sized` already starts each layer at the bottom of its range, and the profile fill already skips a layer once that layer is full. If one copy of every layer does not fit, the cache stays the old uniform slots. When the cache shrinks, it rebuilds the same equal counts. Cutting the list in the middle would put one layer's blob into the next layer's range. A layer split is untouched. It already walks that stage's profile. The shared cache, without `--expert-cache-per-layer`, is unchanged. ## Measured RX 7900 XTX, same binary, same profile, only the slot layout changed. BetterBench cold prefill, thinking on, 5 passes: | prompt | 146 uniform | 194 real size | |---|---|---| | 8,086 | 1,518.9 tok/s | 1,524.7 | | 63,949 | 1,686.4 | 1,690.2 | | 128,136 | 1,661.5 | 1,670.3 | | ~197k | 1,587.9 | 1,587.2 | No depth got slower. Past the 28,928-token chunk a long prompt is dominated by streaming, so the extra cached experts barely show up there, and they did not cost anything. The weights are the same bytes. A GPU-resident expert rounds differently from the CPU copy; the engine already records that as equal perplexity. More slots means more tokens take the GPU path. There is no known performance drawback from this. Decode on the shipped profile at 194 slots was 48.9 tok/s weighted (5 passes). The earlier uniform 146-slot serve was 43.8 (20 passes) on an older binary, so that pair is not a clean comparison. ## Check Start with `--expert-cache-per-layer` on a native pack. The log line should show more slots than N, and the profile fill should fill them. Slot 0 is still checked against the file.
Sur le site
Liens install, modèles, releases.