Issues / #1616
#1616 Auto layer assigns only 1 layer to card dropping prefill due to prompt-chunk clamp
open · @tonydiep · 0 comentarios · En GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsWindows
Descripción
# Summary
On a 4-GPU rig (RTX 5060 Ti 16GB + RTX 3090 24GB + 2x RTX 3060 12GB, no NVLink), prompt processing for
Qwen3.8-Flash-Next IQ3_XXS fell from **576 tok/s (0.1.38)** to **~347 tok/s (0.1.41)**. The dominant cause
is the **prompt-chunk clamp (#448)**: `layer_split: auto` gives a 12 GB 3060 a **single layer**, which
holds only **512 expert-cache slots**, and that card clamps the prompt chunk for every card in the
pipeline. A manual `layer_split` that stops any card owning one layer roughly **doubles prefill** — confirmed
**347 → 826 tok/s (+138%)** on this box.
## The clamp, from the startup log
```
layer split: layers 0-1 (CUDA0 3060), 2 (CUDA1 3060), 3-19 (5060 Ti), 20-47+head (3090)
expert cache slots: CUDA1 512 (0.89 GiB) CUDA2 6442 CUDA3 9125
WARNING: prompt chunk 1280 tokens, not 1792: CUDA1 (RTX 3060) has 512 expert-cache slots,
and a 1792-token chunk borrows 492 of them (it must keep 128, and lend at most 90%)
- prompts read slower than on CUDA0 alone (#448)
```
A 12 GB card owning 1 of 48 layers clamps the chunk to 1280 for the whole pipeline. `--prefill auto` would
allow 8192; the layout only reaches 1280.
## Evidence (matched harness, 3 runs x 3 reps, median; baseline = live config minus --vision)
| arm | change | prefill tok/s | vs baseline | gen tok/s |
| --- | --- | --- | --- | --- |
| A0 | baseline (chunk 1280) | 377.5 | — | 105.4 |
| A2 | 3090+5060Ti only, no 1-layer stage (chunk 8192) | 831.2 | +120% | 92.6 |
| A4 | A2 + STRATA_PF_FUSED=1 | 883.5 | +134% | 94.4 |
| A5 | keep 4 cards, manual split 4,8,24 (chunk 4864) | 868.2 | +130% | 100.6 |
## Workaround that fixed it (confirmed in production)
Keep all four cards but stop any card owning a single layer:
```json
"gpu": [1, 3, 0, 2],
"layer_split": "4,8,24",
"env": { "STRATA_PF_FUSED": "1" }
```
Each 3060 now owns 4 layers (holds ~2048 slots instead of 512). After restart the engine reports
`prompt chunk auto: 6144` (was 1280) and `prompt experts on the fused int8 kernels (STRATA_PF_FUSED=1)`.
Measured on the live server: **prefill 347 -> 826 tok/s (+138%)**, gen 109 -> 98. A residual clamp remains
(a 12 GB 3060 only has ~2048 slots, so an 8192-token chunk borrows 1938 of them and is trimmed to 6144),
but it is far weaker than the 1-layer case.
## Things that did NOT help (so they can be ruled out)
- `--pcie-frac 0.20`: +0.1% (noise) once the chunk is fixed. The earlier "+15%" only appeared at a small chunk.
- `--pipeline-windows 2`: no gain over the 2-card baseline.
- `--prefill auto:32768` (chunk > 8192): no gain; 8192 is the ceiling.
- GPU power caps: already measured as not the cause.
## Questions for maintainers
1. Should `layer_split: auto` avoid assigning a single layer to the smallest-VRAM card? That card then
clamps the chunk for everyone; a manual split is strictly faster on this rig.
## Environment
Engine 0.1.41, model Qwen3.8-Flash-Next IQ3_XXS, 262144 ctx, `--kv q4_0`. Rig: RTX 5060 Ti (PCIe x8),
RTX 3090 (x4), 2x RTX 3060 (x4). Related to #448.En el sitio
Enlaces a install, modelos, releases.