Issues / #1576

#1576 [Bug]: 0.1.41 auto card order starves the smaller card on a VRAM-asymmetric pair (12 GB fast + 22 GB slow)

open · @sankerXD · 1 コメント · GitHub で見る

BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows

本文

**What happened**

v0.1.41's new automatic card order (#1352) reorders our two GPUs so the faster card runs the last stage, and the auto layer split follows. On our pair — a **fast but small** card plus a **slower but big** card — this puts the head/draft stage on the card that can least afford it:

- 0.1.40.x, given order `"gpu": [0,1]`: the 4070 Ti 12 GB ran layers 0-28; the 2080 Ti 22 GB ran layers 29-47 + head + draft, with ~7,000 expert-cache slots (~14 GiB) on the last stage.
- 0.1.41, auto order: the 2080 Ti ran layers 0-16; the **12 GB** 4070 Ti ran layers 17-47 + head + draft + verify, leaving only **1,572 cache slots (3.13 GiB)**. VRAM free hit **0 MiB** during the verify-window captures, and the engine's own warning fired: `a card this small can serve as a helper expert cache instead of a split stage`.

Measured effects:

- Decode: **~58 -> ~44 tok/s**
- Prompt chunk auto-cut 8,192 -> 4,608 (the cache-borrow floor). Prompt ingestion still felt faster overall — the batched file reads and faster-card-last genuinely help there.
- Output drift on a long structured-generation task (SVG drawing; feet misplaced relative to wheels — our established canary for expert-set/numerics drift; IQ3 output is not expected to be bit-identical across runs anyway, but this was a visible quality regression)

`"gpu_order": "as_given"` restored a healthy layout: auto split picked 0-33 / 34-47 with **6,882 slots (13.97 GiB)** on the last stage, ~582 MiB VRAM free, decode and canary output back to normal.

**Suggestion**: when reordering for an auto split, weigh the last stage's extra VRAM cost (head + draft + verify + cache floor) against each card's free VRAM — or keep the given order when the reordered last stage would land below some cache-slot / free-VRAM floor (the "card this small" warning firing right after the auto placement is a ready-made signal). Alternatively keep the reorder but let the split compensate: the reordered run gave the 22 GB card only 17 layers while the 12 GB card carried 31 + head.

**Strata version, GPU, OS**

- Strata 0.1.41, release engine (BUILD.json: sm_75/86/89/120, CUDA 13.0), Windows 11, 64 GB RAM, machine dedicated to the model
- GPU0: RTX 4070 Ti, 12 GB, PCIe gen2 x16
- GPU1: RTX 2080 Ti, 22 GB (VRAM-modded card), PCIe gen3 x4
- Note the faster card also has ~2x the PCIe bandwidth; the auto placement sees links too (#794), which may be part of why the split moved so far (17 vs the usual 26-30).

**Config** (sanitized; pack = Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S)

```json
"gpu": [0, 1], "layer_split": "auto",
"args": ["--expert-cache","auto","--prefill","auto","--spec","4","--mtp","...",
  "--max-context","524288","--rope-scaling","yarn","--rope-scale","2",
  "--kv","int8","--kv-resident","131072","--vision","--vram-reserve-mib","700",
  "--pcie-frac","0.20","--spec-min-p","0.70","--pipeline-windows","2"],
"env": {"STRATA_PREFILL_HELP": "1"}
```

**Engine log** (relevant excerpts)

```
# 0.1.40.x, given order (several runs):
layer split: CUDA1 runs layers 29-47, expert cache 6867 slots (13.77 GiB), 6867 of its 9728 profiled pairs
layer split: CUDA0 runs layers 0-28
# (earlier runs had up to 8030 slots / 16.13 GiB on the last stage)

# first 0.1.41 start, auto order:
layer split: CUDA1 runs layers 17-47, expert cache 1572 slots (3.13 GiB), 1572 of its 15872 profiled pairs
layer split: CUDA0 runs layers 0-16
WARNING: prompt chunk 4608 tokens, not 8192: CUDA1 (NVIDIA GeForce RTX 4070 Ti) has 1572 expert-cache slots, and a 8192-token chunk borrows 1572 of them (it must keep 128, and lend at most 85%) - prompts read slower than on CUDA0 alone (#448)
  a card this small can serve as a helper expert cache instead of a split stage: without --layer-split, with --expert-cache-device1 N (docs/SECOND_GPU.md)
554 MiB of VRAM free with everything loaded
strata verify: capturing the 6-token window (0 MiB of VRAM free)

# same config + "gpu_order": "as_given", 0.1.41:
layer split: CUDA1 runs layers 34-47, expert cache 6882 slots (13.97 GiB), 6882 of its 7168 profiled pairs
layer split: CUDA0 runs layers 0-33
582 MiB of VRAM free with everything loaded
```

**Not tested**: `--batch` (unused); a manual split under the new order; whether the output drift came from the layer-to-card reassignment itself or from the starved cache changing the resident expert set — we only verified that `as_given` restored both speed and canary quality.

関連リンク

インストール・モデル・リリースへの站内リンク。