Issues / #1428

#1428 --layer-split auto on a P40 + RTX 3070 (3070 first) flips from K=43 to K=47 with --kv int8 or --vram-reserve-mib 600, and prefill drops from 458 to 111-113 tok/s (0.1.40.2)

open · @paulhothersall · 0 commentaires · Sur GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quants

Description

**On v0.1.40.2, `auto` with the 3070 first picks K=43 with the default flags (4K prefill 458 tok/s) and K=47 with either `--kv int8` or `--vram-reserve-mib 600` added (prefill 113 and 111 tok/s), one launch each; the cost model predicts 63.1 ms per decode window for K=43 and 61.9 / 62.8 ms for K=47, so it moves layers onto the P40 for about 1-2% of predicted decode while the prompt path gets about 4x slower.** The same flip showed on 0.1.39 and 0.1.40 (0.1.39 also with `--kv-resident 32768`). These flags are in the community benchmark recipe.
Setup: P40 24 GB (cc 6.1) + RTX 3070 8 GB (cc 8.6), PCIe 3.0 x8 each, 64 GB RAM, IQ3_XXS, ctx 36,864, v0.1.40.2 built from the tag (archs 61, 75, 86), 4K prompt, 3 runs, prefill = prompt tokens / TTFT.
| flags added | auto K | prefill 4K | predicted decode window |
|---|---|---|---|
| none | 43 | 458 | 63.1 ms |
| `--kv int8` | **47** | **113** | 61.9 ms |
| `--vram-reserve-mib 600` | **47** | **111** | 62.8 ms |

Measured decode is about the same in all rows (28-34 tok/s). For reference on this pair: a forced K=34 with `STRATA_STAGE_TRIM=1` gives 446 / 799 prefill (4K / 32K).

Questions: 1. Should the search price prompt reading, or need a margin before it moves layers onto the smaller card? 2. Is K=47 meant to be reachable here, with the P40 keeping one layer? Not tested: other card pairs, forced splits near K=43-47 on 0.1.40.2, the `--kv-resident` flag on 0.1.40.2.

Sur le site

Liens install, modèles, releases.