Issues / #1428
#1428 --layer-split auto on a P40 + RTX 3070 (3070 first) flips from K=43 to K=47 with --kv int8 or --vram-reserve-mib 600, and prefill drops from 458 to 111-113 tok/s (0.1.40.2)
open · @paulhothersall · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quants
描述
**On v0.1.40.2, `auto` with the 3070 first picks K=43 with the default flags (4K prefill 458 tok/s) and K=47 with either `--kv int8` or `--vram-reserve-mib 600` added (prefill 113 and 111 tok/s), one launch each; the cost model predicts 63.1 ms per decode window for K=43 and 61.9 / 62.8 ms for K=47, so it moves layers onto the P40 for about 1-2% of predicted decode while the prompt path gets about 4x slower.** The same flip showed on 0.1.39 and 0.1.40 (0.1.39 also with `--kv-resident 32768`). These flags are in the community benchmark recipe. Setup: P40 24 GB (cc 6.1) + RTX 3070 8 GB (cc 8.6), PCIe 3.0 x8 each, 64 GB RAM, IQ3_XXS, ctx 36,864, v0.1.40.2 built from the tag (archs 61, 75, 86), 4K prompt, 3 runs, prefill = prompt tokens / TTFT. | flags added | auto K | prefill 4K | predicted decode window | |---|---|---|---| | none | 43 | 458 | 63.1 ms | | `--kv int8` | **47** | **113** | 61.9 ms | | `--vram-reserve-mib 600` | **47** | **111** | 62.8 ms | Measured decode is about the same in all rows (28-34 tok/s). For reference on this pair: a forced K=34 with `STRATA_STAGE_TRIM=1` gives 446 / 799 prefill (4K / 32K). Questions: 1. Should the search price prompt reading, or need a margin before it moves layers onto the smaller card? 2. Is K=47 meant to be reachable here, with the P40 keeping one layer? Not tested: other card pairs, forced splits near K=43-47 on 0.1.40.2, the `--kv-resident` flag on 0.1.40.2.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。