Issues / #875
#875 P40 + 3070 (mixed Pascal/Ampere pair): `layer_split: auto` vs forced splits on 0.1.39, with hit rates (follow-up to #604)
open · @paulhothersall · 4 commentaires · Sur GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
Description
[p40-3070-logs-v0.1.39.tar.gz](https://github.com/user-attachments/files/33039641/p40-3070-logs-v0.1.39.tar.gz) Follow-up to #604, where you asked for "a P40 + 3070 with `layer_split: auto` against a few forced splits (`--stats` and the hit rates)". Here is that run, on v0.1.39 as released (no patches; CUDA 12.9 engine, `STRATA_EXPERIMENTAL_SM60=1`), IQ3_XXS, 36,864-token context, P40 power-capped at 130 W, DDR4-2400, PCIe 3.0 x8. Three runs per cell with `bench/results/.../strata-bench.py` (unique prefix per run, 256 tokens). Prefill / decode in tok/s; hit = decode expert-cache hit rate range over the runs. | config (3070 first unless noted) | slots 3070 / P40 | chunk | 4K | 32K | hit | |---|---|---|---|---|---| | P40 alone | - / 10,451 | 8192 | 322 / 34 | 339 / 34 | 83-96% | | 3070 alone | 513 / - | 1024 | 144 / 30 | 155 / 37 | 14-29% | | **auto** (K=43) | 1,536 / 2,560 | 4096 | 462 / 35 | 617 / 42 | 39-58% | | auto, P40 first | 8,192 / 736 | 1024 | 193 / 32 | 221 / 37 | 26-42% | | forced K=8 | 2,177 / 10,333 | 5888 | 333 / 39 | 362 / 37 | 89-98% | | forced K=20 | 2,070 / 9,780 | 5120 | 386 / 38 | 491 / 38 | 82-89% | | K=20 + `STRATA_STAGE_TRIM=1` | 3,459 / 10,500 | 8192 | 392 / 42 | 511 / 48 | 90-96% | | K=30 + trim | 2,714 / 9,216 | 6912 | 446 / 40 | 707 / 39 | 52-78% | | K=34 + trim | 2,375 / 7,168 | 6144 | 481 / 40 | 879 / 49 | 50-71% | | K=36 + trim | 2,173 / 6,144 | 6400 | 478 / 38 | 909 / 40 | 45-60% | | K=40 + trim | 1,892 / 4,096 | 5632 | 499 / 39 | 800 / 37 | 34-67% | | K=46 + trim | 1,537 / 1,024 | 2816 | 327 / 34 | 427 / 34 | 25-46% | What I take from it (one rig, three runs, please weigh accordingly): 1. **Does the speed term matter on a mixed pair? Yes, and 0.1.39's choice is good.** It puts the 3070 first and gives it 43 of 48 layers; that is 1.4-1.8x the P40's prefill and 35-42 decode. The best forced split I found (K=34-36 with the trim) reads a 32K prompt about 45% faster than auto (879-909 vs 617), at the same decode. The opposite order is the bad case: with the P40 first the chunk falls to 1024 (the 3070's room sets it, as @gopinath87607 described) and the pair is slower than the P40 alone. 2. **Prefill follows how many layers the faster card runs, not the hit rate.** The hit rate falls from 98% (K=8) to 45-60% (K=34-36) while 32K prefill goes 362 -> 909. Decode is nearly flat (34-42) across all of it, except with the carve. 3. **The carve is the only change that moved decode at an unchanged split** (K=20: 38 -> 42 / 48 tok/s, hit rate 82-89% -> 90-96%, chunk 5120 -> 8192). It is opt-in and needs explicit split points, so `auto` cannot use it. Would you take a change that lets `--layer-split auto` use it (the cost model already counts what trimming frees)? That is where I would look first. 4. **A worry about the cost model:** for auto it printed "the caches hold 4146 of 24576 pairs (~91.7% of the routed mass)", while the decode hit rates measured 39-58%. On IQ2_XS auto chose K=47 (3070 holds 47 layers, 512 slots on the P40), prefill 175-189 and decode 26-27: slower than the P40 alone. With K=36 + trim the same quant gave 527 / 990 prefill. This looks like the mass-curve problem gopinath87607 describes. 5. **`--stats`:** I could not combine it with a split: `--stats` prints only in one-shot mode and `--layer-split` needs `--serve`. For the single P40 I have the breakdown (above, earlier comment); for splits I only have the engine log (slots, chunk, hit rate, the cost model's per-card ms). If there is a flag or `/metrics` field that gives per-stage time on a split in serve mode, tell me and I will rerun. 6. Not helpful here, for the record: `--peer-device` (no P2P between these cards: the log says the prompt path stays on the primary; decode stayed 32-42), `--kv int8`, `--spec 2/3/6`, `STRATA_BF16_TC=1` (all within noise), `--prefill` below auto's choice (-4% to -24%). Still to do on this rig: IQ3_S and Q2_0, concurrency (`--batch`), Hardin22's fork on this pair, a 32 GB-RAM boot, and 2 x P40 (one card on the chipset x4 slot). I will send those as separate notes. Logs, configs and the runner scripts are available if useful. Attachments: the engine logs (`strata-<model>.log`) for each row, the generated configs, and the runner scripts. Related: #584 (placement search, mass curve), #639 / #559 (weight carve), #583 (ring), #642 (resident mode on a split).
Sur le site
Liens install, modèles, releases.