Issues / #1239

#1239 v0.1.40 two-GPU tester results on a P40 + RTX 3070: 150-request soak clean, --pipeline-windows 2 +9% / +15% decode in 5 of 5 pairs

open · @paulhothersall · 0 commentaires · Sur GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsWindows

Description

**On v0.1.40, a P40 + RTX 3070 pair ran a 150-request soak of the resident mode with `--adapt-async 1` with no error, and `--pipeline-windows 2` raised decode by 8.7% (no resident mode) and 15% (resident mode) in 5 of 5 interleaved pairs, with prefill unchanged.** Limit: one rig, 4K prompts, three runs per cell; decode swings 34-47 tok/s inside one launch, so the gain is small next to the spread.

Setup: v0.1.40 (1735d64), CUDA 12.9, `-DCMAKE_CUDA_ARCHITECTURES="61;86"`, Tesla P40 24 GB + RTX 3070 8 GB, PCIe 3.0 x8 each, no P2P, 64 GB RAM, Qwen3.8-Flash-Next IQ3_XXS, `--layer-split 34`, context 36,864, 4K prompts; decode from the engine line per run.

| #859 pairs (5 interleaved, 3 runs each) | decode default | decode `--pipeline-windows 2` | prefill (TTFT, runs 2-3) |
|---|---|---|---|
| no resident mode | 39.0 (sd 3.6) | 42.4 (sd 2.0), +8.7%, higher in 5/5 pairs | 9.01 s vs 9.24 s |
| `--resident-experts` | 31.5 (sd 1.3) | 36.2 (sd 2.6), +15%, higher in 5/5 pairs | 9.61 s vs 9.59 s |

Soak (#848 + #876, `--resident-experts --adapt-async 1`, K=34, back-to-back 4K prompts): 150 requests in 38.5 minutes (632K prompt and 26.8K generated tokens), no crash or engine error, decode median 36.1 tok/s in the first quarter and 36.5 in the last, TTFT 9.6 s both. GPU power medians 99 W (3070) and 66 W (P40, 148 W peak against a 130 W cap); temperatures median 60 C (3070) and 50 C (P40), maxima 63 and 52.
Resident mode on the split (K=34, one launch, 4K decode): none 38.8; `--resident-experts` 34.2; plus `--adapt-async 1` 35.9.
CPU flags (Q2_0, RTX 3070 alone, 6 cores, one launch per arm, 4K decode): default 37.9; `STRATA_Q2_BITPLANE=1` 40.6; `STRATA_ADAPT_LAG=2` 38.9; both 43.9 (+16%). P40 with `--expert-cache 2000`: 34.1 vs 35.7 with both.
Not tested: IQ3_S, other pairs, 3-4 GPUs, 32K prompts for the pairs.

Sur le site

Liens install, modèles, releases.