Issues / #1029
#1029 docs: MULTI_GPU.md lists Pascal as unsupported for the layer split, but 2x Tesla P40 runs it (v0.1.39)
closed · @sarge18 · 1 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentation
Beschreibung
**Observation (v0.1.39, `main` = `6f32ec0`, unmodified):** two Tesla P40 22 GB (compute capability 6.1) ran Flash-Next IQ2_XS with the layer split, although `docs/MULTI_GPU.md` lists cards below compute capability 7.5 as unsupported.
**How:** I installed with `./setup.sh ... --cuda 12 --build --gpu 1` (CUDA 12.4 + g++-13, engine compiled for `sm_61` with `STRATA_EXPERIMENTAL_SM60=1`), then edited the generated config by hand to `"gpu": [0, 1]` and `"layer_split": "auto"` and started `serve/server.py --engine strata --config ...`. I did **not** try `--gpus 0,1` through setup, so I can't say whether setup itself would refuse it.
**Evidence (engine log, two-card runs):**
```
strata generate: layer split across 2 GPUs: CUDA0, then CUDA1 (split auto)
strata generate: layer split auto: K=24 - predicted 75.8 ms per decode window; the caches hold 24576 of 24576 profiled pairs (~100.0% of the routed mass)
strata generate: layer split: CUDA1 holds its weights, session [24, 48) and the head; 18.63 GiB free
```
Median decode for a short prompt went from 19.6 tok/s (one P40) to 34.8 tok/s (both); prompt reading 348 / 388 tok/s (4K / 16K) on one card and 353 / 501 on both. Three runs per cell; the first run of each is much slower (warm-up). All per-run data, engine logs and the config files are in #1028 (a results-only community report).
**Docs:** `docs/MULTI_GPU.md` ("Not supported", and the example listing a GTX 1080 Ti as "not supported - older than the RTX 20 series") and `docs/OLDER_GPUS.md` (which allows `--gpus 0,1` "with a newer card" and says a V100 can share a model with an RTX 30/40 card) read differently for **two Pascal cards**, which neither page mentions. If this is expected to work experimentally, a sentence in either page would help; if it only works by hand-editing the config, that may be worth saying too.
**Caveats:** one machine, one quantisation (IQ2_XS), 32,768 context, three runs; the two-card runs were not NUMA-pinned. Not tested: `--batch` with a layer split, vision, contexts above 32K, other quantisations. Happy to test a specific variant if that helps (for example `--gpus 0,1 --cuda 12` through setup, or IQ3_S).
Related Pascal reports: #875, #395. The wider study (llama.cpp and Ollama on the same weights, accuracy and repeatability): https://github.com/sarge18/p40-llm-engine-bakeoff
Disclosure: prepared with an AI assistant (Claude Sonnet 5.5, medium effort) under my direction; every number above comes from the per-run data in #1028.
Mehr auf der Site
Links zu Install, Modellen, Releases.