Pull requests / #1159

#1159 Community benchmark: 2x Tesla P40 22 GB (Pascal, experimental CUDA 12), single card and layer split

closed · @sarge18 · 0 コメント · GitHub で見る

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentation

本文

Re-opened from #1028, which GitHub closed automatically when `main` was force-pushed. Same two commits, cherry-picked onto the new `main`; the content is unchanged.

Results-only community benchmark report (no engine or docs changes), following `docs/COMMUNITY_BENCHMARKS.md`.

**What:** Strata v0.1.39 (`6f32ec0`, unmodified) on **2x Tesla P40 22 GB (Pascal, sm_61)**, Flash-Next IQ2_XS, 32,768 context, experimental CUDA 12 engine built locally (CUDA 12.4 + g++-13, Ubuntu 26.04 / glibc 2.43). Once on a single P40, once on **both P40s with the layer split** (`"gpu": [0, 1]`).

**Headline (3 runs per cell, medians; per-run values and ranges are in the README):**

| | one P40 | both P40s |
| --- | --- | --- |
| decode, short prompt | 19.6 tok/s | 34.8 tok/s |
| prompt reading, 4K / 16K | 348 / 388 tok/s | 353 / 501 tok/s |

The first run of each configuration includes expert-cache warm-up and is much slower; every run is listed. The engine log for the two-card runs records the split ("layer split across 2 GPUs: CUDA0, then CUDA1", K=24, all 24576 profiled expert pairs resident). Memory, VRAM, temperatures, per-run JSON, engine logs and telemetry are included. A 69-task correctness suite ran twice (long-context recall 3/3 at ~4K, ~12K, ~24K tokens); the README also reports that two greedy runs of identical prompts gave 49 of 69 identical outputs, i.e. not exactly repeatable on this setup.

**Notes**
- `docs/MULTI_GPU.md` lists cards below compute capability 7.5 as unsupported for the layer split, but this ran on two P40s through the config. That is filed separately as #1029, so this PR stays results-only.
- Related Pascal reports: #875 and #395. The setups differ, so the numbers are not directly comparable.
- The benchmark script (`scripts/bench.py`, stdlib only, prompts are regenerated from a fixed seed) and the full study comparing llama.cpp and Ollama on the same weights are in https://github.com/sarge18/p40-llm-engine-bakeoff.
- Happy to re-run other settings (e.g. IQ3_S, a power cap, pinned vs unpinned two-card runs) if useful.

Disclosure: this report and the measurements were produced with an AI assistant (Claude Sonnet 5.5, medium effort) working under my direction; every number in the README is computed from the included per-run data.

関連リンク

インストール・モデル・リリースへの站内リンク。