Pull requests / #1680

#1680 bench: 3x P4, 23 -> 28 tok/s with exact pipeline output

closed · @CC-David-CC · 0 comentarios · En GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Descripción

**3x Tesla P4: 23.000485 -> 28.606868 tok/s (+24.375063%)** on 4K prose, with exact serial/pipeline token equality.

Results-only companion to the minimal correctness fix [#1674](https://github.com/Niko1221/Strata/pull/1674). The pipeline speedup comes from [#1656](https://github.com/Niko1221/Strata/pull/1656); the guard comes from earlier three-P4 work in [#1154](https://github.com/Niko1221/Strata/pull/1154). Both measured arms use #1656 head `06abdfa` plus the exact guard published in #1674.

| Workload | Prompt tokens | Generated tokens | Serial tok/s | Pipeline tok/s | Gain |
| --- | ---: | ---: | ---: | ---: | ---: |
| Prose | 4,096 | 256 | 23.000485 | 28.606868 | +24.375063% |
| Prose | 512 | 256 | 23.377928 | 29.103477 | +24.491258% |
| Code | 512 | 125 | 25.461878 | 43.067806 | +69.146224% |
| Count | 512 | 256 | 26.400734 | 46.913942 | +77.699384% |

Three measured repetitions per arm, medians above. Same binary, expert placement, and prompts; zero prompt reuse. IQ3_XXS weights, int8 KV, split `19,37`, greedy, MTP spec 2, static expert placement, 27 CPU workers. Measured 2026-10-09 with CUDA 12.0 on a dual-Xeon E5-2697 v3 R730xd.

4K prose median whole-request latency: **33.755897 -> 31.510369 seconds (-6.652252%)**. Decode throughput is measured separately from prompt processing and whole-request time.

All **32 fixed-build requests** matched their serial reference, including warmups and forced-rollback/no-guess controls. All 512-token cases also matched the original unmodified serial output.

Adds one report folder with per-run JSON, median/range summaries, exact prompt/output tokens, engine timing log, configuration, asset hashes, and a replay driver. No engine changes. See `bench/results/2026-10-09-community-three-p4-pipeline/README.md`.

Scope: this CUDA IQ3_XXS workload only; Q8, sampling, images, dynamic adaptation, longer contexts, HIP/SYCL, and general answer quality were not tested. Model revision/full GGUF hashes were not recorded; shard sizes and supporting asset hashes are included.

En el sitio

Enlaces a install, modelos, releases.