Pull requests / #1680
#1680 bench: 3x P4, 23 -> 28 tok/s with exact pipeline output
closed · @CC-David-CC · 0 commentaires · Sur GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Description
**3x Tesla P4: 23.000485 -> 28.606868 tok/s (+24.375063%)** on 4K prose, with exact serial/pipeline token equality. Results-only companion to the minimal correctness fix [#1674](https://github.com/Niko1221/Strata/pull/1674). The pipeline speedup comes from [#1656](https://github.com/Niko1221/Strata/pull/1656); the guard comes from earlier three-P4 work in [#1154](https://github.com/Niko1221/Strata/pull/1154). Both measured arms use #1656 head `06abdfa` plus the exact guard published in #1674. | Workload | Prompt tokens | Generated tokens | Serial tok/s | Pipeline tok/s | Gain | | --- | ---: | ---: | ---: | ---: | ---: | | Prose | 4,096 | 256 | 23.000485 | 28.606868 | +24.375063% | | Prose | 512 | 256 | 23.377928 | 29.103477 | +24.491258% | | Code | 512 | 125 | 25.461878 | 43.067806 | +69.146224% | | Count | 512 | 256 | 26.400734 | 46.913942 | +77.699384% | Three measured repetitions per arm, medians above. Same binary, expert placement, and prompts; zero prompt reuse. IQ3_XXS weights, int8 KV, split `19,37`, greedy, MTP spec 2, static expert placement, 27 CPU workers. Measured 2026-10-09 with CUDA 12.0 on a dual-Xeon E5-2697 v3 R730xd. 4K prose median whole-request latency: **33.755897 -> 31.510369 seconds (-6.652252%)**. Decode throughput is measured separately from prompt processing and whole-request time. All **32 fixed-build requests** matched their serial reference, including warmups and forced-rollback/no-guess controls. All 512-token cases also matched the original unmodified serial output. Adds one report folder with per-run JSON, median/range summaries, exact prompt/output tokens, engine timing log, configuration, asset hashes, and a replay driver. No engine changes. See `bench/results/2026-10-09-community-three-p4-pipeline/README.md`. Scope: this CUDA IQ3_XXS workload only; Q8, sampling, images, dynamic adaptation, longer contexts, HIP/SYCL, and general answer quality were not tested. Model revision/full GGUF hashes were not recorded; shard sizes and supporting asset hashes are included.
Sur le site
Liens install, modèles, releases.