Pull requests / #1755

#1755 bench: fleet calibration candidates and cache gains up to 22.6%

open · @CC-David-CC · 0 commentaires · Sur GitHub

BenchmarksNVIDIA / CUDADocumentationLinux

Description

Reports the faster candidate settings from six Linux hosts, with raw measurements, ranges, reproducible configurations and the rejected RTX 4090 candidate.

- Three P4s: PCIe/draft confirmation 26.819 -> 29.356 tok/s (+9.46% versus product defaults).
- Decode-cache tuning: RTX 3070 +3.83%, RX 5500 XT +22.59%, RX 7900 XTX +15.91%. These are separate stage gains, not additive.
- PRO 6000 rebuilt on public 0.1.41: all cache policies pass; default/80/160 swaps measure 236.144/236.212/235.984 tok/s. Keep default cache policy. PCIe/draft confirmation: 231.884 -> 240.647 (+3.78%).
- RTX 4090 candidate rejected: full-configuration check was 8.57% slower despite apparent stage wins.

[Full report, exact candidate settings and evidence](https://github.com/CC-David-CC/Strata-a5500/blob/8ec15dc9fe0f55e0199cc7a8b67f958cfe14fee5/bench/results/2026-10-09-fleet-calibration/README.md).

Results only; no engine or default changes. Three rounds for PCIe/draft confirmation, two for worker/cache candidates. Related: #906, #1332, #1337, #566; corrected three-P4 build from #1674.

Sur le site

Liens install, modèles, releases.