Pull requests / #1137

#1137 Community results: Strata 0.1.40 Q4/Q8, MTP and ngram on RTX PRO 6000

closed · @CC-David-CC · 0 comments · View on GitHub

BenchmarksNVIDIA / CUDA

Description

![Strata 0.1.40: Q4/Q8 with MTP, ngram and both](https://raw.githubusercontent.com/CC-David-CC/Strata-a5500/bacfc4d5723456bbbea59a6cb6a68f972b30e601/bench/results/2026-10-06-community-rtxpro-v0140/ngram.png)

NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96 GB, Ryzen 9 7950X, 128 GB RAM. Unmodified Strata v0.1.40.

8,192 input + 512 output tokens per request; one request at a time; FP16 KV; 16,384-token allocation. Total time includes prefill, excludes model loading.

| Model / path | Output tok/s | Prefill s | Total s | Effective tok/s |
|---|---:|---:|---:|---:|
| Unsloth Q4_K_XL, No speculation | 129.9 | 3.115 | 7.057 | 72.6 |
| Unsloth Q4_K_XL, MTP T4 | 257.6 | 3.035 | 5.024 | 101.9 |
| Unsloth Q4_K_XL, Ngram only | 133.8 | 2.990 | 6.817 | 75.1 |
| Unsloth Q4_K_XL, MTP + ngram | 255.4 | 3.017 | 5.023 | 102.0 |
| Unsloth Q8_0, No speculation | 82.6 | 12.790 | 18.989 | 27.0 |
| Unsloth Q8_0, MTP T4 | 140.2 | 12.812 | 16.465 | 31.1 |
| Unsloth Q8_0, Ngram only | 81.6 | 12.782 | 19.056 | 26.9 |
| Unsloth Q8_0, MTP + ngram | 133.8 | 12.761 | 16.589 | 30.9 |

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.