Pull requests / #1124

#1124 Community results: Strata 0.1.40 Q4/Q8 on RTX PRO 6000

closed · draft · @CC-David-CC · 0 コメント · GitHub で見る

BenchmarksNVIDIA / CUDA

本文

![Strata 0.1.40 Q4/Q8 results on RTX PRO 6000](https://raw.githubusercontent.com/CC-David-CC/Strata-a5500/69dfa9284f5d4fce4ba6b209159a1a347af419aa/bench/results/2026-10-06-community-rtxpro-v0140/overview.png)

## Strata 0.1.40 results

Measured with the unmodified v0.1.40 engine on an NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96 GB, Ryzen 9 7950X and 128 GB RAM.

Each request processed 8,192 input tokens and generated 512 output tokens. One request at a time, FP16 KV, 16,384-token allocation, ngram off. Model loading is excluded.

| Model / mode | Output tok/s | Prefill | Total request | Effective tok/s |
|---|---:|---:|---:|---:|
| Unsloth Q4_K_XL, MTP T4 | 257.6 | 3.035 s | 5.024 s | 101.9 |
| Unsloth Q4_K_XL, no MTP | 129.9 | 3.115 s | 7.057 s | 72.6 |
| Unsloth Q8_0, MTP T4 | 140.2 | 12.812 s | 16.465 s | 31.1 |
| Unsloth Q8_0, no MTP | 82.6 | 12.790 s | 18.989 s | 27.0 |

Native Q8 PLE loaded successfully in both modes. All four 0.1.40 cases completed the full input/output counts and passed the small generated-code function check.

関連リンク

インストール・モデル・リリースへの站内リンク。