Pull requests / #1171

#1171 docs: minor prefill config change results in faster prefill

closed · @CC-David-CC · 0 评论 · 在 GitHub 查看

BenchmarksNVIDIA / CUDADocumentation

描述

![Q8 prefill results](https://raw.githubusercontent.com/CC-David-CC/Strata-a5500/a810bf5e5f8cde4de2c6136bff90a637249ab657/bench/results/2026-10-06-q8-prefill8192-rtxpro/overview.png)

Changing `--prefill 1024` to `--prefill 8192` on unmodified **Strata 0.1.40** reduced Q8 prefill time at **32,768 input tokens**:

| Metric | Chunk 1,024 | Chunk 8,192 |
|---|---:|---:|
| Prefill | 52.87 s | **7.90 s (6.70x faster)** |
| Total request | 59.71 s | **14.85 s** |
| Output throughput | 149.8 tok/s | 147.3 tok/s |

**Hardware:** RTX PRO 6000 Blackwell Workstation Edition 96 GB, 400 W; Ryzen 9 7950X; 128 GB RAM.

**Model/settings:** Unsloth Qwen3.8-Flash-Next Q8_0, FP16 KV, MTP T4, ngram off, 1,024 output tokens, 15,472 GPU expert slots, 56 GiB resident expert budget, prefill borrowing disabled. Same weights, expert cache and adaptation settings in both runs.

[Configuration and measurements](https://github.com/CC-David-CC/Strata-a5500/blob/a810bf5e5f8cde4de2c6136bff90a637249ab657/bench/results/2026-10-06-q8-prefill8192-rtxpro/README.md)

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。