Issues / #857

#857 "parallel": 2 lowers total throughput on a single 48 GB card (100% expert-cache hits): keep MTP drafts in batch windows?

closed · @bobvious · 5 commentaires · Sur GitHub

NVIDIA / CUDAModels & quantsWindows

Description

Could batch windows keep MTP drafts? On one card where the experts fit, `"parallel": 2` lowers total throughput, because each slot decodes one token per window.

v0.1.39, one RTX PRO 5000 48 GB, Xeon 6148 (AVX2 kernels), Qwen3.8-Flash-Next IQ3_S, `--spec 4 --mtp`, 262K ctx. Decode expert cache hits 100%. Rates are streamed tokens per second, total across all requests:

| | no `parallel` | `parallel: 2` |
|---|---|---|
| 1 request | 97 | 64 |
| 2 at once | 103 | 61 |
| 4 at once | 100 | 71 |

`strata batch: 800 windows, avg 1.99 rows, 16.71 ms/window = run 15.88 (CUDA0 GPU-reach wait 11.34 + CPU experts 3.05) ...`

Happy to test a build.

Sur le site

Liens install, modèles, releases.