Issues / #857
#857 "parallel": 2 lowers total throughput on a single 48 GB card (100% expert-cache hits): keep MTP drafts in batch windows?
closed · @bobvious · 5 评论 · 在 GitHub 查看
NVIDIA / CUDAModels & quantsWindows
描述
Could batch windows keep MTP drafts? On one card where the experts fit, `"parallel": 2` lowers total throughput, because each slot decodes one token per window. v0.1.39, one RTX PRO 5000 48 GB, Xeon 6148 (AVX2 kernels), Qwen3.8-Flash-Next IQ3_S, `--spec 4 --mtp`, 262K ctx. Decode expert cache hits 100%. Rates are streamed tokens per second, total across all requests: | | no `parallel` | `parallel: 2` | |---|---|---| | 1 request | 97 | 64 | | 2 at once | 103 | 61 | | 4 at once | 100 | 71 | `strata batch: 800 windows, avg 1.99 rows, 16.71 ms/window = run 15.88 (CUDA0 GPU-reach wait 11.34 + CPU experts 3.05) ...` Happy to test a build.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。