Issues / #857
#857 "parallel": 2 lowers total throughput on a single 48 GB card (100% expert-cache hits): keep MTP drafts in batch windows?
closed · @bobvious · 5 comentários · No GitHub
NVIDIA / CUDAModels & quantsWindows
Descrição
Could batch windows keep MTP drafts? On one card where the experts fit, `"parallel": 2` lowers total throughput, because each slot decodes one token per window. v0.1.39, one RTX PRO 5000 48 GB, Xeon 6148 (AVX2 kernels), Qwen3.8-Flash-Next IQ3_S, `--spec 4 --mtp`, 262K ctx. Decode expert cache hits 100%. Rates are streamed tokens per second, total across all requests: | | no `parallel` | `parallel: 2` | |---|---|---| | 1 request | 97 | 64 | | 2 at once | 103 | 61 | | 4 at once | 100 | 71 | `strata batch: 800 windows, avg 1.99 rows, 16.71 ms/window = run 15.88 (CUDA0 GPU-reach wait 11.34 + CPU experts 3.05) ...` Happy to test a build.
No site
Links install, modelos, releases.