Pull requests / #846
#846 Keep MTP drafting in concurrent batch slots
closed · draft · @rkcth · 0 comentarios · En GitHub
BenchmarksServer & APIAMD / HIPWindows
Descripción
## Problem Batch slots currently verify one token per request, even when solo decoding uses MTP drafts. On a single GPU this can give up much of the solo decode speed as soon as a second request arrives. ## Change With `MULTI_CONCURRENCY=TRUE`, each active batch slot drafts one token and verifies the current token plus that proposal in a shared window. Each slot keeps private draft KV/state while borrowing the immutable MTP weights and head. The verifier commits only the accepted prefix for each slot; more slots rotate through its eight-row windows. The existing batch path remains the default. The diff also bounds captured batch-graph layouts, documents the opt-in setting and hardware limits, and removes evaluation-only probes and diagnostics. The comments focus on shared ownership, accepted-prefix state, and slot rotation. ## Validation - HIP build passed on an AMD R9700; `python -m unittest serve.test_parallel` passed (11 tests). - Earlier guarded R9700 trials on the same implementation at 8K prompt / 1024 output tokens, matched expert cache and MTP on both sides: two concurrent requests produced identical text and improved aggregate output from 27.48 to 33.63 tokens/s (+22.4%). Three requests produced identical text at 30.72 to 31.60 tokens/s. Four produced identical text but slowed from 30.68 to 29.06 tokens/s. These are single runs, not a cross-hardware throughput claim. - Earlier guarded trials also covered nine rotating slots, two 115K-token prompts with KV streaming, fixed-seed sampled output, cancellation/replacement, and graph eviction. The PR cleanup removed only test switches and internal notes; it did not get a separate GPU throughput run. ## Limits / review focus Grouped MTP is currently single-GPU only and verifies one proposal per slot. It has **not** been tested on 16 GB cards or NVIDIA. Per-slot draft state and buffers reduce VRAM available to the expert cache; a smaller card may see less benefit or fail to fit the requested slot count. A longer mixed-workload soak and cross-hardware checks are still needed before enabling this by default.
En el sitio
Enlaces a install, modelos, releases.