Pull requests / #846

#846 Keep MTP drafting in concurrent batch slots

closed · draft · @rkcth · 0 comentarios · En GitHub

BenchmarksServer & APIAMD / HIPWindows

Descripción

## Problem

Batch slots currently verify one token per request, even when solo decoding uses MTP drafts. On a single GPU this can give up much of the solo decode speed as soon as a second request arrives.

## Change

With `MULTI_CONCURRENCY=TRUE`, each active batch slot drafts one token and verifies the current token plus that proposal in a shared window. Each slot keeps private draft KV/state while borrowing the immutable MTP weights and head. The verifier commits only the accepted prefix for each slot; more slots rotate through its eight-row windows. The existing batch path remains the default.

The diff also bounds captured batch-graph layouts, documents the opt-in setting and hardware limits, and removes evaluation-only probes and diagnostics. The comments focus on shared ownership, accepted-prefix state, and slot rotation.

## Validation

- HIP build passed on an AMD R9700; `python -m unittest serve.test_parallel` passed (11 tests).
- Earlier guarded R9700 trials on the same implementation at 8K prompt / 1024 output tokens, matched expert cache and MTP on both sides: two concurrent requests produced identical text and improved aggregate output from 27.48 to 33.63 tokens/s (+22.4%). Three requests produced identical text at 30.72 to 31.60 tokens/s. Four produced identical text but slowed from 30.68 to 29.06 tokens/s. These are single runs, not a cross-hardware throughput claim.
- Earlier guarded trials also covered nine rotating slots, two 115K-token prompts with KV streaming, fixed-seed sampled output, cancellation/replacement, and graph eviction. The PR cleanup removed only test switches and internal notes; it did not get a separate GPU throughput run.

## Limits / review focus

Grouped MTP is currently single-GPU only and verifies one proposal per slot. It has **not** been tested on 16 GB cards or NVIDIA. Per-slot draft state and buffers reduce VRAM available to the expert cache; a smaller card may see less benefit or fail to fit the requested slot count. A longer mixed-workload soak and cross-hardware checks are still needed before enabling this by default.

En el sitio

Enlaces a install, modelos, releases.