贡献 / #1290
#1290 Fix shared MTP slot initialization and speculative row bounds
closed · draft · @jeremiahritchey · 0 评论 · 去 GitHub 看
BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quantsLinux
说明
With `--batch 4 --batch-mtp` and a native draft-vocabulary head, the first slot admission fails with `mtp: unsupported native MMVQ GGML type`. The slot borrows the head pointer but leaves its quantization type at `-1`. This fixes that shared initialization in the existing upstream multislot MTP path; it does not import the older experimental implementation.
The slot also inherits the shared Q4 projections, their offsets, per-stream normalization setting and host vocabulary mapping. Only the owning drafter frees shared projections. At a request's output or context limit, the batch stages one row instead of a speculative second row; this prevents `verify: slot ... runs past its context` at the final cell. The actual row count is used for residual copies and subsequent drafts.
Repeated live HTTP runs also exposed VRAM exhaustion while instantiating batch graphs: the previous 64-layout cache could outgrow the available reserve. Limit it to 16 layouts using the existing LRU eviction. The tested 24 GB configuration uses `--vram-reserve-mib 1536`; the 700 MiB reserve with 64 retained layouts failed after two successful benchmark rounds.
`--batch-mtp` remains opt-in and single-GPU, with one proposal per slot. Additive `INFO batch_mtp=0/1` and an idle-time acceptance count make activation observable. The new GPU regression checks four-slot token parity, short output limits and the final context cell; the existing interleave test covers conversation reuse, preemption and slot-to-solo continuation.
Validation:
- CUDA Release build passed (CUDA 13.4, architecture 89, Linux x86_64).
- 55 focused Python server tests passed: `serve.test_parallel`, `serve.test_slots`, `serve.test_control_cancel`, `serve.test_restart_waiters`.
- Python compilation and `git diff --check` passed.
- GPU regressions passed on an RTX 4090 (24 GB), EPYC 7532 (32 physical cores), 247 GiB RAM, Linux/CUDA 13.4, IQ3_S model with its native Q5_K head:
- `batch_mtp_test.py`: all four slots equal solo for 64 tokens each; limits 1–4 equal solo prefixes; all four slots reach the final context cell. Passed with the native draft head and again with `--mtp-q4 all --mtp-hnorm stream`. Both runs accepted batch drafts.
- `batch_interleave_test.py --max-new 100 --long 4096`: token parity, next-turn reuse, turn checkpoints, BYIELD/resume, and BSTOP followed by solo resume passed.
- `batch_test.py --batch 4 --n 4 --max-new 64 --keys "temperature=0.7 top_k=20 seed=123"`: all four slots equal solo, 96 of 154 batch proposals accepted.
- The test configuration used `--prefill 1024 --vram-reserve-mib 1536` to leave room for the parity tests' owned prompt buffers. The native/Q4 boundary tests override context to 512; the interleave and sampled tests used context 8192. These are correctness runs, not performance measurements.
- Deployed on the same machine with the existing 262144-token context, int8 KV, 32768 resident KV and four slots; `INFO batch_mtp=1` confirmed. Final build `ae706fa` passed repeated four-client HTTP runs, six queued requests, mixed long/short prompts, solo/batch client disconnects, conversation reuse, Responses API and tool-call checks. The run captured 36 batch layouts, exercising eviction beyond the new 16-layout limit, and finished with about 700 MiB free on the GPU.
HTTP measurement on that machine/configuration: one warm-up round, then three rounds of four concurrent requests × 256 tokens, temperature 0, thinking disabled. Median aggregate throughput **110.7 tokens/s**, median time to first token **1.17 s**. This is a local workload measurement, not a general speedup claim.
Reproduction commands with an existing model config containing `--mtp`, `--spec >= 2` and enough VRAM reserved for owned prompt buffers:
```sh
python tools/batch_mtp_test.py --exe build/strata --config strata-model.json
python tools/batch_mtp_test.py --exe build/strata --config strata-model.json \
--extra "--mtp-q4 all --mtp-hnorm stream"
python tools/batch_test.py --exe build/strata --config strata-model.json --batch 4 --n 4 \
--keys "temperature=0.7 top_k=20 seed=123" \
--extra "--batch-mtp --pcie-frac 0 --adapt-every 1000000"
python tools/batch_interleave_test.py --exe build/strata --config strata-model.json \
--extra "--batch-mtp --pcie-frac 0 --adapt-every 1000000 --no-prefill-borrow"
```
本站相关内容
相关页面的快捷入口。