Issues / #1121

#1121 --batch-mtp kills the engine on an IQ3_S (GSQ-RCO) pack: mtp: unsupported native MMVQ GGML type

open · @Benderyu · 7 评论 · 在 GitHub 查看

BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentationWindows

描述

# `--batch-mtp` kills the engine on an IQ3_S (GSQ-RCO) pack: `mtp: unsupported native MMVQ GGML type`

**Summary.** With `"parallel": 4` and `--batch-mtp` on a single RTX 4090, the first concurrent batch
admission ends the engine (exit code 1) and every client gets a 503. `docs/BATCHING.md` says a
`--batch-mtp` run that cannot work should *say so and batch as usual*, so this looks like a missing
fallback rather than just an unsupported combination.

## Environment

- **Engine:** 0.1.40 (`engine/BUILD.json`); Python server from `main` at `82f46a8` (v0.1.40.1)
- **GPU:** RTX 4090 48 GB (modded), driver 617.14, CUDA 13.0 build, PCIe Gen4 x16
- **CPU / RAM:** i9-13900K (8P + 16E), 128 GB DDR5-6000
- **OS:** Windows 11
- **Model:** `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, **IQ3_S**, 2 shards, full-RAM mode
  (`--expert-cache auto` 鈫?18,280 expert slots, 34.76 GiB of VRAM), vision on

## Configuration (abridged)

```jsonc
"parallel": 4,
"args": [
  "--pack", "D:\\Strata-data\\packs\\iq3_s",
  "--native", "...\\IQ3_S-00001-of-00002.gguf",
  "--ple-gguf", "...\\IQ3_S-00002-of-00002.gguf",
  "--expert-profile", "data\\expert-profile.bin", "--expert-cache", "auto",
  "--prefill", "auto:32768",
  "--mtp", "D:\\Strata-data\\mtp\\rt", "--spec", "4", "--spec-min-p", "0.5",
  "--batch-mtp",                       // 鈫?added for this test; all other runs omit it
  "--max-context", "262144", "--kv", "int8", "--kv-resident", "32768"
]
```

## Steps to reproduce

1. Take a working config that already has `"parallel": 4`, `--mtp <dir>` and `--spec 4`.
2. Add `--batch-mtp` to its engine arguments and start the server.
   The engine **accepts** it 鈥?at startup it prints
   `strata mtp: shared draft weights, 50 MiB of private state and buffers` (once per slot, 脳4) and
   `strata verify: batch windows of up to 4 sequences (layers [0, 48))`.
3. Send two `POST /v1/chat/completions` at the same time (any prompt; I used one ~4K-token prompt,
   300 output tokens, `temperature: 0`).

## Actual

The engine prints, once, then exits:

```
strata batch: MTP admission for slot 0 failed: mtp: unsupported native MMVQ GGML type
```

Both clients receive HTTP 503:

```json
{"error": {"type": "server_error", "message": "the engine stopped unexpectedly (exit code 1); the next request restarts it"}}
{"error": {"type": "server_error", "message": "the engine stopped while this request waited; it was not sent; the next request restarts it"}}
```

Completion rate per run (two runs, identical):

| Clients at once | With `--batch-mtp` | Same config, `--batch-mtp` removed |
| ---: | --- | --- |
| 1 | 1/1 ok, 86.6 tok/s | 1/1 ok, 86.6 tok/s |
| 2 | **0/2 ok, engine exit 1** | 2/2 ok, 133.0 tok/s aggregate |
| 4 | **0/4 ok, engine exit 1** | 4/4 ok, 178.5 tok/s aggregate |

It fails on the **first** admission into slot 0, so it is deterministic, not a load or memory race.
Without `--batch-mtp` the same configuration ran 20+ concurrent requests across the three sizes with
zero failures.

## Expected

`docs/BATCHING.md` (lines 22-27):

> On one GPU with MTP (`--mtp` and `--spec`), `--batch-mtp` 鈥?lets each batch slot verify one MTP
> proposal per window. 鈥?**If it cannot run (one slot, no `--mtp`, a layer split or helper GPU) the
> engine says so and batches as usual.**

So the slots should fall back to decoding one token per window (the 0.1.39 behaviour) instead of
taking the engine down, and ideally the reason should appear at start-up rather than on the first
admission.

## Notes that may help

- The pack is the **GSQ-RCO IQ3_S**: per ISTA-DASLab's model card its `ffn_down_exps` lands on
  **Q2_0 / IQ4_NL** (the 640-row shape rules out block-256 K/I formats) and the gate/up matrices are
  IQ3_S / IQ3_XXS. The failure names a missing **native MMVQ GGML type**, which suggests the
  batch-MTP draft path needs a native MMVQ type this pack does not carry 鈥?the same class of gap as
  IQ3_M's Q5_0 downs in `docs/ORCA.md`.
- Also active in this config (probably unrelated): `--kv-resident 32768`, `--ple-io direct`, and the
  experimental speed projection control vector.
- `--batch-mtp` is the one feature that would make `"parallel"` free on a single large card, so a
  graceful fallback (or a clear start-up refusal) would be very welcome.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。