Issues / #1121
#1121 --batch-mtp kills the engine on an IQ3_S (GSQ-RCO) pack: mtp: unsupported native MMVQ GGML type
open · @Benderyu · 7 comentarios · En GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentationWindows
Descripción
# `--batch-mtp` kills the engine on an IQ3_S (GSQ-RCO) pack: `mtp: unsupported native MMVQ GGML type`
**Summary.** With `"parallel": 4` and `--batch-mtp` on a single RTX 4090, the first concurrent batch
admission ends the engine (exit code 1) and every client gets a 503. `docs/BATCHING.md` says a
`--batch-mtp` run that cannot work should *say so and batch as usual*, so this looks like a missing
fallback rather than just an unsupported combination.
## Environment
- **Engine:** 0.1.40 (`engine/BUILD.json`); Python server from `main` at `82f46a8` (v0.1.40.1)
- **GPU:** RTX 4090 48 GB (modded), driver 617.14, CUDA 13.0 build, PCIe Gen4 x16
- **CPU / RAM:** i9-13900K (8P + 16E), 128 GB DDR5-6000
- **OS:** Windows 11
- **Model:** `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, **IQ3_S**, 2 shards, full-RAM mode
(`--expert-cache auto` 鈫?18,280 expert slots, 34.76 GiB of VRAM), vision on
## Configuration (abridged)
```jsonc
"parallel": 4,
"args": [
"--pack", "D:\\Strata-data\\packs\\iq3_s",
"--native", "...\\IQ3_S-00001-of-00002.gguf",
"--ple-gguf", "...\\IQ3_S-00002-of-00002.gguf",
"--expert-profile", "data\\expert-profile.bin", "--expert-cache", "auto",
"--prefill", "auto:32768",
"--mtp", "D:\\Strata-data\\mtp\\rt", "--spec", "4", "--spec-min-p", "0.5",
"--batch-mtp", // 鈫?added for this test; all other runs omit it
"--max-context", "262144", "--kv", "int8", "--kv-resident", "32768"
]
```
## Steps to reproduce
1. Take a working config that already has `"parallel": 4`, `--mtp <dir>` and `--spec 4`.
2. Add `--batch-mtp` to its engine arguments and start the server.
The engine **accepts** it 鈥?at startup it prints
`strata mtp: shared draft weights, 50 MiB of private state and buffers` (once per slot, 脳4) and
`strata verify: batch windows of up to 4 sequences (layers [0, 48))`.
3. Send two `POST /v1/chat/completions` at the same time (any prompt; I used one ~4K-token prompt,
300 output tokens, `temperature: 0`).
## Actual
The engine prints, once, then exits:
```
strata batch: MTP admission for slot 0 failed: mtp: unsupported native MMVQ GGML type
```
Both clients receive HTTP 503:
```json
{"error": {"type": "server_error", "message": "the engine stopped unexpectedly (exit code 1); the next request restarts it"}}
{"error": {"type": "server_error", "message": "the engine stopped while this request waited; it was not sent; the next request restarts it"}}
```
Completion rate per run (two runs, identical):
| Clients at once | With `--batch-mtp` | Same config, `--batch-mtp` removed |
| ---: | --- | --- |
| 1 | 1/1 ok, 86.6 tok/s | 1/1 ok, 86.6 tok/s |
| 2 | **0/2 ok, engine exit 1** | 2/2 ok, 133.0 tok/s aggregate |
| 4 | **0/4 ok, engine exit 1** | 4/4 ok, 178.5 tok/s aggregate |
It fails on the **first** admission into slot 0, so it is deterministic, not a load or memory race.
Without `--batch-mtp` the same configuration ran 20+ concurrent requests across the three sizes with
zero failures.
## Expected
`docs/BATCHING.md` (lines 22-27):
> On one GPU with MTP (`--mtp` and `--spec`), `--batch-mtp` 鈥?lets each batch slot verify one MTP
> proposal per window. 鈥?**If it cannot run (one slot, no `--mtp`, a layer split or helper GPU) the
> engine says so and batches as usual.**
So the slots should fall back to decoding one token per window (the 0.1.39 behaviour) instead of
taking the engine down, and ideally the reason should appear at start-up rather than on the first
admission.
## Notes that may help
- The pack is the **GSQ-RCO IQ3_S**: per ISTA-DASLab's model card its `ffn_down_exps` lands on
**Q2_0 / IQ4_NL** (the 640-row shape rules out block-256 K/I formats) and the gate/up matrices are
IQ3_S / IQ3_XXS. The failure names a missing **native MMVQ GGML type**, which suggests the
batch-MTP draft path needs a native MMVQ type this pack does not carry 鈥?the same class of gap as
IQ3_M's Q5_0 downs in `docs/ORCA.md`.
- Also active in this config (probably unrelated): `--kv-resident 32768`, `--ple-io direct`, and the
experimental speed projection control vector.
- `--batch-mtp` is the one feature that would make `"parallel"` free on a single large card, so a
graceful fallback (or a clear start-up refusal) would be very welcome.
En el sitio
Enlaces a install, modelos, releases.