Issues / #1644
#1644 Windows HIP (RX 7900 XT, gfx1100): every generation fails with "mtp prefill: unspecified launch failure" on a GSQ-RCO IQ3_XXS native pack (0.1.41 and 0.1.40.4)
open · @johnmiu · 0 Kommentare · Auf GitHub
Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
Beschreibung
## Summary
On Windows 11 with a Radeon RX 7900 XT (gfx1100), Strata loads the full IQ3_XXS model, serves `/v1/models` and `/health`, and passes `strata-device --selftest` — but **no generation ever completes**. The first request returns HTTP 400 with:
```json
{"error": {"type": "invalid_request_error", "message": "mtp prefill: unspecified launch failure"}}
```
The engine's own GPU kernel faults inside the **MTP draft layer's prompt pass**, about 0.5 s into the first request. The same fault reproduces with **both engine 0.1.41 and 0.1.40.4** (identical symptom, identical final log line), so it does not look like a 0.1.41 regression.
Because a native (IQ) pack requires the MTP drafter, there is **no configuration that reaches a generated token** — the engine refuses to start without `--mtp`/`--spec`:
```
strata generate: C:\...\packs\iq3_xxs is a native (IQ) pack: it needs --native SHARD1, --spec T (T >= 2) and --prefill CHUNK
```
This is the same shape as #273 (gfx1100, Linux, 0.1.29 — "no working configuration"), now on Windows with the ready-made HIP engine.
## Environment
* OS: Windows 11 (10.0.26100)
* GPU: AMD Radeon RX 7900 XT, 20 GB, gfx1100, wave32 (the box also has an integrated Radeon gfx1036 — pinned out with `HIP_VISIBLE_DEVICES=1`, see below)
* AMD display driver: `32.0.31041.1004`
* CPU: AMD Ryzen 7 7800X3D (AVX-512); RAM: 62 GiB usable (64 GB)
* Strata: git clone of main; engine from the release asset `strata-windows-x64-hip.zip` — tested at **0.1.41** and **0.1.40.4**. Both report `"rocm": "10.2.0a20260930"`, `"hipblaslt_version": 100500`, archs `gfx1100,gfx1101,gfx1102,gfx1200,gfx1201,gfx1030,gfx1151`
* Model: family `qwen`, size `IQ3_XXS` (GSQ-RCO, ISTA-DASLab), context 65536, KV int8, `--spec 4 --spec-min-p 0.5`, `--prefill auto`, `--expert-cache auto`
* Server: `0.0.0.0:8080`, API key set
## What works
* setup's own PC check: `IQ3_XXS needs ~60 GB RAM: fits`
* `strata-device --selftest` (with `HIP_VISIBLE_DEVICES=1`): `device 0: AMD Radeon RX 7900 XT` → `strata-device selftest OK`
* Model load: `experts loaded: 39.97 GiB at 10.45 GiB/s`, `filling the GPU's expert cache (6601 experts, 10.68 GiB of VRAM)`, `session is up (engine 0.1.41)`
* API: `/health`, `/v1/models`, `/v1/status` all correct; HTTP 401 without the key; reachable from the LAN
## Reproduce
1. `START-HERE.bat --family qwen --model IQ3_XXS --host 0.0.0.0 --api-key <key> --yes` (or `serve/server.py --engine strata --config strata-iq3_xxs.json --port 8080` with the config setup wrote)
2. wait for `ready: http://127.0.0.1:8080/v1`
3. send one chat completion (`/v1/chat/completions`, any prompt, any length — 12 tokens fails the same as anything else)
Result: HTTP 400 `mtp prefill: unspecified launch failure`; the engine then exits (code 1) and the server restarts it (`the engine had stopped (exit code 1); starting it again`). Fully deterministic — every start that received a request failed at exactly the same point.
## Engine log at the failure
(engine 0.1.40.4 above; the 0.1.41 run is identical apart from the version string)
```
strata generate: session is up (engine 0.1.40.4)
strata generate: token graph hit path: 6596 resident experts, decided on the device
strata serve: prompt chunk auto: 8192 tokens, a 96-slot ring
strata serve: the prompt path borrows 2105 CUDA0 cache slots (3.37 GiB)
strata hc: CUDA0: the hyper-connection read runs as staged (the norm per token and stream, the down projection's activations staged ahead by cp.async); checked bit for bit against the plain read on this card (STRATA_HC_SPLIT=1 or 0 for the earlier ones)
strata verify: window up to 6 tokens, 69.6 MiB of device buffers
strata mtp: draft head over 106299 tokens (178.4 MiB)
strata serve: 2235 MiB of VRAM free with everything loaded
strata verify: capturing the 6-token window (2215 MiB of VRAM free)
strata verify: captured the 6-token window (upload no error, sync no error)
strata serve: mtp prefill: unspecified launch failure
```
The shape of it: the prompt read completes, the verify window is **captured successfully — including its upload and its sync** (`upload no error, sync no error`) — and then the *draft layer's* prompt pass is the first thing that faults. The server keeps reporting the model as loaded and keeps serving `/v1/models`; it is the engine that dies.
## Configurations tried — all fail identically at the same line
* engine 0.1.41 (default settings)
* engine 0.1.40.4
* no hipBLASLt tuning table (`env` with `STRATA_HIPBLASLT_TUNING` removed → plain hipBLAS for the prompt's dense products)
* `STRATA_MTP_PREFILL_SYNC=1` (the per-group loop, HIP's default) and `=0` (the E-4 batched queue)
* `--vram-reserve-mib 6144` — 5.9 GiB of VRAM free at the failure point instead of 2.2 GiB
* `STRATA_NO_BATCH_KV_STEP=1` + `STRATA_NO_MULTI_GR=1` (the #783 per-token fallbacks)
* `STRATA_STAGER_SSD=0`, `STRATA_HC_SPLIT=1`, `STRATA_PLE_BATCH=0`
* `--spec 2` (a 4-token verify window) and `--spec 4`
* `--prefill 4096` and `--prefill auto` (8192)
* integrated Radeon hidden (`HIP_VISIBLE_DEVICES=1`, so the 7900 XT is HIP device 0 — the engine logs `GPU 0: AMD Radeon RX 7900 XT (gfx1100)` and never mentions gfx1036)
So it is not VRAM pressure, not the hipBLASLt table, not the prefill chunk size, not the verify-window size, and not device selection.
## Notes
* `--prefill auto` picks 8192 on this box; [#1630](https://github.com/Niko1221/Strata/issues/1630) reports a Windows-HIP-7900-XTX problem at that chunk size, but `--prefill 4096` changes nothing for this fault.
* Related reports: [#1630](https://github.com/Niko1221/Strata/issues/1630), [#461](https://github.com/Niko1221/Strata/issues/461), [#273](https://github.com/Niko1221/Strata/issues/273). [#953](https://github.com/Niko1221/Strata/issues/953) (UD-IQ4_XS on a 7900 XTX under Windows, working) suggests the non-native-pack path is fine.
* The box is idle and dedicated to this test — happy to run an instrumented build, a diagnostic env, or a different model size if that helps narrow it.
* Full `strata-iq3_xxs.log` (~10 KB, every run retained) available on request.
Mehr auf der Site
Links zu Install, Modellen, Releases.