反馈 / #968
#968 0.1.39: UD-IQ4_XS prompts read as garbage on an RTX 5090 (sm_120) through the MMQ prompt path; STRATA_PREFILL_MMQ=0 or an sm_86 card reads them correctly
open · @signalnine · 6 评论 · 去 GitHub 看
BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsDocumentationLinux
说明
On an RTX 5090, any prompt longer than a few dozen tokens comes out garbled for the model when Strata reads it through the MMQ prompt path with Unsloth's UD-IQ4_XS. The model calls the input "corrupted or fragmented", misses a planted fact, and summarizes code as the wrong kind of program. Decoding is fine: 1,500-token answers to short prompts are coherent and correct. `STRATA_PREFILL_MMQ=0` fixes it at the same prompt speed, and the same binary on an RTX 3090 (sm_86) reads the same prompts correctly with MMQ on.
## Setup
- Engine 0.1.39, source build from `main` 6f32ec0 by setup (`BUILD.json`: archs `[86, 120]`, CUDA 13.2), `STRATA_MMQ_KQUANTS=OFF` (setup's default)
- RTX 5090 32 GB (sm_120) + RTX 3090 24 GB (sm_86), driver 580.119.02, Ryzen 9 9900X, 128 GB, Linux 6.18
- `--family unsloth --model UD-IQ4_XS --gguf-dir` on the three shards (SHA-256s match the table in `docs/UNSLOTH_Q4.md`), packed by setup with `--compat-bf16`
- Config as setup wrote it: `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp ... --max-context 32768 --kv int8`, no images
## Repro
Any passage plus a question about it, temperature 0, thinking off:
```sh
txt=$(head -c 1200 docs/MULTI_GPU.md | tr '\n' ' ')
curl -s http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d "$(jq -n --arg f "$txt" \
'{model:"x",messages:[{role:"user",content:("Here is a passage:\n\n"+$f+"\n\nIn one sentence, what is this passage about? Then quote its first five words exactly.")}],max_tokens:100,temperature:0,reasoning_effort:"none"}')" \
| jq -r .choices[0].message.content
```
5090, MMQ on (default), 344-token prompt:
> The text provided is a **corrupted or nonsensical string** that appears to be a mix of: 1. **Technical jargon** (e.g., "GPU," "CUDA," ...
5090 with `STRATA_PREFILL_MMQ=0`, and the 3090 with MMQ on, same prompt:
> This passage explains how Strata implements pipeline parallelism to run a single model across multiple NVIDIA GPUs by splitting layers and managing expert caches. Strata on two or three GPUs
## What I measured
| GPU | MMQ | 113 tok | 344 tok | 19,440-token needle (a passphrase planted mid-file) |
| --- | --- | --- | --- | --- |
| 5090 | on | "your message was cut off" | "a corrupted or nonsensical string" | missed: "incomplete or corrupted" |
| 5090 + 3090 split (auto, K=47) | on | "a prompt injection attempt" | "your message was cut off" | missed: "corrupted or fragmented" |
| 5090 | off | correct | correct | found |
| 3090 | on | correct | correct | found |
- Prompts of 29-41 tokens come out right on the split with MMQ on, so the break sits somewhere between 41 and 113 tokens.
- Prompt speed with MMQ off is the same as with it on: 690 vs 695 tok/s on the same 19K prompt on the 5090.
- Decode with MMQ off: 121 tok/s on a 1,500-token code answer, MTP accepting 1,024 of 1,291 drafts.
The pack's experts are IQ3_S gate/up (IQ4_XS in one layer) with IQ4_NL down (Q8_0 in five layers). I only have this pack, so I can't say whether the native IQ packs hit it too. #402 ran UD-IQ4_XS on an sm_120 card at 0.1.31, but its needle check covered UD-Q4_K_XL only. #420 and #954 are other sm_120 problems in the MMQ prompt path.
Workaround: `"env": {"STRATA_PREFILL_MMQ": "0"}` in `strata-<model>.json`, or the variable in the server's environment.
本站相关内容
相关页面的快捷入口。