Issues / #1553
#1553 IQ3_XXS tier: engine + server exit silently on ~161k-token prompts at --max-context 262144 (IQ3_S / UD-IQ4_XS unaffected on the same box)
open · @noahark · 4 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
Beschreibung
## Summary The IQ3_XXS tier dies silently — the engine process and the `serve.server` wrapper both exit with **no diagnostic anywhere** — when a request arrives with a ~161k-token prompt, with `--max-context 262144` (inside the model's trained range, no rope scaling). Reproduced 2/2 with full engine restarts between attempts. The other two tiers on the same box, same `--max-context`, same request bodies, are unaffected — including one that shares the pack + 2-shard GSQ-RCO layout. ## Environment - Windows 10 Enterprise LTSC 2021 (19044), dual Xeon E5-2696 v3, 128 GB DDR3L-1600 - Tesla V100-PCIE-32GB (sm_70, TCC), single GPU, driver 582.78 - Strata **v0.1.40.3**, official prebuilt `strata-windows-x64-cuda12.zip` - Tier installed 2026-10-03 (`--model IQ3_XXS --family qwen --vision gpu --context 131072`); for this test only `--max-context` was raised to 262144 in the tier args. Full engine args: ``` --pack D:\strata-data\packs\iq3_xxs --native D:\models\strata\IQ3_XXS\Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf --ple-gguf D:\models\strata\IQ3_XXS\Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00002-of-00002.gguf --expert-profile data\expert-profile.bin --expert-cache auto --prefill auto --spec 4 --mtp D:\strata-data\mtp\rt --max-context 262144 --kv int8 --kv-resident 32768 --vision --vram-reserve-mib 700 --pcie-frac 0.33 --spec-min-p 0.70 ``` ## Repro 1. Start the tier; `/v1/models` reports `n_ctx: 262144`; load is normal (experts 39.97 GiB into RAM, GPU expert cache 14,504 experts / 23.54 GiB — VRAM sits at ~31.7 of 32 GB after load). 2. A ~1k-token chat completion works normally (52-55 tok/s decode). 3. Send a chat completion with a **~161k-token prompt** (synthetic Chinese document summary workload; the exact body tokenizes to 161,633): the connection is reset mid-request, the port stops listening, and both `strata.exe` and the `serve.server` wrapper process are gone. ## Threshold bracket Ascending ladder on a fresh instance — all of these pass: **47,975 / 96,545 / 129,119 / 151,478** prompt tokens (decode 53.4 / 69.4 / 73.3 / 61.2 tok/s). The failure lands between **151,478 and 161,633**. ## Diagnostics — there are none - server stderr log: empty. - server stdout log: ends at the previous request's completion line; the dying request produces **not even the "reading the prompt" progress line**. - Windows Application event log / WER: **no record** for strata.exe at the crash timestamps. ## Isolation on the same box (same `--max-context 262144`, same request bodies) | tier | expert source | installed | 161,633-tok prompt | 258,139-tok prompt | |---|---|---|---|---| | IQ3_XXS | pack `experts.bin` + 2-shard GSQ-RCO (`--native`+`--ple-gguf`) | 2026-10-03 | **engine dies (2/2)** | n/a | | IQ3_S | pack `experts.bin` + 2-shard GSQ-RCO (`--native`+`--ple-gguf`) | 2026-10-05 | OK — 57.8 tok/s | OK — 51.3 tok/s | | UD-IQ4_XS | GGUF-direct 3-shard (`--native` only) | 2026-10-05 | OK — 34.6 tok/s | OK — 36.9 tok/s | So it is **not** the pack path per se (IQ3_S packs and runs 258k fine), not the 2-shard GSQ-RCO mechanism (IQ3_S uses the same layout), and not the prompt size per se (UD/S serve the identical body). What remains specific to the failing tier from the outside: the oldest pack build (2026-10-03, created by the then-current setup version — the other two packs are two days newer), and/or the XXS tier's VRAM headroom at 262144 (it sits closest to the 32 GB ceiling; a prompt-scaled buffer could tip it over without an OOM message reaching the logs). Happy to rerun with an instrumented build, a freshly rebuilt IQ3_XXS pack, `--kv q4_0`, or a capped `--expert-cache` if any of those would discriminate. *(Bench protocol: the synthetic-doc summary workload from the community bench kit — 400-token generation, `reasoning_budget_tokens` 400, seed 42, server-side timings. Raw runs available on request.)*
Mehr auf der Site
Links zu Install, Modellen, Releases.