Issues / #871

#871 Fresh prompts over --short-read decode !!!! (token 0) when the GPU holds 100% of the experts (48 GB card)

closed · @etzerg · 2 Kommentare · Auf GitHub

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsWindows

Beschreibung

<!-- Draft for https://github.com/Niko1221/Strata/issues/new — not yet posted. -->

# Fresh prompts over `--short-read` decode `!!!!` (token 0) when the GPU holds 100% of the experts (48 GB card)

## Summary

On a card large enough to hold **every** expert of IQ2_XS in VRAM, any fresh prompt whose new text exceeds
`--short-read` (default 64 tokens) answers a run of `!` (token 0). Shorter prompts, and requests that hit the
prompt cache, answer normally. Capping the expert cache below 100% residency makes it go away with no other
change. This looks like the all-resident ("zero-doorbell", #646) path on the **batched prompt read**, the same
family as #792, which notes that BATCHING.md was measured only on cards below 100% residency. A 24 GB card never
reaches full residency with IQ2_XS, which would explain why this has not been reported.

## Environment

- Strata `6f32ec0`, ready-made engine **0.1.39** (sm_75/86/89/120, CUDA 13.0), Windows 11 Pro 26200
- GPU: RTX 4090 **48 GB** (memory-modded board), compute capability 8.9, driver 591.86
- CPU: Intel i7-14700K (AVX2, no AVX-512, 8P+12E, `--pool-workers 13` chosen by setup); RAM: 64 GB
- Model: Qwen3.8-Flash-Next GSQ-RCO **IQ2_XS** (original family)

## Reproduction (untouched install)

1. Fresh clone, `START-HERE.bat --yes` (no other flags). Setup's settings:
   `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 131072 --kv int8 --kv-resident 32768 --pool-workers 13`
2. Startup log: `filling the GPU's expert cache (24576 experts, 33.02 GiB of VRAM)` /
   `expert cache 24576 slots, 33.02 GiB of VRAM`
3. Send a fresh ~350-token user message (any content; a random prefix avoids the prompt cache),
   `reasoning_effort: "none"`.

Result: 4/4 requests answer `!!!!!!!!…`. Sampling makes no difference (temperature 0, 0.7, absent).
The serve log for the bad requests shows every draft accepted:

```
strata serve: prompt 354 tokens = 0 reused + 354 read in 390 ms (907.6 tok/s), 40 generated in 219 ms (182.8 tok/s), drafts accepted 18 of 18, 1 checkpoints
```

## What narrows it down

**Threshold is exactly `--short-read`.** Same prompt truncated word by word, fresh each time:

| prompt tokens (incl. template) | result |
| ---: | --- |
| 28, 35, 70 | ok |
| 72, 75, 76, 80, 92, 137, 237, 354 | `!!!!` |

**The decode path is fine.**
- `--short-read 1000000` (every fresh part through the decode windows): 0/3 bad, but a 40K prompt reads at 352 tok/s.
- Repeating the same prompt (prompt cache hit, only the tail read through the decode path) answers correctly, even
  though the reused KV came from the batched read. So the KV written by the batched read looks fine; what breaks
  is the handoff from the batched read to the first decoded token.

**It depends on residency, not on anything else.** Same config, only `--expert-cache` changed, no other workaround:

| `--expert-cache` | slots the engine allocated | fresh 354-token prompt |
| --- | ---: | --- |
| auto | 24576 (100%) | 4/4 bad |
| 24000 | 24576 (rounded up to 100%) | 4/4 bad |
| 23000 | 24076 (98%) | 0/4 bad |
| 20000 | 20939 (85%) | 0/4 bad |

With 23000: 40K prompt 3,780 tok/s, decode 195–217 tok/s, which is what we run now.

**Ruled out (still bad with each):** low-RAM resident mode, `--low-ram mmap`, normal RAM mode; `--vision` on/off;
KV streaming on/off (`--max-context 32768`, no `--kv-resident`); `--spec 2`; `--pcie-frac 0`;
`STRATA_RESIDENT_PIN=0`, `STRATA_ARGMAX_MULTI=0`, `STRATA_PF_FUSED=0`, `STRATA_PLE_BATCH=0`.
`CUDA_LAUNCH_BLOCKING=1` makes the engine exit with 0xC0000409 (stall dump written).

## Two smaller points

- `--expert-cache 24000` rounds up to the full 24,576, so a user picking a "just under" value to avoid the
  all-resident path lands back on it. Some note or a hard cap below the total would help.
- With a 48 GB card the default `auto` always fills the cache, so every fresh prompt over 64 tokens fails out of the
  box. Most clients (agents with a system prompt plus tool schemas) never send one under 64 tokens.

Happy to run a debug build or `STRATA_*` switches on this machine if useful.

Mehr auf der Site

Links zu Install, Modellen, Releases.