Issues / #1008

#1008 Title: 0.1.39 decode is ~15–23% slower than 0.1.38 on AMD (gfx1201/HIP, 2× R9700); prefill unchanged

closed · @uptou888 · 3 Kommentare · Auf GitHub

AMD / HIPModels & quantsLinux

Beschreibung

Environment: same box, same model, same config for both versions. Linux + AMD HIP (ROCm from host),
2× Radeon AI PRO R9700 32GB (gfx1201), 3970x, 128 GB RAM. Model Qwen3.8-Flash-Next IQ3_S (local pack) + MTP.
Args: --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 262144 --kv int8 --kv-resident 32768,
STRATA_HIP_WMMA=1 STRATA_SELECT_WMMA=1, "parallel": 1.

decode (128-token runs, 30K/60K/120K/239K prompts):
  0.1.39: 63.7 / 53.0 / 54.3 / 49.9 t/s      0.1.38: 72.3 / 64.3 / 67.1 / 58.5 t/s   (-15% … -23%)
prefill (same prompts):
  0.1.39: 3290 / 3882 / 4289 / 4264 t/s      0.1.38: 3282 / 3872 / 4276 / 4251 t/s   (identical)
chat: reason 54.2 → 60.6 t/s, creative 52.1 → 55.7 t/s, repeat 95.7 → 97.6 t/s.
Draft acceptance unchanged (88/60/59% vs 90/61/61%) ⇒ not a speculation effect.
Cross-check: a second machine (same GPUs, 0.1.38) reproduces 72.9 / 64.4 / 67.1 / 58.4 ⇒ the regression is version-specific.

Your notes say the zero-doorbell verify runs only when every expert of a layer is in VRAM and that big cards were
not measured; here every expert is in VRAM (2×32 GB) and the AMD path is untested by you ⇒ looks like a regression there.

Also, with "parallel": 2 the engine exits 1 on the same box:
  [strata] the engine stopped unexpectedly (exit code 1). Its last log line: strata verify: captured the batch window over slots 0,1 …
(Can split into a separate issue if you prefer.)

Happy to run variants (STRATA_RING_BYTES=0 / --resident-budget-gib instead of --kv-resident / a debug build) — box available.

Mehr auf der Site

Links zu Install, Modellen, Releases.