Issues / #1008
#1008 Title: 0.1.39 decode is ~15–23% slower than 0.1.38 on AMD (gfx1201/HIP, 2× R9700); prefill unchanged
closed · @uptou888 · 3 commentaires · Sur GitHub
Description
Environment: same box, same model, same config for both versions. Linux + AMD HIP (ROCm from host), 2× Radeon AI PRO R9700 32GB (gfx1201), 3970x, 128 GB RAM. Model Qwen3.8-Flash-Next IQ3_S (local pack) + MTP. Args: --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 262144 --kv int8 --kv-resident 32768, STRATA_HIP_WMMA=1 STRATA_SELECT_WMMA=1, "parallel": 1. decode (128-token runs, 30K/60K/120K/239K prompts): 0.1.39: 63.7 / 53.0 / 54.3 / 49.9 t/s 0.1.38: 72.3 / 64.3 / 67.1 / 58.5 t/s (-15% … -23%) prefill (same prompts): 0.1.39: 3290 / 3882 / 4289 / 4264 t/s 0.1.38: 3282 / 3872 / 4276 / 4251 t/s (identical) chat: reason 54.2 → 60.6 t/s, creative 52.1 → 55.7 t/s, repeat 95.7 → 97.6 t/s. Draft acceptance unchanged (88/60/59% vs 90/61/61%) ⇒ not a speculation effect. Cross-check: a second machine (same GPUs, 0.1.38) reproduces 72.9 / 64.4 / 67.1 / 58.4 ⇒ the regression is version-specific. Your notes say the zero-doorbell verify runs only when every expert of a layer is in VRAM and that big cards were not measured; here every expert is in VRAM (2×32 GB) and the AMD path is untested by you ⇒ looks like a regression there. Also, with "parallel": 2 the engine exits 1 on the same box: [strata] the engine stopped unexpectedly (exit code 1). Its last log line: strata verify: captured the batch window over slots 0,1 … (Can split into a separate issue if you prefer.) Happy to run variants (STRATA_RING_BYTES=0 / --resident-budget-gib instead of --kv-resident / a debug build) — box available.
Sur le site
Liens install, modèles, releases.