Issues / #1606
#1606 Intel Arc A770 (Alchemist, i915): prefill serialized (~47 tok/s, GPU idle, 1 CPU core), decode CPU-bound (all 14 cores at 90%, no VNNI)
open · @Javalopes · 0 commentaires · Sur GitHub
BenchmarksModels & quantsDocumentation
Description
Summary On an Intel Arc A770 (Alchemist DG2) with the i915 driver, the prefill phase is severely serialized: the GPU sits idle (RING_MODE: idle, Awake? 0), only 1-2 CPU cores do any work, and the prefill runs at ~47 tok/s. Expected on similar hardware from docs/INTEL.md is 400+ tok/s. Decode works correctly (uses all 14 physical cores), but is CPU-bound at ~13.5 tok/s because the CPU lacks AVX-512/VNNI. This is a follow-up to #1248 (compilation failures on A-series, now fixed in v0.1.40.3+). System GPU: Intel Arc A770 16 GB (Alchemist / DG2, PCI ID 8086:56a0) Driver: i915 (forced: i915.force_probe=56a0 xe.force_probe=!56a0) CPU: Intel Xeon E5-2680 v4 (14C/28T, Broadwell, 2.4-3.3 GHz, no AVX-512, no VNNI, AVX2 only) RAM: 62.6 GB DDR4-2400 quad-channel SSD: Fanxiang S500PRO 256GB NVMe (~1500 MB/s buffered reads) OS: CachyOS (Arch-based), kernel 7.2.9-1-cachyos, oneAPI DPC++ 2026.1.1 Compute runtime: intel-compute-runtime 26.35.39758.10-1.1 Strata: v0.1.41 (also tested v0.1.40.3), SYCL backend, native build Kernel cmdline includes i915.force_probe=56a0 xe.force_probe=!56a0 because on this kernel the xe driver has "ext_intel_free_memory is not supported" and --expert-cache auto reports 0 slots. With i915, the engine works, but the prefill issue below appears. Config --pack /home/javal/Strata-data/packs/iq3_xxs --native .../Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf --ple-gguf .../Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00002-of-00002.gguf --expert-profile /home/javal/Strata/data/expert-profile.bin --expert-cache 4096 --prefill auto --spec 4 --spec-min-p 0.5 --mtp /home/javal/Strata-data/mtp/rt --max-context 32768 --kv int8 --ple-io direct --vram-reserve-mib 300 5480 experts in VRAM (8.89 GiB). VRAM in use at idle: 15.70 GiB / 15.90 GiB. Env: STRATA_VERIFY_DEVICE_PLAN=1, STRATA_VERIFY_NO_HOST=0, STRATA_STAGER_THREADS=12, SYCL_CACHE_PERSISTENT=0, ZES_ENABLE_SYSMAN=1, ONEAPI_DEVICE_SELECTOR=level_zero:gpu, UR_L0_ENABLE_RELAXED_ALLOCATION_LIMITS=1. FP64 emulation vars present or absent — same results. Problem 1 - Prefill is serialized, GPU is idle Test: 5,756-token prompt, 200-token output. Engine log: [strata] reading the prompt: 5,756 tokens, 10 s so far [strata] reading the prompt: 5,756 tokens, 20 s so far ... [strata] reading the prompt: 2,048 of 5,756 tokens, 46 s so far [strata] reading the prompt: 2,048 of 5,756 tokens, 56 s so far ... [strata] reading the prompt: 4,096 of 5,756 tokens, 83 s so far ... [strata] answering: 6 of max 200 tokens, 14.0 tok/s, 124 s Total prefill: 5,756 tokens in ~123 s = ~47 tok/s (vs 424+ tok/s on the B70 in docs/INTEL.md). Pattern: 30-40 s stalls between 2048-token chunks with no visible progress. i915 engine info captured during prefill (cat /sys/kernel/debug/dri/1/i915_engine_info): rcs0 Awake? 0 RING_MODE: 0x00001200 [idle] RING_HEAD: 0x00000000 RING_TAIL: 0x00000000 ACTHD: 0x00000000_00000000 Requests: (empty) The GPU compute engine is off. Not slow — idle. nvtop during prefill: GPU util: 25%, frequency 717 MHz (max 2400 MHz) CPU: 97% of one core (out of 28 logical CPUs) Thread dump during prefill (ps -T -p <engine> -o tid,stat,wchan:25,pcpu): TID STAT WCHAN %CPU 51113 R<l - 35.4 <- main thread, 35% only ... 51606 S<l futex_wait 30.2 <- 13 workers, all in futex_wait 51607 S<l futex_wait 30.2 ... (13 workers 51606-51618, all 30.2%) ... 51631 S<l anon_pipe_read.cold 0.0 <- blocked on pipe 51632 S<l hrtimer_nanosleep 0.0 The main thread is at 35% (not 100%). The 13 expert-pool workers are all parked in futex_wait. One thread (51631) is stuck in anon_pipe_read.cold — a blocked pipe read. iotop during prefill showed no significant NVMe activity from the strata process. Problem 2 - Decode is CPU-bound (expected, no VNNI), but well parallelized mpstat during decode (mpstat -P ALL 1 45, summary): Average: all 45.25 0.26 3.35 0.11 ... 50.98 Average: 0 89.44 0.00 4.24 0.02 ... 5.98 <- core 0 at 89% Average: 1 90.56 0.00 0.00 0.13 ... 9.09 <- core 1 at 90% ... Average: 13 90.60 0.00 0.00 0.02 ... 9.20 <- core 13 at 90% Average: 14 0.11 0.02 0.29 2.43 ... 96.61 <- HT thread idle ... Average: 27 0.35 0.00 0.89 1.95 ... 96.81 <- HT thread idle All 14 physical cores at ~90%, HT threads idle. nvtop confirms: CPU 1378%, GPU spikes 0-100% at 2400 MHz. Decode at 13.4-16.7 tok/s - consistent with docs/INTEL.md for the A750 (10-15 tok/s), and CPU-bound because the Xeon lacks AVX-512 and VNNI (Broadwell, 2016). Flags tested --host-core first -> WORSE (decode 13.1 vs 16.7 tok/s) --kv-resident 32768 -> WORSE (decode 13.4 vs 16.7; prefill unchanged) --pcie-frac 0 -> NEUTRAL --ple-io mmap -> NEUTRAL vs direct --ple-io ram -> OOM (28 GB table + 47 GB arena > 62 GB RAM) --expert-cache 5120 -> OOM --expert-cache 4096 -> BASELINE (5480 experts, 8.89 GiB VRAM) FP64 emulation vars -> NEUTRAL Version comparison v0.1.40.3: Prefill 5,756 tokens: ~123 s (~47 tok/s) Decode 800 tokens, cache 4096: 13.4 tok/s VRAM idle (cache 4096): 15.70 GiB v0.1.41: Prefill 5,756 tokens: ~123 s (~47 tok/s) Decode 800 tokens, cache 4096: 13.7-13.9 tok/s VRAM idle (cache 4096): 15.70 GiB No regression between versions for this hardware. Release notes mention +22% decode on A750 in 0.1.41 — not observed on A770 (both use DG2 ACM-G10). Hypothesis The prefill thread (51631) blocked in anon_pipe_read.cold, combined with the GPU being idle and only the main thread working, suggests the CPU-GPU doorbell/handshake channel is not delivering during prefill on Alchemist. Possible causes: The pipe that feeds work to the GPU is only populated by one thread (the main one), and that thread stalls in anon_pipe_read waiting for the previous chunk's completion signal The 13 expert workers park in futex_wait because no work is dispatched The GPU gets no commands because the pipe never delivers Cycle repeats per 2048-token chunk -> the visible 30-40 s stalls If a Battlemage card (B70, B580) shares this code path but works, the difference may be in the Level Zero doorbell implementation on Alchemist (i915 vs xe) or in the number of in-flight requests the DG2 can hold. Request Could you (or anyone with Alchemist hardware) reproduce this? Specifically: Large prompt (>4K tokens) on A750 / A770 / A580 with i915 Capture i915_engine_info during prefill — is the GPU also idle? Thread dump — is a thread blocked in anon_pipe_read? I have direct SSH access to the box and can test patches. Additional context The xe driver has a separate issue (ext_intel_free_memory is not supported) that makes --expert-cache auto fail with 0 slots. Worth a separate report — this forced us to use i915. With --expert-cache 4096 (5480 experts), the GPU hits 98% VRAM. A larger cache OOMs. mpstat data and raw i915_engine_info available on request.
Sur le site
Liens install, modèles, releases.