Issues / #1302
#1302 0.1.40 SYCL port on 2x Arc Pro B65: expert-cache fill stalls (1-core spin + xe "Timedout job" + device coredump); 0.1.38-based engine works on the same card
open · @james151br · 1 comments · View on GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
Description
## Environment
| | |
|---|---|
| GPUs | 2x Intel Arc Pro B65 32 GB (8086:e222, Battlemage G31), PCIe Gen4 x16, `xe` driver, Level Zero V2 (U/R 1.14.37020) |
| Box | Proxmox VM, Ubuntu 26.04.1 LTS, kernel 7.0.0-38-generic, AMD EPYC 7452 (no AVX-512), 96 GB RAM, SATA SSD |
| oneAPI | DPC++/C++ compiler 2026.1.1 (native host build, no Docker) |
| Build | tag `v0.1.40.1` (latest release; the engine self-reports `0.1.40-sycl` — the port's internal version, since the `.1` hotfixes are serve-layer only) + `sycl/` from PR #1111 (commit `3f37281`), clean — no local patches |
| CMake | `cmake -S sycl -B build-sycl -G Ninja -DCMAKE_CXX_COMPILER=icpx -DCMAKE_C_COMPILER=icx -DCMAKE_BUILD_TYPE=Release` — no AOT (JIT SPIR-V) |
| Model | Qwen3.8-Flash-Next IQ2_XS (GSQ-RCO native pack + PLE), shipped `expert-profile.bin` (24576 ranked pairs) |
| Flags | `--expert-cache auto` `--prefill auto` `--spec 4` `--mtp` `--max-context 131072` `--kv int8` `--pcie-frac 0.55` `--spec-min-p 0.70` |
| Serve | `sycl/serve/server_intel.py --engine strata --config ... --host 0.0.0.0 --port 8095` under systemd |
GPU0 of the two serves a different model (vLLM, ~30 GB of its 32 GB VRAM in use); the Strata tests ran on GPU1, the empty card.
## What happens
Two runs of the clean 0.1.40 SYCL build, both fail, in different phases:
**Run A — no `ONEAPI_DEVICE_SELECTOR`** (engine defaulted to `level_zero:0` = GPU0, the busy card):
- Loaded normally and reached `session is up (engine 0.1.40-sycl)` (`strata serve: 627 MiB of VRAM free with everything loaded`).
- First request (short prompt, 64 tokens):
```
strata serve: verify: timed out at layer 1; its GPU waits were released but the GPU did not finish within 5 s (#267)
terminate called after throwing an instance of 'sycl::_V1::exception'
what(): level_zero backend failed with error: 20 (UR_RESULT_ERROR_DEVICE_LOST)
```
→ engine exits (code -6), systemd restarts it. `xe` faults on GPU0 in the kernel log.
**Run B — `ONEAPI_DEVICE_SELECTOR=level_zero:1`** (GPU1, the correct/empty card):
- `strata generate: loaded 33.02 GiB at 0.76 GiB/s`, then the last line ever written to the engine log:
```
strata generate: expert cache auto: 25.63 GiB free, 700 MiB reserved (+143 MiB for the draft head) -> 17631 slots
```
- ~30 s later, the kernel log on that card:
```
Oct 06 22:06:33 xe 0000:03:00.0: [drm] Tile0: GT0: Fault response: Unsuccessful -EINVAL
Oct 06 22:06:33 xe 0000:03:00.0: [drm] Tile0: GT0: Timedout job: seqno=4294967200, lrc_seqno=4294967200, guc_id=6, flags=0x20 in strata [225676]
Oct 06 22:06:33 xe 0000:03:00.0: [drm] Xe device coredump has been created
Oct 06 22:06:33 xe 0000:03:00.0: [drm] Check your /sys/class/drm/card2/device/devcoredump/data
```
- The engine process then spins at ~95% on one core indefinitely (health endpoint never answers); I stopped the service ~11 min after the fill line. Never reaches `session is up`.
**Control — same card, 0.1.38-based engine** (community SYCL build; engine self-reports the port's `0.1.0` placeholder, upstream 0.1.38 `99f3dbd`):
- `experts loaded: 33.02 GiB at 0.32 GiB/s (116 s so far)` → `filling the GPU's expert cache (18426 experts, 24.75 GiB of VRAM)` → `session is up` **3 seconds later**.
- Serves steadily since (~44 tok/s decode, 128K context).
So on this card: the 0.1.38-based SYCL engine is ~2 min to ready and stable; the 0.1.40 SYCL port stalls in the expert-cache-fill region and the `xe` driver reports a timed-out job plus a device coredump.
**Low-confidence data point (not a clean build):** a separate 0.1.40.1 AOT build (`STRATA_SYCL_AOT=bmg-g21`) that an LLM assembled with a few manual `sed` patches over the port (3 files: `sycl/src/kernels/cuda/fused_gr.dp.cpp`, `core/expert_source.cpp`, `prefill/gemm.dp.cpp`) showed the same stall after the same log line, with `strace -c` showing 61,728 `sched_yield` calls / 99.98% of all syscalls over 6 s — a pure spin loop, no I/O. Corroborating only.
## Also: GPU selection is not honored on the serve path
On the 0.1.40 SYCL port, `--gpu-pci` / `STRATA_GPU_PCI` from the config is ignored (no code path in `sycl/serve/` or `sycl/src/` consumes it — the 0.1.38-based engine does honor it). Without `ONEAPI_DEVICE_SELECTOR` in the environment the engine takes `level_zero:0`; on a two-card box that silently picks the first card — in my case the one already serving another model (Run A above). INTEL.md notes that `sycl/serve/strata-sycl.sh` forwards the host's selector, but launching `server_intel.py` directly (systemd) requires setting `ONEAPI_DEVICE_SELECTOR` on the unit.
## Related
- #1054 (2x B70 `--layer-split` startup deadlock — same "host thread spins" signature; I have no layer split)
- #867 (2x B60 "never rang (graph finished)" on the host-mirror path)
- #267 (verify-window stall → unrecoverable device-lost state; Run A hit the `#267` message)
- #955 (B65 Gen4 on patched v0.1.40 — works for the reporter; I'm on the clean tag + PR #1111)
Happy to share the device coredump (`/sys/class/drm/card2/device/devcoredump/data`), `sycl-ls` output, or run further probes/bisect if you point me at the likely culprits between the 0.1.38 port and the 0.1.40 port.Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.