Issues / #1302
#1302 0.1.40 SYCL port on 2x Arc Pro B65: expert-cache fill stalls (1-core spin + xe "Timedout job" + device coredump); 0.1.38-based engine works on the same card
open · @james151br · 1 commentaires · Sur GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
Description
## Environment
| | |
|---|---|
| GPUs | 2x Intel Arc Pro B65 32 GB (8086:e222, Battlemage G31), PCIe Gen4 x16, `xe` driver, Level Zero V2 (U/R 1.14.37020) |
| Box | Proxmox VM, Ubuntu 26.04.1 LTS, kernel 7.0.0-38-generic, AMD EPYC 7452 (no AVX-512), 96 GB RAM, SATA SSD |
| oneAPI | DPC++/C++ compiler 2026.1.1 (native host build, no Docker) |
| Build | tag `v0.1.40.1` (latest release; the engine self-reports `0.1.40-sycl` — the port's internal version, since the `.1` hotfixes are serve-layer only) + `sycl/` from PR #1111 (commit `3f37281`), clean — no local patches |
| CMake | `cmake -S sycl -B build-sycl -G Ninja -DCMAKE_CXX_COMPILER=icpx -DCMAKE_C_COMPILER=icx -DCMAKE_BUILD_TYPE=Release` — no AOT (JIT SPIR-V) |
| Model | Qwen3.8-Flash-Next IQ2_XS (GSQ-RCO native pack + PLE), shipped `expert-profile.bin` (24576 ranked pairs) |
| Flags | `--expert-cache auto` `--prefill auto` `--spec 4` `--mtp` `--max-context 131072` `--kv int8` `--pcie-frac 0.55` `--spec-min-p 0.70` |
| Serve | `sycl/serve/server_intel.py --engine strata --config ... --host 0.0.0.0 --port 8095` under systemd |
GPU0 of the two serves a different model (vLLM, ~30 GB of its 32 GB VRAM in use); the Strata tests ran on GPU1, the empty card.
## What happens
Two runs of the clean 0.1.40 SYCL build, both fail, in different phases:
**Run A — no `ONEAPI_DEVICE_SELECTOR`** (engine defaulted to `level_zero:0` = GPU0, the busy card):
- Loaded normally and reached `session is up (engine 0.1.40-sycl)` (`strata serve: 627 MiB of VRAM free with everything loaded`).
- First request (short prompt, 64 tokens):
```
strata serve: verify: timed out at layer 1; its GPU waits were released but the GPU did not finish within 5 s (#267)
terminate called after throwing an instance of 'sycl::_V1::exception'
what(): level_zero backend failed with error: 20 (UR_RESULT_ERROR_DEVICE_LOST)
```
→ engine exits (code -6), systemd restarts it. `xe` faults on GPU0 in the kernel log.
**Run B — `ONEAPI_DEVICE_SELECTOR=level_zero:1`** (GPU1, the correct/empty card):
- `strata generate: loaded 33.02 GiB at 0.76 GiB/s`, then the last line ever written to the engine log:
```
strata generate: expert cache auto: 25.63 GiB free, 700 MiB reserved (+143 MiB for the draft head) -> 17631 slots
```
- ~30 s later, the kernel log on that card:
```
Oct 06 22:06:33 xe 0000:03:00.0: [drm] Tile0: GT0: Fault response: Unsuccessful -EINVAL
Oct 06 22:06:33 xe 0000:03:00.0: [drm] Tile0: GT0: Timedout job: seqno=4294967200, lrc_seqno=4294967200, guc_id=6, flags=0x20 in strata [225676]
Oct 06 22:06:33 xe 0000:03:00.0: [drm] Xe device coredump has been created
Oct 06 22:06:33 xe 0000:03:00.0: [drm] Check your /sys/class/drm/card2/device/devcoredump/data
```
- The engine process then spins at ~95% on one core indefinitely (health endpoint never answers); I stopped the service ~11 min after the fill line. Never reaches `session is up`.
**Control — same card, 0.1.38-based engine** (community SYCL build; engine self-reports the port's `0.1.0` placeholder, upstream 0.1.38 `99f3dbd`):
- `experts loaded: 33.02 GiB at 0.32 GiB/s (116 s so far)` → `filling the GPU's expert cache (18426 experts, 24.75 GiB of VRAM)` → `session is up` **3 seconds later**.
- Serves steadily since (~44 tok/s decode, 128K context).
So on this card: the 0.1.38-based SYCL engine is ~2 min to ready and stable; the 0.1.40 SYCL port stalls in the expert-cache-fill region and the `xe` driver reports a timed-out job plus a device coredump.
**Low-confidence data point (not a clean build):** a separate 0.1.40.1 AOT build (`STRATA_SYCL_AOT=bmg-g21`) that an LLM assembled with a few manual `sed` patches over the port (3 files: `sycl/src/kernels/cuda/fused_gr.dp.cpp`, `core/expert_source.cpp`, `prefill/gemm.dp.cpp`) showed the same stall after the same log line, with `strace -c` showing 61,728 `sched_yield` calls / 99.98% of all syscalls over 6 s — a pure spin loop, no I/O. Corroborating only.
## Also: GPU selection is not honored on the serve path
On the 0.1.40 SYCL port, `--gpu-pci` / `STRATA_GPU_PCI` from the config is ignored (no code path in `sycl/serve/` or `sycl/src/` consumes it — the 0.1.38-based engine does honor it). Without `ONEAPI_DEVICE_SELECTOR` in the environment the engine takes `level_zero:0`; on a two-card box that silently picks the first card — in my case the one already serving another model (Run A above). INTEL.md notes that `sycl/serve/strata-sycl.sh` forwards the host's selector, but launching `server_intel.py` directly (systemd) requires setting `ONEAPI_DEVICE_SELECTOR` on the unit.
## Related
- #1054 (2x B70 `--layer-split` startup deadlock — same "host thread spins" signature; I have no layer split)
- #867 (2x B60 "never rang (graph finished)" on the host-mirror path)
- #267 (verify-window stall → unrecoverable device-lost state; Run A hit the `#267` message)
- #955 (B65 Gen4 on patched v0.1.40 — works for the reporter; I'm on the clean tag + PR #1111)
Happy to share the device coredump (`/sys/class/drm/card2/device/devcoredump/data`), `sycl-ls` output, or run further probes/bisect if you point me at the likely culprits between the 0.1.38 port and the 0.1.40 port.Sur le site
Liens install, modèles, releases.