Issues / #867
#867 Intel SYCL port on 2x Arc Pro B60: results, and "never rang (graph finished)" on the host-mirror path
closed · @LocalXPU · 4 comments · View on GitHub
BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quants
Description
A B60 result for the port is already in #423 (one card, Q2_0, a fork with its own experiments, from maxious). This one is two cards, the 0.1.39 port with the release's two build fixes, both IQ-quant models, `--layer-split`. On one B60 with host-mirrored experts the Coder stopped at layer 15 until a one-line fix (#866). maxious's fork has fixed the same bug in `verify.cpp`; `session.cpp` still has it there. **Card:** 2x Intel Arc Pro B60 24 GB, PCI `8086:e211`, `xe` driver, Level Zero V2. Ryzen 5 5600, 64 GB DDR4, PCIe 3.0 x8 per card, Ubuntu 24.04, kernel 6.17. **Driver:** compute runtime NEO 26.09.37435.12 (`intel-opencl-icd`, `libze-intel-gpu1`, `intel-ocloc`), `libze1` 1.28.0. **`sycl-ls`:** ``` [level_zero:gpu][level_zero:0] Intel(R) oneAPI Unified Runtime over Level-Zero V2, Intel(R) Arc(TM) Pro B60 Graphics 20.1.0 [1.14.37435+12] [level_zero:gpu][level_zero:1] Intel(R) oneAPI Unified Runtime over Level-Zero V2, Intel(R) Arc(TM) Pro B60 Graphics 20.1.0 [1.14.37435+12] [opencl:gpu][opencl:1] Intel(R) OpenCL Graphics, Intel(R) Arc(TM) Pro B60 Graphics OpenCL 3.0 NEO [26.09.37435.12] [opencl:gpu][opencl:2] Intel(R) OpenCL Graphics, Intel(R) Arc(TM) Pro B60 Graphics OpenCL 3.0 NEO [26.09.37435.12] ``` **oneAPI:** DPC++ 2026.1.1 (`intel-oneapi-compiler-dpcpp-cpp 2026.1.1-325`), oneMKL 2026.1.0, `ocloc ids bmg-g21` = 20.1.0. AOT `-DSTRATA_SYCL_AOT=bmg-g21`, `-j4`, no Docker. **Engine:** `6f32ec0` plus two changes: the compile fix (same as #784 / #809) and the ring-wait fix (#866). Reports `0.1.39-sycl`. **Model and flags:** Coder IQ1_M and Flash-Next IQ2_XS, served through `sycl/serve/server_intel.py`. `--expert-cache auto --stream-experts --prefill auto --spec 4 --spec-min-p 0.5 --mtp <rt> --max-context 8192 --kv int8 --vram-reserve-mib 1536 --no-prefill-borrow --layer-split 24` (Coder) or `25` (IQ2_XS). Environment: `ONEAPI_DEVICE_SELECTOR=level_zero:0,1 SYCL_CACHE_PERSISTENT=0 STRATA_MIRROR_MIB=6144 STRATA_VERIFY_DEVICE_PLAN=1 STRATA_STAGE_TRIM=1 STRATA_VERIFY_NO_HOST=1` (the last only after the log said `100% of the experts resident`). **Numbers** (server `timings`, greedy, one request at a time, every expert in VRAM): | | decode tok/s | prompt tok/s | | --- | --- | --- | | Coder IQ1_M, 2x B60, short prompt | 56.6 (52.8 on the cold first request) | | | Coder IQ1_M, 2x B60, 1,922 tokens of context, then 2,129 | 57.2 after the 1.9K context | 425 and 428 | | IQ2_XS, 2x B60, short prompt | 58.6 and 61.3 | | | IQ2_XS, 2x B60, 1,922 tokens of context, then 2,129 | 55.5 after the 1.9K context | 445 and 436 | | Coder IQ1_M, one B60, 4,042 of 12,288 experts mirrored | 11.9 | | Both models answered coding questions correctly (code executed against known outputs) and held a three-turn conversation. Per-request numbers, VRAM per card, configs and engine logs: #866 (`bench/results/2026-10-04-community-2x-arc-pro-b60/`). Not tested: streaming, tool calls, the Anthropic route, anything beyond 2.2K tokens of context, `--layer-split auto`. **The bug.** On one B60 the Coder stopped at layer 15 on every try with `verify: layer 15 never rang (graph finished)`. The port turned `cudaStreamQuery`'s answer into `DPCT_CHECK_ERROR(cs->ext_oneapi_empty())`, which is always 0, so any per-layer wait over 2 ms was reported as a finished graph. The all-resident path does not wait per layer, which is why the B70 runs never met it. The fix is #866. A system fence in `strata_spin_pause()` (#667) made no difference here, before or after. **Still slow, not fixed:** with the fix, one B60 and mirrored experts decodes at 11.9 tok/s. The stage table shows about 3.9 ms per layer in `wait for rings`, and it scales with `kSpinMax` in `sycl/include/strata/sycl_doorbell.hpp` (25.1 tok/s at 2,000, 28.4 at 200, same text). The GPU does not seem to see the host's flag store during the spin. I did not try `STRATA_VERIFY_COHERENT=1`. **Other things I hit:** - `setvars.sh` stops at `OCL_ICD_FILENAMES: unbound variable` in a shell with `set -u`. - `setup_intel.py` lists the B60 as `e221`; these cards are `e211` (#866). - On a `--layer-split`, the startup warning `N experts are neither in VRAM nor mirrored` counts the other card's layers, and a 6 GiB pinned mirror is allocated that the split does not use. - `iq_parity` needs fixtures (`tools/iq_fixture.py`, with `PYTHONPATH=third_party/llama.cpp/gguf-py`). One run of it did not finish in 300 s on card 0; my timeout killed it and the card did a GT reset (`Schedule disable failed to respond`, recovered in 14 ms, no reboot). The same binary then passed in under 90 s. #809's note about a pinned-host to pageable-host copy hanging the copy engine may be related; I did not check. - `ctest` on one B60: 21 of 25 pass. The four: `iq_parity` (fixtures; passes for all 10 formats once they exist), `ple_parity` (needs the Q2_0 shard), `s2_expert_grouped_parity` (4 bitwise differences, host-reference check within tolerance), `conversation_snapshot_test` (`munmap_chunk(): invalid pointer`, the one #809 reports fixed).
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.