Issues / #867

#867 Intel SYCL port on 2x Arc Pro B60: results, and "never rang (graph finished)" on the host-mirror path

closed · @LocalXPU · 4 Kommentare · Auf GitHub

BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quants

Beschreibung

A B60 result for the port is already in #423 (one card, Q2_0, a fork with its own experiments, from maxious). This one is two cards, the 0.1.39 port with the release's two build fixes, both IQ-quant models, `--layer-split`. On one B60 with host-mirrored experts the Coder stopped at layer 15 until a one-line fix (#866). maxious's fork has fixed the same bug in `verify.cpp`; `session.cpp` still has it there.

**Card:** 2x Intel Arc Pro B60 24 GB, PCI `8086:e211`, `xe` driver, Level Zero V2. Ryzen 5 5600, 64 GB DDR4, PCIe 3.0 x8 per card, Ubuntu 24.04, kernel 6.17.

**Driver:** compute runtime NEO 26.09.37435.12 (`intel-opencl-icd`, `libze-intel-gpu1`, `intel-ocloc`), `libze1` 1.28.0.

**`sycl-ls`:**

```
[level_zero:gpu][level_zero:0] Intel(R) oneAPI Unified Runtime over Level-Zero V2, Intel(R) Arc(TM) Pro B60 Graphics 20.1.0 [1.14.37435+12]
[level_zero:gpu][level_zero:1] Intel(R) oneAPI Unified Runtime over Level-Zero V2, Intel(R) Arc(TM) Pro B60 Graphics 20.1.0 [1.14.37435+12]
[opencl:gpu][opencl:1] Intel(R) OpenCL Graphics, Intel(R) Arc(TM) Pro B60 Graphics OpenCL 3.0 NEO  [26.09.37435.12]
[opencl:gpu][opencl:2] Intel(R) OpenCL Graphics, Intel(R) Arc(TM) Pro B60 Graphics OpenCL 3.0 NEO  [26.09.37435.12]
```

**oneAPI:** DPC++ 2026.1.1 (`intel-oneapi-compiler-dpcpp-cpp 2026.1.1-325`), oneMKL 2026.1.0, `ocloc ids bmg-g21` = 20.1.0. AOT `-DSTRATA_SYCL_AOT=bmg-g21`, `-j4`, no Docker.

**Engine:** `6f32ec0` plus two changes: the compile fix (same as #784 / #809) and the ring-wait fix (#866). Reports `0.1.39-sycl`.

**Model and flags:** Coder IQ1_M and Flash-Next IQ2_XS, served through `sycl/serve/server_intel.py`. `--expert-cache auto --stream-experts --prefill auto --spec 4 --spec-min-p 0.5 --mtp <rt> --max-context 8192 --kv int8 --vram-reserve-mib 1536 --no-prefill-borrow --layer-split 24` (Coder) or `25` (IQ2_XS). Environment: `ONEAPI_DEVICE_SELECTOR=level_zero:0,1 SYCL_CACHE_PERSISTENT=0 STRATA_MIRROR_MIB=6144 STRATA_VERIFY_DEVICE_PLAN=1 STRATA_STAGE_TRIM=1 STRATA_VERIFY_NO_HOST=1` (the last only after the log said `100% of the experts resident`).

**Numbers** (server `timings`, greedy, one request at a time, every expert in VRAM):

| | decode tok/s | prompt tok/s |
| --- | --- | --- |
| Coder IQ1_M, 2x B60, short prompt | 56.6 (52.8 on the cold first request) | |
| Coder IQ1_M, 2x B60, 1,922 tokens of context, then 2,129 | 57.2 after the 1.9K context | 425 and 428 |
| IQ2_XS, 2x B60, short prompt | 58.6 and 61.3 | |
| IQ2_XS, 2x B60, 1,922 tokens of context, then 2,129 | 55.5 after the 1.9K context | 445 and 436 |
| Coder IQ1_M, one B60, 4,042 of 12,288 experts mirrored | 11.9 | |

Both models answered coding questions correctly (code executed against known outputs) and held a three-turn conversation. Per-request numbers, VRAM per card, configs and engine logs: #866 (`bench/results/2026-10-04-community-2x-arc-pro-b60/`). Not tested: streaming, tool calls, the Anthropic route, anything beyond 2.2K tokens of context, `--layer-split auto`.

**The bug.** On one B60 the Coder stopped at layer 15 on every try with `verify: layer 15 never rang (graph finished)`. The port turned `cudaStreamQuery`'s answer into `DPCT_CHECK_ERROR(cs->ext_oneapi_empty())`, which is always 0, so any per-layer wait over 2 ms was reported as a finished graph. The all-resident path does not wait per layer, which is why the B70 runs never met it. The fix is #866. A system fence in `strata_spin_pause()` (#667) made no difference here, before or after.

**Still slow, not fixed:** with the fix, one B60 and mirrored experts decodes at 11.9 tok/s. The stage table shows about 3.9 ms per layer in `wait for rings`, and it scales with `kSpinMax` in `sycl/include/strata/sycl_doorbell.hpp` (25.1 tok/s at 2,000, 28.4 at 200, same text). The GPU does not seem to see the host's flag store during the spin. I did not try `STRATA_VERIFY_COHERENT=1`.

**Other things I hit:**

- `setvars.sh` stops at `OCL_ICD_FILENAMES: unbound variable` in a shell with `set -u`.
- `setup_intel.py` lists the B60 as `e221`; these cards are `e211` (#866).
- On a `--layer-split`, the startup warning `N experts are neither in VRAM nor mirrored` counts the other card's layers, and a 6 GiB pinned mirror is allocated that the split does not use.
- `iq_parity` needs fixtures (`tools/iq_fixture.py`, with `PYTHONPATH=third_party/llama.cpp/gguf-py`). One run of it did not finish in 300 s on card 0; my timeout killed it and the card did a GT reset (`Schedule disable failed to respond`, recovered in 14 ms, no reboot). The same binary then passed in under 90 s. #809's note about a pinned-host to pageable-host copy hanging the copy engine may be related; I did not check.
- `ctest` on one B60: 21 of 25 pass. The four: `iq_parity` (fixtures; passes for all 10 formats once they exist), `ple_parity` (needs the Q2_0 shard), `s2_expert_grouped_parity` (4 bitwise differences, host-reference check within tolerance), `conversation_snapshot_test` (`munmap_chunk(): invalid pointer`, the one #809 reports fixed).

Mehr auf der Site

Links zu Install, Modellen, Releases.