Issues / #870
#870 Intel Arc Pro B60 (e211), 24 GB: first run on 0.1.39 + 4 fixes for current main
closed · @Magh97 · 2 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPModels & quantsDocumentation
Beschreibung
# Arc Pro B60 (e211), 24 GB: first run on 0.1.39 + 4 fixes for current main
## Summary
- First run of the SYCL port on an **Arc Pro B60 (PCI ID `8086:e211`, 24 GB, ASRock)**. It works end to end
(OpenAI + Anthropic APIs, streaming, tool calls) after **4 small fixes** for drift since the last re-migration
(`5047172` "sycl: port 0.1.39"; current main is 88 commits ahead of it).
- Coder IQ1_M, 32K context, INT8 KV, MTP + `--spec 4`: **31 tok/s** decode at setup's defaults, **33-36 tok/s**
after a small expert-cache/reserve change; a 6,029-token prompt read at **657 tok/s**. Kernel parity tests: **8/8
pass on the card** (byte-exact `quantize_act`, `sampler`).
- Two cards in the box (B60 + a B580): the engine must be pointed with
`ONEAPI_DEVICE_SELECTOR=level_zero:1`; the container image pins `level_zero:0`, and setup's config `"gpu": 1`
does not reach the SYCL path. Details below.
Patch with all 4 fixes: `strata-b60-fixes.patch` (4 files, +14/-8), attached. `sycl-ls` output attached.
## Environment
| | |
|---|---|
| GPU 0 (compute) | Arc Pro B60, `8086:e211`, 24.0 GB, subsystem ASRock `6023`, `xe` driver |
| GPU 1 (display) | Arc B580, `8086:e20b`, 12.0 GB, `xe` driver |
| OS / kernel | CachyOS (Arch), `7.2.8-1-cachyos` |
| CPU / RAM | AMD Ryzen 9 5900X, 62 GB |
| Compute runtime | `intel-compute-runtime` 26.35.39758.10, `level-zero-loader` 1.32.0 (Level Zero reports 1.17.39758 host, `+10` in the image) |
| oneAPI (host) | 2026.0.1.27; `icpx` 2026.0.0.20260331 (used to build and to run the binary natively) |
| Image | `strata-sycl-dev` built from `sycl/tools/Dockerfile` with **podman 6.1.2** (via the Docker shim) |
| Build | cmake 4.4.3, ninja 1.13.2, `-DSTRATA_SYCL_AOT=bmg-g21`, 0 errors |
| PCIe | both cards Gen4 x8; engine probe measured **13.6 GB/s host->device** |
```
$ ONEAPI_DEVICE_SELECTOR=level_zero:* sycl-ls
[level_zero:gpu] Intel(R) oneAPI Unified Runtime over Level-Zero V2, Intel(R) Arc(TM) B580 Graphics 20.1.0 [1.17.39758]
[level_zero:gpu] Intel(R) oneAPI Unified Runtime over Level-Zero V2, Intel(R) Arc(TM) Pro B60 Graphics 20.1.0 [1.17.39758]
```
## The 4 fixes current main needs (in `strata-b60-fixes.patch`)
1. **`sycl/setup_intel.py` - the card is `e211`, not `e221`.** `INTEL_ARC` knows `e221` as "Arc Pro B60"; this
card reports `e211` (lspci: `Battlemage G21 [Arc Pro B60] [8086:e211]`). Without a table entry it is named
"Intel GPU e211 (xe)" and sized **32.0 GB** from the BAR instead of 24 GB. Fix: `"e211": ("Arc Pro B60", 24.0)`.
2. **Upstream #626 (thread affinity)** changed `SessionLoopScratch::pinned_core` from a raw int to
`strata::kernels::cpu::ThreadAffinity`. Ported: the field type + `pool.hpp` include in
`sycl/include/strata/core/session.hpp`, and in `sycl/src/core/session.cpp` `pinned = pinned_core.valid` and
`pinned_core = {}` (was `pinned = true`, `pinned_core = -1`).
3. **Upstream PR #559 (batch slots / per-stage dense weights)** changed `NativeDense::load`:
`int64_t layer_lo, int64_t layer_hi` and the `outside()` name filter. Ported into
`sycl/src/core/native_dense.cpp` (compile error without it: "out-of-line definition of 'load' does not match").
4. **`sycl/setup_intel.py` - `write_run_script` now takes 4 arguments.** Current `setup.py:4385` calls
`write_run_script(tag, cfg_path, port, cfg.get("open_browser") is not False)`; the port's replacement took 3 and
died with `TypeError` **after** the 58 GB download and both pack steps, i.e. only at the very end. Fix: accept
`open_browser=True` and pass it through. (The port's "stop with a message instead of writing a wrong config"
guard does not cover this shape of drift.)
## Build and tests
```sh
source /opt/intel/oneapi/setvars.sh
cmake -S sycl -B build-sycl-aot -G Ninja -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx -DSTRATA_SYCL_AOT=bmg-g21
cmake --build build-sycl-aot --target strata -j 24
```
Built and linked with 0 errors; AOT `bmg-g21` runs on the B60. Kernel parity on the card
(`ONEAPI_DEVICE_SELECTOR=level_zero:1`): `quantize_act`, `s_gemv`, `rope`, `router_top10`, `dequant_s2`, `kv_q8`,
`sampler`, `gdn` - **8/8 pass**. `quantize_act` reports 0 bad bytes / 0 differing over 9 distributions; `sampler`
0 of 1632 draws differ. No device loss, no hang.
## End to end
`./setup.sh --backend sycl --yes --family coder --model IQ1_M --context 32768 --no-start`, then
`ONEAPI_DEVICE_SELECTOR=level_zero:1 ./run-coder-iq1_m.sh`. Correct Python on every test; OpenAI
`/v1/chat/completions` (streaming + `tools`), Anthropic `/v1/messages` and the web app all work. Engine log:
```
strata generate: GPU 0: Intel(R) Arc(TM) Pro B60 Graphics, compute capability 20.1
strata generate: PCIe probe: 13.6 GB/s host->device (best of 13.3 13.2 13.6 13.6) -> pcie_frac 0.37
strata generate: expert cache 8742 slots, 16.65 GiB of VRAM; policy is PROFILE, ranked by routing frequency, no eviction.
strata generate: 3546 of 3546 experts missing from VRAM mirrored in pinned host memory (6.77 GiB); the GPU reads them over PCIe
strata serve: 584 MiB of VRAM free with everything loaded
```
### Decode on the B60 (Coder IQ1_M, 32K, INT8 KV, MTP + spec 4)
Same prompt ("Explain what a red-black tree...", 400 tokens, thinking off) and the same draft acceptance
(~72-73%, 244-245 of 335-341) in every row:
| expert cache / reserve | experts resident | host mirror | VRAM free | decode |
|---|---|---|---|---|
| `--expert-cache auto --vram-reserve-mib 1024` (setup default) | 8,238 / 15.68 GiB | 4,050 / 7.73 GiB | 1,540 MiB | 31.0-31.3 tok/s |
| `--expert-cache 8400 --vram-reserve-mib 1800` | 8,742 / 16.65 GiB | 3,546 / 6.77 GiB | 584 MiB | 33.3-33.9 tok/s |
| `--expert-cache 8400 --vram-reserve-mib 1024` | 9,128 / 17.39 GiB | 3,160 / 6.03 GiB | 380 MiB | 35.4-35.8 tok/s |
- Short code bursts reach 40+ tok/s (fewer prose drafts rejected); the first request after a start is slower.
- **~2 tok/s per GiB** of expert set removed from the PCIe mirror on this card - more than the B70 numbers in
`docs/INTEL.md` suggest, and worth knowing for other 24 GB cards.
- A 6,029-token prompt: **656.6 tok/s** read (9.18 s), then 32.6 tok/s on a 50-token answer.
### Cache sizing observations (may be worth a doc line)
- `--expert-cache N` is a **budget** (`N x max_blob`), not a slot count: with a native pack + profile the result is
capped by `free VRAM - --vram-reserve-mib`, so `N = 8400` and `N = 8600` both landed at ~9,130 slots and
~380 MiB free. `--vram-reserve-mib` is the effective margin knob.
- Free VRAM at the end is below the reserve (post-cache allocations - prompt buffers etc. - come from the same
pool), so the reserve does not mean "free at the end".
## Notes
- **Two Intel GPUs:** with a B580 (`level_zero:0`) also present, the image's `ENV ONEAPI_DEVICE_SELECTOR=level_zero:0`
and setup's config `"gpu": 1` disagree: the engine must get `ONEAPI_DEVICE_SELECTOR=level_zero:1` from the
environment or it targets the 12 GB card. `sycl/setup_intel.py`/`strata-sycl.sh` could map the config's `gpu`
index to the selector.
- **`xe_telemetry` link reading:** `/sys/class/drm/card1/device/current_link_speed` and `current_link_width` report
`2.5 GT/s / x1` on this kernel for the GPU device (and for the downstream bridge in front of it), while the
upstream switch port reports `16.0 GT/s / x8` and the engine probe measures 13.6 GB/s. The Monitor tab would show
the wrong PCIe link.
- The runtime image was built and run with **podman** (Docker shim) without problems - `docs/INTEL_ARC.md` says
"Docker"; a one-line "podman works" note may help Arch users.
- Running the binary natively (host oneAPI 2026.0, no container) also works; `quantize_act` is byte-exact on the B60.
## Not tested
Context above 32K on this card; prompts longer than ~6K; `--layer-split` on the mixed B580+B60; `--peer-device`
(stub); images.
[strata-b60-fixes.patch](https://github.com/user-attachments/files/33039120/strata-b60-fixes.patch)
[strata-b60-sycl-ls.txt](https://github.com/user-attachments/files/33039121/strata-b60-sycl-ls.txt)Mehr auf der Site
Links zu Install, Modellen, Releases.