Issues / #1410

#1410 Support Intel Arc Pro B50 (Xe2 / BMG-G21) as a tested SYCL target — 16 GB at 70 W is the card users are actually buying

open · @ditronicos · 0 コメント · GitHub で見る

BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

本文

## TL;DR

The Intel Arc Pro B50 is the cheapest card that can plausibly host a mid-size MoE
model **entirely in VRAM** (16 GB), draws **70 W with no external power connector**,
is sold in volume with a 3-year warranty, and is already supported by the SYCL
backend's AOT target `bmg-g21` — but only as *experimental*, with no maintainer
hardware, no CI coverage, and detection that rejects it.

I have one in a 24/7 inference box and I'd like to help turn it into a target you
can actually say "supported" about. This issue is the technical case + a concrete
list of what I think it would take.

## Why this card, specifically

Intel Arc Pro B50 (launched Q3'25, Battlemage / Xe2-HPG, TSMC N5, device ID `8086:e212`):

| Spec | Arc Pro **B50** | Arc Pro B60 | Arc Pro B70 |
|---|:---|:---|:---|
| Xe2 cores / RT units | **16 / 16** | 20 / 20 | 32 / 32 |
| XMX engines | **128** | 160 | 256 |
| Memory | **16 GB GDDR6** | 24 GB | 32 GB |
| Bus / bandwidth | **128-bit, 224 GB/s** | 456 GB/s | 608 GB/s |
| FP32 (dense) | **10.65 TFLOPS** | 12.3 | — |
| INT8 peak | **170 pTOPS** | 197 | 367 |
| Total board power | **70 W, no connector** | 120–200 W | ~230 W |
| Interface | **PCIe 5.0 x8** | 5.0 x8 | 5.0 x8 |
| ECC | **yes (per ARK)** | yes | yes |
| Launch price | **$299 SEP / ~$349 street** | $499 | ~$999 |

Physical/engineering points that matter for always-on inference boxes, not just
workstations: **168 × 69 mm, dual-slot, half-length/half-height, 330 g**, minimum
PSU 280 W, and **no power cable at all** — it runs off the slot. It also carries
**4 × DisplayPort 2.1 (UHBR 13.5)** and **two Xe Media engines** (H.264/HEVC/VP9/AV1
encode+decode, incl. 10-bit 4:2:2 decode) — useful if a box does transcode and
inference, irrelevant if it's headless. API surface: oneAPI / Level Zero / OpenCL 3.0 /
Vulkan 1.4 / OpenVINO / IPEX.

Positioning vs. the cards people quote instead: it competes with AMD Radeon Pro
W7500 (8 GB) and NVIDIA RTX A1000/A2000 (8/12 GB) at a similar or lower price with
**double-to-triple the VRAM**, and Phoronix's launch-day Linux testing put it at
**~1.47 × an RTX A1000** across an OpenGL/Vulkan/OpenCL mix on the open-source
stack. The interesting part is that it does that *without* a proprietary kernel
module: it runs on the upstream **`xe`** KMD with `bmg_guc_70.bin` / `bmg_huc.bin`.

For a quantised-MoE engine like Strata the relevant trade is:
**16 GB of local VRAM and 224 GB/s, at 70 W, for ~$350.** Bandwidth is the card's
weak point — decode is memory-bound, and ~224 GB/s is why an 8 B Q4 sits around the
mid-20s tok/s rather than the 60–70 tok/s of a 16 GB card with 512–624 GB/s. But
224 GB/s of *VRAM* still beats spilling experts over a PCIe link or DDR5, and 16 GB
is the size that makes "everything on the card" reachable at IQ2–Q4 for the model
sizes people actually self-host.

## What happens when I try it with Strata today

Environment (Linux, upstream `xe` driver bound, `intel-compute-runtime` 26.22,
`oneapi-level-zero`, `ocloc` present — no `icpx`/oneAPI DPC++ compiler installed):

- The SYCL/Intel backend **does** list `bmg-g21` as an AOT target, so the path exists.
- It is documented as **experimental**, and the docs/maintainer notes say it is
  untested on Arc hardware. On v0.1.39 I have no numbers, because I could not get
  to "running" without first solving the items below — i.e. the failure mode is
  silent-ish, which is exactly the thing to fix.
- **Toolchain friction:** `--backend sycl` needs oneAPI/`icpx`. My distro ships the
  *runtime* but not the compiler, so the Docker route is the only documented path
  and it is the one users fall over — and per #1113 the shipped port does not
  currently survive a clean build, so a first-time Arc user hits two walls in a row.
- **Device detection:** the device-selection logic does not recognise `8086:e212`,
  so the card is rejected rather than reported as "known, unsupported, want?"
  That converts a would-have-been-15-minutes test into a source read.

## Why I think it's worth your time (the interesting engineering part)

1. **It is a legitimate MoE test case, not a gaming card in disguise.** 16 GB local
   + one PCIe hop is exactly the topology where expert-streaming policies get
   decided. Any tuning you do for a 224 GB/s device with a finite 16 GB budget
   transfers to the APUs and the 12–16 GB RDNA cards your users already have.
2. **The upstream stack is already there.** Xe2 on the `xe` KMD, Level Zero, oneAPI —
   no vendor DKMS module to babysit. That is rare and it makes CI feasible.
3. **SR-IOV.** The Arc Pro line does hardware vGPU via standard SR-IOV with no
   per-seat licence, and the `xe` driver gained VF support around kernel 6.17. That
   means a single B50 can be shared across VMs/containers, which changes the
   economics of a shared Strata endpoint. (Caveat: the number of usable VFs has been
   firmware- and driver-version-dependent — worth documenting rather than asserting.)
4. **Power/frequency envelope as a first-class constraint.** A 70 W slot-powered
   card that stays in a box 24/7 is the deployment target most self-hosters are
   actually building toward. If Strata has an optimal configuration on a B50, that
   configuration is a selling point.

## Related open threads (so this is not another "any update?")

The Intel side of Strata is clearly being worked on — PR #1111 and the closed
#809 series show Arc work landing, and #1390/#1295 are pushing a native Windows
Arc path. But searching the tracker there is **no Arc Pro B50 thread**, and the
open items read as one systemic problem seen from five angles:

- **#1113** — the v0.1.40.1 SYCL port "does not build as shipped" (compile errors +
  undefined refs). Until that is resolved, no Arc card of any generation has
  reproducible numbers, which is exactly why the B50 row is empty rather than bad.
- **#1248** — SYCL build fails on Arc A770 (Alchemist), same family, same root path.
- **#1054 / #1302** — Arc Pro B70 `--layer-split` deadlock and B65 expert-cache fill
  stall: multi-card and expert-residency behaviour on Xe2 is still unresolved.
- **#692 / #573 / #299** — A770 / B580 / "will it work on Intel Arc" requests, with
  the last two closed without a support statement. That ambiguity is what costs
  users money and goodwill; a one-line matrix answer fixes it.
- **#868** documented the B60's PCI id (`e211`) in the Intel setup path. **Is `e212`
  (B50) expected to work through that same path?** If yes, the fix here may be
  documentation plus one detection entry, not new code — and I can verify that.

So this issue is deliberately narrow: not "please support Intel", but "the B50 has a
row in the matrix, detection knows the device id, and one honest install page exists".

## What I'm asking for (checked = I can do it / test it)

- [ ] **Device recognition, not rejection.** Allowlist `8086:e212` (BMG-G21) alongside
      `e211`, with an explicit `experimental on Arc, unsupported on consumer Arc`
      message instead of a generic "no supported device found".
- [ ] **A pointer to #1113.** If the current SYCL port cannot be built from a clean
      checkout, say so at the top of the Intel docs — that single line would have
      saved me the attempt and is the reason there is no B50 data yet.
- [ ] **A tested-matrix entry.** Even "B50 + SYCL + IQ2_XS: works, X tok/s, Y tok/s
      prefill, kernel ≥ 6.14, CR 26.22" is worth more than the current silence.
      I will run and maintain that row.
- [ ] **Pre-built `bmg-g21` AOT artifact** in the release, so users do not need
      `icpx` on the host to get value out of the SYCL backend.
- [ ] **Docs: the one honest SYCL install page** — kernel floor for Xe2, firmware
      files, compute-runtime version pairing, and a `sycl-ls`/Level Zero sanity
      check. There is a lot of stale guidance out there (including Intel-side
      tooling that has moved on), and community guides that send people in circles.
- [ ] **Prefill/decode backend note.** On this hardware in llama.cpp, SYCL and
      Vulkan have inverted strengths (SYCL much better token generation, Vulkan
      better prompt processing). If Strata ever exposes both paths on Xe2, that
      trade should be documented rather than discovered.
- [ ] **PCIe-width awareness in the streaming policy.** Explicitly model link width
      (x16/x8/x4, and OCulink/eGPU topologies that negotiate x4) when deciding what
      to keep resident vs. stream — a silent x4 negotiation is the difference between
      usable and not, and it is invisible to the engine today.
- [ ] Optional, if you want a real signal: **one CI job** that compiles the `bmg-g21`
      target and runs a smoke generate, so the experimental path cannot rot quietly.
      (Compilation coverage alone already catches most regressions and needs no GPU
      runner — for runtime you can pass a VF to a runner.)

## What I can offer

A physical Arc Pro B50 on an always-on Linux box (upstream `xe`, current compute
runtime), with: the model/pack set already downloaded and quantised at several
levels; `llama-bench`-style harnesses and a reproducible prefill/decode measurement
script; the ability to test `--backend sycl` builds (host `icpx` install or
containerised, whichever you prefer); and time to bisect. I'd rather send you
numbers and a working matrix row than another "any update?" comment.

If there is an existing internal Arc/Xe branch, or a reason `bmg-g21` is deliberately
frozen, say so and I'll close this — the goal is a clear statement of whether Intel
Xe2 is on the roadmap or not, for the several of us who picked this card partly
because the SYCL target already nominally existed.

## Reference data

- Intel ARK — Arc Pro B50: 16 Xe2 cores, 128 XMX, 2600 MHz max, 10.65 TFLOPS FP32,
  170 pTOPS INT8, 16 GB GDDR6 @ 14 Gbps / 128-bit / 224 GB/s, PCIe 5.0 x8, TBP 70 W,
  DP 2.1 UHBR 13.5 ×4, `8086:e212`
- Intel Arc Pro B50 data sheet / B-series press deck (SEP $299; B60 24 GB / 456 GB/s)
- Phoronix, *Intel Arc Pro B50 Linux Performance Benchmarks* (launch-day Linux
  numbers, ~1.47 × RTX A1000 aggregate, Level Zero / Compute Runtime working)
- Intel dgpu-docs — Xe KMD supported GPU table (B50 = `E212`, Xe2, kernel 6.11+ line)
- TechPowerUp GPU database entry for Arc Pro B50 (BMG-G21, 2600 MHz, 16 GB)

> Figures marked "estimated"/"reported" are community- or bandwidth-model-derived,
> not measured by me; I'll replace them with measured values from my box in a
> follow-up comment rather than put unverified numbers in the main table.

関連リンク

インストール・モデル・リリースへの站内リンク。