Pull requests / #667

#667 Add Intel XPU backend: Qwen decode on Arc Pro B60 via SYCL

closed · @the-talos · 0 评论 · 在 GitHub 查看

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

描述

## What this adds

A third backend next to CUDA and HIP, enabled with `STRATA_ENABLE_XPU=ON` (icpx/icx, oneAPI 2025.3). The CUDA kernels are reused unchanged: `tools/xpu/rewrite_cuda.py` rewrites launch syntax and declarations to SYCL at build time, and `include/strata/xpu_compat/` implements the CUDA runtime API the engine leans on (graphs, streams, events, mapped memory, cublas via oneMKL) on SYCL/Level Zero. `docs/INTEL_XPU.md` documents the design and the build.

## Numbers

Qwen3.8-Flash-Next GSQ-RCO Q2_0 on one Intel Arc Pro B60 (23.3 GB), expert cache auto (14,570 slots, 18.8 GiB), context 1024, `--spec 2 --kv int8`:

- prefill 127 tokens at 73.9 tok/s, time to first token 2.1 s
- greedy decode through the speculative verify window at 9.4 tok/s (64 tokens), exit 0
- in-run gates pass: cache slot verified bit for bit, every layer doorbell rings

## Porting notes worth reviewing

- Graph execs deep-copy the captured op list; the engine destroys the capture right after instantiate, as CUDA allows.
- `cudaStreamQuery` is an honest non-blocking idle check on the in-order queue; the verify loop polls it while spinning on doorbells.
- Async copies from unpinned host pages (the mmap-ed GGUF shards) stage through a `malloc_host` bounce with an in-order `host_task` fill; Level Zero faults on raw pages and strands every later op on the stream.
- The doorbell spin kernels poll host-mapped flags. Pure volatile read loops over host USM do not see CPU stores on this stack, so `__nanosleep` (called by every `strata_spin_pause()` site each iteration) issues a `sycl::atomic_fence(seq_cst, system)`. Without it the verify window times out at the first flag wait, every run.
- Performance is scalar so far: `__dp4a` and shuffles are correct subgroup implementations, XMX paths are future work.

I have run the greedy smoke test and the measurement above on the B60 container only; the CUDA and HIP builds are untouched (the new code compiles in only under `STRATA_ENABLE_XPU`).

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。