Pull requests / #667
#667 Add Intel XPU backend: Qwen decode on Arc Pro B60 via SYCL
closed · @the-talos · 0 comentarios · En GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Descripción
## What this adds A third backend next to CUDA and HIP, enabled with `STRATA_ENABLE_XPU=ON` (icpx/icx, oneAPI 2025.3). The CUDA kernels are reused unchanged: `tools/xpu/rewrite_cuda.py` rewrites launch syntax and declarations to SYCL at build time, and `include/strata/xpu_compat/` implements the CUDA runtime API the engine leans on (graphs, streams, events, mapped memory, cublas via oneMKL) on SYCL/Level Zero. `docs/INTEL_XPU.md` documents the design and the build. ## Numbers Qwen3.8-Flash-Next GSQ-RCO Q2_0 on one Intel Arc Pro B60 (23.3 GB), expert cache auto (14,570 slots, 18.8 GiB), context 1024, `--spec 2 --kv int8`: - prefill 127 tokens at 73.9 tok/s, time to first token 2.1 s - greedy decode through the speculative verify window at 9.4 tok/s (64 tokens), exit 0 - in-run gates pass: cache slot verified bit for bit, every layer doorbell rings ## Porting notes worth reviewing - Graph execs deep-copy the captured op list; the engine destroys the capture right after instantiate, as CUDA allows. - `cudaStreamQuery` is an honest non-blocking idle check on the in-order queue; the verify loop polls it while spinning on doorbells. - Async copies from unpinned host pages (the mmap-ed GGUF shards) stage through a `malloc_host` bounce with an in-order `host_task` fill; Level Zero faults on raw pages and strands every later op on the stream. - The doorbell spin kernels poll host-mapped flags. Pure volatile read loops over host USM do not see CPU stores on this stack, so `__nanosleep` (called by every `strata_spin_pause()` site each iteration) issues a `sycl::atomic_fence(seq_cst, system)`. Without it the verify window times out at the first flag wait, every run. - Performance is scalar so far: `__dp4a` and shuffles are correct subgroup implementations, XMX paths are future work. I have run the greedy smoke test and the measurement above on the B60 container only; the CUDA and HIP builds are untouched (the new code compiles in only under `STRATA_ENABLE_XPU`).
En el sitio
Enlaces a install, modelos, releases.