Pull requests / #1390

#1390 Windows: native Intel Arc build and setup (OpenCL), plus three Windows-SDK macro fixes

open · @demetree · 0 comentários · No GitHub

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Descrição

### What



Native Windows support for the SYCL port (`setup --backend sycl`): no Docker, the engine is built from source with

Intel oneAPI (`sycl\tools\build.bat`) and run directly (`sycl\serve\strata-sycl.bat`). Plus the fixes the port needs

to compile on Windows at all, and the no-command-graph paths ΓÇö this Windows driver exposes no user-mode Level Zero

adapter, so the OpenCL adapter has no Graph extension.



Replaces #1295: that branch predates 0.1.40.2, and the release merged most of its forward-port (and has its own

`STRATA_VERIFY_EAGER` window replay), so this one carries the Windows deltas only ΓÇö 4 commits.



## Measured



Arc Pro B70 32 GB, Windows 10, driver 32.0.101.8976, i9-9900 / 64 GB, conda-forge `dpcpp_win-64` 2026.1.1, OpenCL

backend, `STRATA_VERIFY_EAGER=1`. Coder IQ1_M, 12,288/12,288 experts resident, `--stream-experts`, INT8 KV, greedy.

Each row is this branch's engine (0.1.40-sycl) against the previous 0.1.39-based build of this port, run

**interleaved** from the command line (v1, v2, v1, v2, …), which is how the numbers in docs/INTEL.md were taken:



| | 0.1.39-sycl | 0.1.40-sycl (this PR) | Linux / Level Zero / graphs |

|---|---|---|---|

| decode, 256 tokens, `--spec 4 --mtp`, 3 pairs | 69.6-69.9 (median 69.7) | **77.2-77.3 (median 77.3)** | 78.2 |

| decode, 128 tokens after a 2,048-token prompt, 2 pairs | 37.6 | **38.2-41.1 (median 39.6)** | ΓÇö |

| decode, 256 tokens, `--spec 2` (no draft layer), 3 pairs | 53.4-53.9 (median 53.6) | **22.0-22.9 (median 22.3)** | 33.6 |

| prompt reading, 2,048 tokens | 8.0-14.5 tok/s | **71-73 tok/s** | 790-1,002 tok/s |

| `quantize_act_parity --selftest`, `sampler_parity` | byte-exact, 0 failures | byte-exact, 0 failures | ΓÇö |



**Two traps in measuring this backend, both worth passing on.** With a 2,048-token prompt the engine is not

run-to-run deterministic: the same binary diverges at token 60 and its state hash differs between two runs. And the

web app answers a `temperature: 0.0` request with its own Chat defaults (temperature 1.0, top_p 0.95), so a "greedy"

run through it is sampled. The kernel-level self-tests and the CLI's `--greedy` are the checks that hold.



**And short runs are not comparable on this backend.** The same build and mode gave 48.9 and

67.5 tok/s on two consecutive 128-token runs, and a 155-token answer through the API contradicted a 256-token run by

25%. My first pass at this table used those short runs and wrongly read as a ~25% regression; interleaved medians over

256 tokens are the numbers above.



**The one reproducible regression is `--spec 2`** ΓÇö 22.3 against 53.6 tok/s, interleaved, same card and flags. It is

the no-draft-layer path, so it is either the drafter-free window or the suffix drafter's cost; not attributed.



`STRATA_VERIFY_PROFILE=1` puts 24 of a 65 ms 4-token window in the GDN hyper-connection read, nearly flat from T=2

(23.6 ms) to T=6 (25.6 ms): a per-layer *fixed* cost of ~0.5 ms over 48 layers ΓÇö kernel dispatch, not arithmetic.

`sycl-ls` lists only `opencl:gpu` here, so there is neither the Graph extension nor sysman's free-VRAM query.



## Two things a fresh Windows build needs, both found by building it clean



- **The DLLs to start at all.** `strata.exe` statically imports `libmmd.dll` (Intel's runtime, beside `icx`), and the

  UR adapters load `ur_loader.dll` / `ur_win_proxy_loader.dll` from beside the exe. Without those three staged the

  engine exits `0xC0000135` (STATUS_DLL_NOT_FOUND) before printing anything ΓÇö it only ran in my earlier build

  directory by accident. `sycl/CMakeLists.txt` now stages all three with the rest.

- **`sycl/tools/build.bat` finds `libircmt.lib` through the build dir's CMakeCache** when the shell has no `icx` on

  PATH, so the script works from a plain prompt; otherwise the link fails `LNK1104`.



## Three compile bugs in the port's own code



`<sycl/sycl.hpp>` reaches the Windows SDK, whose headers define `OUT` (minwindef.h), `small` (rpcndr.h) and `near`

(windef.h) as macros. An identifier with one of those names is macro-expanded away:



- `sycl/src/kernels/cuda/verify_kernels.dp.cpp`: `template <bool OUT>` loses its parameter, so every use becomes

  `error: expected expression`. Renamed `WRITE_OUT`.

- `sycl/src/program/generate.cpp`: `const int64_t small` becomes `const int64_t char`; `memmem` is POSIX (now

  `std::search` over the same bytes).

- `sycl/src/kernels/sampler_parity.cpp`: `const bool near`.



## Without command graphs



`STRATA_VERIFY_EAGER=1` (set by `sycl/serve\strata-sycl.bat`) replays the window, the commit and the drafter's

chains on the queue instead of launching captured graphs:



- `record_commit()` is the commit graph's body; `capture_*()` record it where there is a graph backend.

- The `capture_round` / `capture_step` / `capture_prefill` bodies became `record_*()` functions the drafter

  replays. `prepare_chain()` refuses: its per-step events only exist inside a recorded graph.

- Eager replay **refuses a host-served expert plan** instead of running it ΓÇö the mid-record waits expire into stale

  plans, which is silently wrong rather than slow.

- oneMKL SYCL BLAS has no OpenCL Xe2 backend and the failed attempt poisons the queue, so the prompt GEMM is a

  plain SYCL kernel there - and the first version of that kernel read the whole weight matrix once **per row** (a

  2,048-row chunk of one 2,560 x 4,096 projection moved 86 GB of weights). `fallback_gemm_tiled` blocks a 64 x 64

  corner of Y per work-group and walks K in 16-wide tiles with the operands in local memory: **8.0-14.5 -> 71-73

  tok/s** on a 2,048-token prompt, decode unchanged. Same arithmetic - both kernels accumulate over k in ascending

  order with the fma spelled out - and `STRATA_FALLBACK_SELFTEST=1` runs both over a problem that crosses every tile

  edge and prints the bitwise difference: **0 of 4,189 values, largest 0**.

- `sampler_parity`'s fixture 18 skips its graph half without the extension instead of ending the process.



## Also



- Setup names the Intel Arc on Windows (display adapters + the display-class registry's 64-bit VRAM, which is the

  true VRAM ΓÇö WMI's `AdapterRAM` stops at 4 GB), and `--backend sycl` hands over to `sycl/setup_intel.py` there too.

- `sycl/CMakeLists.txt`: IntelLLVM / MSVC flag branches (`/fp:precise` is load-bearing ΓÇö `quantize_act_parity` goes

  from byte-exact to 303k mismatches without it), installer-free oneMKL, and the runtime DLLs staged beside

  `strata.exe`.

- `GgufExpertSource` reads at an offset through `OVERLAPPED` (there is no `pread` on Windows).

- The free-VRAM query needs Level Zero sysman, so setup writes `STRATA_DEVICE_FREE_MIB` / `_TOTAL_MIB` from the

  registry, and keeps them across a rewrite when detection flakes.



## The spin bound on Windows, and what it decides

0.1.40.3 made the device's spin bound on host flags a build option (`STRATA_SYCL_SPIN_MAX`) and gave 20,000 only to a
`bmg` **AOT** build, 2,000,000 to everything else. A JIT build of a B70 is not an AOT build, so it got the long bound,
and on this backend that costs a lot: same build, one `-DSTRATA_SYCL_SPIN_MAX` apart, Coder IQ1_M with `--ple-gguf`,
256 greedy tokens after a 2,048-token prompt, medians of two interleaved runs.

| mode | | 2,000,000 | 20,000 |
|---|---|---|---|
| `--spec 4 --mtp` (what setup writes) | tok/s | 26.2 | **45.4** |
| | tokens per round | 2.28 (73 of 126 drafts accepted) | **4.19 (99 of 99)** |
| `--spec 2` | tok/s | **68.8** | 21.8 |
| | tokens per round | **2.17** (2 suffix windows drafted) | 1.03 (**0 drafted**) |
| `--spec 4`, no `--mtp` | tok/s | **69.7** | 18.9 |

What the bound decides is **whether speculation works at all**, not how long a window costs. A window that spins on a
host flag it cannot see runs out and goes on without what it was waiting for: harmless when what it wanted was a plan it
already has (the window's waits, with the MTP record path having published it), ruinous when it was a draft. So with
`--mtp` the short bound is both faster and better, and without it the suffix drafter's own waits run out inside 20,000,
it produces no draft at all, every round verifies a single token and decode is 3.7x slower - silently, since the answer
is still right.

Hence: 20,000 is the default on Windows, because the port's own configuration is `--spec 4 --mtp`;
`sycl/serve/strata-sycl.bat` warns on stderr when the engine is started without `--mtp` and names the override; and
`-DSTRATA_SYCL_SPIN_MAX=...` still wins. Upstream report with the numbers: **#1473** - including the correction that
keying the carve-out on the card would not fix it, since one constant cannot serve both kinds of wait.

The `--spec 2` regression in **#1397** is a different thing (both those builds carry the 20,000 bound, and the SYCL
sources differ by 4 lines); its evidence, the tokens-per-round signature and the one-line test I would want run are in
the comment there.

## Limits



- Needs an admin account for VS Build Tools and the Intel driver; the conda + VS + PyPI/NuGet oneMKL pieces install

  per-user.

- A model whose experts do not all fit needs a RAM mirror (`STRATA_MIRROR_MIB`) or all-resident sizing; IQ3_XXS

  (17,687/24,576 experts resident here) stalls without one.

- Windows-SDK macro names have to be avoided as identifiers in the port.



Tests: `tools/test_setup_intel_windows.py` (new), `tools/test_setup_sycl.py`, `tools/test_setup_amd.py`,

`tools/test_setup_choices.py` ΓÇö 87 pass, no GPU and no network.

No site

Links install, modelos, releases.