Pull requests / #1390
#1390 Windows: native Intel Arc build and setup (OpenCL), plus three Windows-SDK macro fixes
open · @demetree · 0 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Beschreibung
### What Native Windows support for the SYCL port (`setup --backend sycl`): no Docker, the engine is built from source with Intel oneAPI (`sycl\tools\build.bat`) and run directly (`sycl\serve\strata-sycl.bat`). Plus the fixes the port needs to compile on Windows at all, and the no-command-graph paths — this Windows driver exposes no user-mode Level Zero adapter, so the OpenCL adapter has no Graph extension. Replaces #1295: that branch predates 0.1.40.2, and the release merged most of its forward-port (and has its own `STRATA_VERIFY_EAGER` window replay), so this one carries the Windows deltas only — 4 commits. ## Measured Arc Pro B70 32 GB, Windows 10, driver 32.0.101.8976, i9-9900 / 64 GB, conda-forge `dpcpp_win-64` 2026.1.1, OpenCL backend, `STRATA_VERIFY_EAGER=1`. Coder IQ1_M, 12,288/12,288 experts resident, `--stream-experts`, INT8 KV, greedy. Each row is this branch's engine (0.1.40-sycl) against the previous 0.1.39-based build of this port, run **interleaved** from the command line (v1, v2, v1, v2, …), which is how the numbers in docs/INTEL.md were taken: | | 0.1.39-sycl | 0.1.40-sycl (this PR) | Linux / Level Zero / graphs | |---|---|---|---| | decode, 256 tokens, `--spec 4 --mtp`, 3 pairs | 69.6-69.9 (median 69.7) | **77.2-77.3 (median 77.3)** | 78.2 | | decode, 128 tokens after a 2,048-token prompt, 2 pairs | 37.6 | **38.2-41.1 (median 39.6)** | — | | decode, 256 tokens, `--spec 2` (no draft layer), 3 pairs | 53.4-53.9 (median 53.6) | **22.0-22.9 (median 22.3)** | 33.6 | | prompt reading, 2,048 tokens | 8.0-14.5 tok/s | **71-73 tok/s** | 790-1,002 tok/s | | `quantize_act_parity --selftest`, `sampler_parity` | byte-exact, 0 failures | byte-exact, 0 failures | — | **Two traps in measuring this backend, both worth passing on.** With a 2,048-token prompt the engine is not run-to-run deterministic: the same binary diverges at token 60 and its state hash differs between two runs. And the web app answers a `temperature: 0.0` request with its own Chat defaults (temperature 1.0, top_p 0.95), so a "greedy" run through it is sampled. The kernel-level self-tests and the CLI's `--greedy` are the checks that hold. **And short runs are not comparable on this backend.** The same build and mode gave 48.9 and 67.5 tok/s on two consecutive 128-token runs, and a 155-token answer through the API contradicted a 256-token run by 25%. My first pass at this table used those short runs and wrongly read as a ~25% regression; interleaved medians over 256 tokens are the numbers above. **The one reproducible regression is `--spec 2`** — 22.3 against 53.6 tok/s, interleaved, same card and flags. It is the no-draft-layer path, so it is either the drafter-free window or the suffix drafter's cost; not attributed. `STRATA_VERIFY_PROFILE=1` puts 24 of a 65 ms 4-token window in the GDN hyper-connection read, nearly flat from T=2 (23.6 ms) to T=6 (25.6 ms): a per-layer *fixed* cost of ~0.5 ms over 48 layers — kernel dispatch, not arithmetic. `sycl-ls` lists only `opencl:gpu` here, so there is neither the Graph extension nor sysman's free-VRAM query. ## Two things a fresh Windows build needs, both found by building it clean - **The DLLs to start at all.** `strata.exe` statically imports `libmmd.dll` (Intel's runtime, beside `icx`), and the UR adapters load `ur_loader.dll` / `ur_win_proxy_loader.dll` from beside the exe. Without those three staged the engine exits `0xC0000135` (STATUS_DLL_NOT_FOUND) before printing anything — it only ran in my earlier build directory by accident. `sycl/CMakeLists.txt` now stages all three with the rest. - **`sycl/tools/build.bat` finds `libircmt.lib` through the build dir's CMakeCache** when the shell has no `icx` on PATH, so the script works from a plain prompt; otherwise the link fails `LNK1104`. ## Three compile bugs in the port's own code `<sycl/sycl.hpp>` reaches the Windows SDK, whose headers define `OUT` (minwindef.h), `small` (rpcndr.h) and `near` (windef.h) as macros. An identifier with one of those names is macro-expanded away: - `sycl/src/kernels/cuda/verify_kernels.dp.cpp`: `template <bool OUT>` loses its parameter, so every use becomes `error: expected expression`. Renamed `WRITE_OUT`. - `sycl/src/program/generate.cpp`: `const int64_t small` becomes `const int64_t char`; `memmem` is POSIX (now `std::search` over the same bytes). - `sycl/src/kernels/sampler_parity.cpp`: `const bool near`. ## Without command graphs `STRATA_VERIFY_EAGER=1` (set by `sycl/serve\strata-sycl.bat`) replays the window, the commit and the drafter's chains on the queue instead of launching captured graphs: - `record_commit()` is the commit graph's body; `capture_*()` record it where there is a graph backend. - The `capture_round` / `capture_step` / `capture_prefill` bodies became `record_*()` functions the drafter replays. `prepare_chain()` refuses: its per-step events only exist inside a recorded graph. - Eager replay **refuses a host-served expert plan** instead of running it — the mid-record waits expire into stale plans, which is silently wrong rather than slow. - oneMKL SYCL BLAS has no OpenCL Xe2 backend and the failed attempt poisons the queue, so the prompt GEMM is a plain SYCL kernel there - and the first version of that kernel read the whole weight matrix once **per row** (a 2,048-row chunk of one 2,560 x 4,096 projection moved 86 GB of weights). `fallback_gemm_tiled` blocks a 64 x 64 corner of Y per work-group and walks K in 16-wide tiles with the operands in local memory: **8.0-14.5 -> 71-73 tok/s** on a 2,048-token prompt, decode unchanged. Same arithmetic - both kernels accumulate over k in ascending order with the fma spelled out - and `STRATA_FALLBACK_SELFTEST=1` runs both over a problem that crosses every tile edge and prints the bitwise difference: **0 of 4,189 values, largest 0**. - `sampler_parity`'s fixture 18 skips its graph half without the extension instead of ending the process. ## Also - Setup names the Intel Arc on Windows (display adapters + the display-class registry's 64-bit VRAM, which is the true VRAM — WMI's `AdapterRAM` stops at 4 GB), and `--backend sycl` hands over to `sycl/setup_intel.py` there too. - `sycl/CMakeLists.txt`: IntelLLVM / MSVC flag branches (`/fp:precise` is load-bearing — `quantize_act_parity` goes from byte-exact to 303k mismatches without it), installer-free oneMKL, and the runtime DLLs staged beside `strata.exe`. - `GgufExpertSource` reads at an offset through `OVERLAPPED` (there is no `pread` on Windows). - The free-VRAM query needs Level Zero sysman, so setup writes `STRATA_DEVICE_FREE_MIB` / `_TOTAL_MIB` from the registry, and keeps them across a rewrite when detection flakes. ## The spin bound on Windows, and what it decides 0.1.40.3 made the device's spin bound on host flags a build option (`STRATA_SYCL_SPIN_MAX`) and gave 20,000 only to a `bmg` **AOT** build, 2,000,000 to everything else. A JIT build of a B70 is not an AOT build, so it got the long bound, and on this backend that costs a lot: same build, one `-DSTRATA_SYCL_SPIN_MAX` apart, Coder IQ1_M with `--ple-gguf`, 256 greedy tokens after a 2,048-token prompt, medians of two interleaved runs. | mode | | 2,000,000 | 20,000 | |---|---|---|---| | `--spec 4 --mtp` (what setup writes) | tok/s | 26.2 | **45.4** | | | tokens per round | 2.28 (73 of 126 drafts accepted) | **4.19 (99 of 99)** | | `--spec 2` | tok/s | **68.8** | 21.8 | | | tokens per round | **2.17** (2 suffix windows drafted) | 1.03 (**0 drafted**) | | `--spec 4`, no `--mtp` | tok/s | **69.7** | 18.9 | What the bound decides is **whether speculation works at all**, not how long a window costs. A window that spins on a host flag it cannot see runs out and goes on without what it was waiting for: harmless when what it wanted was a plan it already has (the window's waits, with the MTP record path having published it), ruinous when it was a draft. So with `--mtp` the short bound is both faster and better, and without it the suffix drafter's own waits run out inside 20,000, it produces no draft at all, every round verifies a single token and decode is 3.7x slower - silently, since the answer is still right. Hence: 20,000 is the default on Windows, because the port's own configuration is `--spec 4 --mtp`; `sycl/serve/strata-sycl.bat` warns on stderr when the engine is started without `--mtp` and names the override; and `-DSTRATA_SYCL_SPIN_MAX=...` still wins. Upstream report with the numbers: **#1473** - including the correction that keying the carve-out on the card would not fix it, since one constant cannot serve both kinds of wait. The `--spec 2` regression in **#1397** is a different thing (both those builds carry the 20,000 bound, and the SYCL sources differ by 4 lines); its evidence, the tokens-per-round signature and the one-line test I would want run are in the comment there. ## Limits - Needs an admin account for VS Build Tools and the Intel driver; the conda + VS + PyPI/NuGet oneMKL pieces install per-user. - A model whose experts do not all fit needs a RAM mirror (`STRATA_MIRROR_MIB`) or all-resident sizing; IQ3_XXS (17,687/24,576 experts resident here) stalls without one. - Windows-SDK macro names have to be avoided as identifiers in the port. Tests: `tools/test_setup_intel_windows.py` (new), `tools/test_setup_sycl.py`, `tools/test_setup_amd.py`, `tools/test_setup_choices.py` — 87 pass, no GPU and no network.
Mehr auf der Site
Links zu Install, Modellen, Releases.