Issues / #1552
#1552 Native Windows 11 Port for Intel Arc B580 (Strata v0.1.40.4)
open · @xlwang1188 · 0 commentaires · Sur GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
Description
# Native Windows 11 Port for Intel Arc B580 (Strata v0.1.40.4)
## Architecture, Source Adaptations, Build Instructions, and Benchmark Report
---
## 1. Executive Summary
Upstream **Strata v0.1.40.4** officially introduced support for Intel Arc graphics via SYCL, but explicitly noted:
> *"Intel Arc is Linux only for now."*
This document provides a comprehensive technical breakdown of our complete port and optimization of **Strata v0.1.40.4** running natively on **Microsoft Windows 11** with an **Intel Arc B580 (12 GB GDDR6, Xe2 Battlemage)** GPU.
### Key Highlights of This Port
1. **100% Native Windows Execution**: No Docker, WSL2, or Linux VM required. Direct Level-Zero (`ze_loader.dll`) hardware acceleration with Ahead-Of-Time (AOT) compilation targeting Intel `bmg-g21`.
2. **256K Context Window**: Configured with `--max-context 262144` and `--kv-resident 32768` (INT8 KV Cache), making it fully capable of running local agentic workflows.
3. **Multimodal Vision Integration**: Native CPU-based image encoder (`strata-vision.exe`) utilizing `qwen2vl` clip weights, with **0 bytes of VRAM overhead**, leaving maximum VRAM for resident LLM MoE experts.
4. **Native Hardware Telemetry**: Live PCIe bandwidth, GPU load, VRAM, board temperature, and CPU/Disk I/O monitoring using Intel PresentMon Service and Level-Zero APIs.
5. **Production Performance**:
- Standard dialogue decoding: **21.0 tok/s** (Expert Cache Hit Rate: **75.8%**, MTP Speculation Acceptance: **84.8%**).
- Extreme Long-Context Stress Test: **74,213 prompt tokens + 3,649 output tokens** sustained for **237 seconds** at **17.8 tok/s** (Zero memory leaks or degradations).
---
## 2. Test Environment & Hardware Specification
| Component | Specification |
| :--- | :--- |
| **Operating System** | Windows 11 Pro 64-bit (Build 24H2 / 26H2, WDDM 3.2) |
| **GPU** | Intel Arc B580 (Xe2 Battlemage, `bmg-g21`), 12 GB GDDR6 (12,706,643,968 bytes) |
| **Host CPU** | AMD Ryzen 5 5600 (6 Cores / 12 Threads, Zen 3, AVX2 enabled) |
| **System Memory** | 64 GB DDR4-3200 (Dual Channel, pinned USM mirror pool ~39 GB) |
| **Storage** | 1 TB NVMe PCIe 4.0 SSD (`D:\projects\strata\Strata`) |
| **GPU Driver** | Intel Arc Software & Drivers (Level-Zero / OpenCL 32.0.101.8991) |
| **Compiler / SDK** | Intel oneAPI Base Toolkit (DPC++ / `icx` 2026.1) + MSVC 19.4x (Visual Studio 2022) |
| **Model** | Qwen 3.8-Flash-Next (125B MoE, IQ2_XS quantization, 24,614 total experts, 48 layers) |
---
## 3. Upstream Code Differences & Technical Adaptations
To bring upstream Linux-centric SYCL code to Windows 11 MSVC, several major obstacles across header collisions, memory subsystem design, and telemetry were resolved.
### 3.1 Windows SDK Macro & Header Collisions
When compiling SYCL kernels on Windows using the Intel DPC++ driver (`icx` in MSVC-compatible mode), Windows SDK headers (`windows.h`, `minwindef.h`, `rpcndr.h`) inject global macros that break C++ template metaprogramming and standard variables.
* **Macro `OUT` and `IN` Collisions**:
- *Location*: [`sycl/src/kernels/cuda/verify_kernels.dp.cpp`](file:///D:/projects/strata/Strata/sycl/src/kernels/cuda/verify_kernels.dp.cpp)
- *Problem*: Windows `<minwindef.h>` defines `#define OUT` and `#define IN`. In upstream Strata, speculative verification templates define `template <bool OUT>`. Under MSVC, this expands to `template <bool >`, triggering compile-time syntax errors.
- *Fix*: Injected explicit macro guards before kernel template declarations:
```cpp
#ifdef OUT
#undef OUT
#endif
#ifdef IN
#undef IN
#endif
```
* **Macro `small` Collision**:
- *Location*: [`sycl/src/program/generate.cpp`](file:///D:/projects/strata/Strata/sycl/src/program/generate.cpp)
- *Problem*: Windows RPC header `<rpcndr.h>` defines `#define small char`. In upstream `generate.cpp`, `const int64_t small = old_rule()` was interpreted as `const int64_t char = old_rule()`, breaking compilation.
- *Fix*: Added `#undef small` guard after system header imports.
* **Missing GNU Extension `memmem`**:
- *Location*: [`sycl/src/program/generate.cpp`](file:///D:/projects/strata/Strata/sycl/src/program/generate.cpp)
- *Problem*: Linux glibc provides `memmem()` for binary substring searches, but the Microsoft Visual C++ CRT does not.
- *Fix*: Implemented an inline Windows fallback:
```cpp
#if defined(_WIN32)
static const void* memmem_win(const void* haystack, size_t haystacklen,
const void* needle, size_t needlelen) {
if (needlelen == 0) return haystack;
if (haystacklen < needlelen) return nullptr;
const char* h = static_cast<const char*>(haystack);
const char* n = static_cast<const char*>(needle);
for (size_t i = 0; i <= haystacklen - needlelen; ++i) {
if (h[i] == n[0] && std::memcmp(h + i, n, needlelen) == 0) {
return h + i;
}
}
return nullptr;
}
#define memmem memmem_win
#endif
```
---
### 3.2 Host-to-Device Memory Mirroring & Chunking
* **4 GiB Single Allocation Limitation on Windows WDDM**:
- *Problem*: Although the Intel Arc B580 has 12 GB of VRAM and the system has 64 GB of RAM, attempting to allocate the entire ~39 GB cold-expert mirror as a single contiguous USM host buffer (`sycl::malloc_host`) causes the Windows Level-Zero runtime to return `nullptr`.
- *Fix*: Refactored the host memory mirror in [`sycl/src/core/gguf_expert_source.cpp`](file:///D:/projects/strata/Strata/sycl/src/core/gguf_expert_source.cpp) into chunked allocations capped at 4 GiB (`kChunkMax = 4ull << 30`). Maintained a two-tier pointer index table `mirror_ptr_[(layer * n_expert + expert)]` that seamlessly resolves physical expert addresses across chunk boundaries.
* **Direct File I/O Mapping**:
- Replaced Linux-specific `pread()` and POSIX `open()` calls with 64-bit Windows file offsets using `_lseeki64()` and `_read()`.
---
### 3.3 256K Context Architecture & Expert Slots Allocation
* **Expert Slots Expansion**:
- Upstream v0.1.40.4 optimized SYCL memory pool alignment. Combined with our DXGI dedicated VRAM query (`IDXGIAdapter3::QueryVideoMemoryInfo`), the engine reliably reserves **2,808 resident expert slots** in GPU VRAM (3.78 GiB) on the 12 GB card, up from 2,741 in previous iterations.
* **256K Context Configuration**:
- Configured in [`strata-iq2_xs.json`](file:///D:/projects/strata/Strata/strata-iq2_xs.json):
- `--max-context 262144`
- `--kv-resident 32768` (high-speed VRAM resident window)
- INT8 quantized KV cache (`"kv": "int8"`)
- Dynamic host streaming for the remainder of the 256K token context.
---
### 3.4 Zero-VRAM Multimodal Vision Integration
* **Challenge**: Standard vision LLM architectures allocate 2 to 4 GB of GPU VRAM for the vision transformer encoder (ViT/CLIP), which would catastrophically starve the 12 GB GPU, shrinking resident MoE experts and destroying decode throughput.
* **Solution**:
- Built an independent native CLI encoder: [`build-sycl-aot/strata-vision.exe`](file:///D:/projects/strata/Strata/build-sycl-aot/strata-vision.exe) compiled from [`tools/vision/strata_vision.cpp`](file:///D:/projects/strata/Strata/tools/vision/strata_vision.cpp).
- Uses AVX2 multi-threaded CPU execution to encode input images and project them into textual token embeddings.
- Runs in **~8.0 seconds** per image on the 6-core Ryzen 5 5600 with **0 bytes of GPU VRAM consumed**.
- Enabled `"images": true` in `serve/server.py` and connected directly to the OpenAI `/v1/chat/completions` multimodal schema.
---
### 3.5 Real-Time Telemetry Pipeline for Windows
* Upstream Strata supported NVML (NVIDIA) and Linux sysfs (AMD). For Intel GPUs on Windows, it showed `not readable`.
* Implemented native Windows dual-engine telemetry in [`serve/telemetry.py`](file:///D:/projects/strata/Strata/serve/telemetry.py):
1. **Intel Level-Zero Loader (`ze_loader.dll`)**: Queries PCIe properties (`PciSpeed.gen`, `PciSpeed.width`) and engine activity fallbacks.
2. **Intel PresentMon Shared Service (`PresentMonAPI2.dll`)**: Dynamic query sampling for GPU temperature, package power, real-time VRAM allocation, and utilization.
3. **Frontend Dashboard Fix**: Resolved a variable scoping omission in [`serve/web/app.js`](file:///D:/projects/strata/Strata/serve/web/app.js) (`const gen = hw.gpu_pcie_gen_max || hw.gpu_pcie_gen;`), ensuring the real-time PCIe bandwidth, CPU load, and disk read/write sparklines update without JavaScript exceptions.
---
## 4. Step-by-Step Compilation & Build Guide
### Prerequisites
1. **Windows 11 (64-bit)**
2. **Microsoft Visual Studio 2022** (Desktop development with C++, MSVC v143, Windows 11 SDK)
3. **Intel oneAPI Base Toolkit 2026.1**:
- Intel oneAPI DPC++/C++ Compiler (`icx`, `icpx`)
- Intel oneAPI Math Kernel Library (oneMKL)
- Intel Graphics Compute Runtime (`ocloc`, `ze_loader.dll`)
4. **Ninja Build Tool & CMake 3.26+**
### Step 1: Environment Initialization
Launch an x64 Native Tools Command Prompt or set the oneAPI environment variables:
```cmd
call "C:\Program Files (x86)\Intel\oneAPI\setvars.bat"
```
Ensure runtime DLLs (`libhwloc-15.dll`, `ur_adapter_level_zero.dll`) from `C:\Program Files (x86)\Intel\oneAPI\2026.1\bin` are accessible in the system `PATH`.
### Step 2: Configure & AOT Compile Strata Engine
Run [`sycl/build_windows.bat`](file:///D:/projects/strata/Strata/sycl/build_windows.bat) or invoke CMake directly:
```cmd
cmake -S sycl -B build-sycl-aot -G Ninja ^
-DCMAKE_BUILD_TYPE=Release ^
-DCMAKE_C_COMPILER=icx ^
-DCMAKE_CXX_COMPILER=icx ^
-DSTRATA_SYCL_AOT=bmg-g21 ^
-DSTRATA_SYCL_PARITY=OFF
cmake --build build-sycl-aot --target strata --config Release
```
*Key Compiler Flags Applied*:
* `/EHsc /O2 /MD /utf-8`: Standard MSVC exception handling and runtime linking.
* `-fsycl -fsycl-targets=spir64_gen -Xsycl-target-backend "-device bmg-g21"`: Targets Intel Xe2 Battlemage explicitly, generating native binary kernels to eliminate runtime JIT pauses.
### Step 3: Compile Vision Encoder (`strata-vision.exe`)
```cmd
cmake -S tools/vision -B build-vision -G Ninja ^
-DCMAKE_BUILD_TYPE=Release ^
-DCMAKE_CXX_COMPILER=cl
cmake --build build-vision --config Release
copy build-vision\strata_vision.exe build-sycl-aot\strata-vision.exe
```
### Step 4: Verification of Build Artifacts
Ensure the following binaries exist in `build-sycl-aot\`:
* `strata.exe` (~73.6 MB)
* `strata-vision.exe` (~12.4 MB)
---
## 5. Preliminary Benchmark & Performance Results
### 5.1 Test Methodology
Testing was conducted on the live server using `qwen3.8-flash-next-iq2_xs` with speculative decoding enabled (`--spec 6 --mtp-max 4 --lookup 3`). Metrics were gathered directly from `/metrics` and the hardware telemetry pipeline.
### 5.2 Benchmark Results Table
| Workload Scenario | Context (Prompt / Output) | Decode Speed | Cache Hit Rate | MTP Acceptance | Observations |
| :--- | :--- | :--- | :--- | :--- | :--- |
| **Short Dialogue** | 81 prompt / 100 out | **21.0 tok/s** | **75.8%** | **84.8% (67/79)** | Snappy interactive chat; instant time-to-first-token. |
| **Medium Context** | 59,272 prompt / 98 out | **20.4 tok/s** | **74.0%** | **85.0% (68/80)** | KV cache remains responsive over PCIe streaming. |
| **Heavy Stress Test** | **74,213 prompt / 3,649 out** | **17.8 tok/s** | **66.2%** | **68.6% (1892/2757)** | **237.7 seconds continuous generation**. Zero memory leaks, steady 60°C GPU temp. |
| **Multimodal Vision** | 1 synthetic test image | ~8.01s total | N/A | N/A | CPU inference; **0 MB VRAM consumed**; exact shape/color recognition. |
### 5.3 Comparative Analysis
```mermaid
xychart-beta
title "Decode Speed Comparison (tokens/second)"
x-axis ["llama.cpp / Ollama (Offload)", "Strata v0.1.3x (Initial Win11)", "Strata v0.1.40.4 (Optimized Win11)"]
y-axis "Tokens / Second" 0 --> 25
bar [2.5, 17.1, 21.0]
```
1. **Versus llama.cpp / Ollama**:
- Sur le site
Liens install, modèles, releases.