Pull requests / #784

#784 sycl: follow #626 (ThreadAffinity) and #559 (NativeDense::load layer range) so v0.1.39 compiles

closed · @sarashin65 · 0 comentários · No GitHub

NVIDIA / CUDADocumentationLinux

Descrição

## What this fixes

The `v0.1.39` tag does not compile on the Intel (SYCL) side. The port went in as 5047172 ("sycl: port 0.1.39"); two engine-side changes were merged into the release branch after it (`cb97e6c`, `1189218`) and nothing in `sycl/` was updated for them.

Three compile errors, 3 files, 11 insertions / 5 deletions:

| Engine change | What it broke in `sycl/` |
|---|---|
| #626 `ec35cf0` - the thread affinity became a value (`pin_current_thread()` returns `strata::kernels::cpu::ThreadAffinity`, `restore_thread_affinity()` takes one) | `sycl/include/strata/core/session.hpp` still stored a `long long` and did not include the header that names the type; `sycl/src/core/session.cpp` still assigned the returned struct to it and passed the `long long` back. `session.cpp:786` / `:801` |
| #559 `ba707d8` - the two layer-range arguments were added to `NativeDense::load` | the out-of-line definition in `sycl/src/core/native_dense.cpp:87` was still the 4-argument one, so it matched no declaration in `include/strata/core/native_dense.hpp:24` |

A plain `cmake --build` stops at the first one:

```
sycl/src/core/session.cpp:786:23: error: assigning to 'long long' from incompatible type 'ThreadAffinity'
sycl/src/core/session.cpp:801:55: error: reference to type 'const ThreadAffinity' could not bind to an lvalue of type 'long long'
sycl/src/core/native_dense.cpp:87:19: error: out-of-line definition of 'load' does not match any declaration in 'strata::core::NativeDense'
```

The third error is independent of the first two: it appears with `-- -k 0`, or once the first two are fixed.

Both files are corrected by mirroring the CUDA sources: the same member type, the same `outside()` predicate, and the same place it is applied in the tensor loop (`if (!eligible(tensor, include_ple_key) || outside(tensor.name)) continue;`).

Three notes for review:

- `pinned` now follows CUDA as well: it is `pinned_core.valid` (`src/core/session.cpp:550`), so a pin that failed is not "restored". Nothing changes when the pin succeeds.
- No caller passes `layer_lo`/`layer_hi` yet (`src/program/generate.cpp:2268,2596`; `sycl/src/program/generate.cpp:2366,2750`), so `outside()` is never true today. This restores the build; it does not change what the engine computes.
- `docs/INTEL.md` keeps the copies up to date by re-migration. This is the same change a re-migration would merge in, so a later re-migration should merge over it cleanly (no `sycl/tools/fixups.py` entry needed).

For anyone reading the port's own history: the "0 errors, 157 targets" in `docs/INTEL_ARC.md` was written on the branch before those two merges (`da3e9a1`), so it is not a contradiction.

## How it was checked

- Arch Linux, Intel oneAPI DPC++ 2026.0.0 + oneMKL 2026.0 (the port's docs used 2026.1.1).
- `cmake -S sycl -B build-sycl -G Ninja -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx`, SPIR-V (JIT).
- `cmake --build build-sycl --target strata -- -k 0` -> **0 errors, links**.

The two mismatches are a type and a declaration mismatch, so the compiler version does not matter.

## Not covered

- The other targets (kernel parity tests, benches) and the AOT path were not built; only `--target strata`.
- Not run on an Arc in this PR.
- Not part of this PR: declarations in 10 shared headers that the port does not define yet (`core/layer.hpp` `moe_route_window`, the `const QsaState* draft` overloads in `core/conversation_snapshot.hpp`, `kernels/bf16_gemv.hpp`, `kernels/elementwise.hpp`, `kernels/fused_gr.hpp`, `kernels/iq_kernels.hpp`, `kernels/ple.hpp`, `kernels/shared_expert.hpp`, `prefill/kernels.hpp`, `prefill/prefill.hpp`). Nothing in `sycl/` calls them, so it links. I can list them in an issue.

No site

Links install, modelos, releases.