Issues / #1397

#1397 SYCL port on Windows/OpenCL: 0.1.40-sycl decodes at 22.3 tok/s with --spec 2 where 0.1.39-sycl does 53.6 (and short runs vary 40%)

open · @demetree · 2 Kommentare · Auf GitHub

BenchmarksSetup & installModels & quantsDocumentationWindows

Beschreibung

### Summary

On an Arc Pro B70 (32 GB) with the SYCL port on **Windows 10 / OpenCL** - a backend without command graphs, since the
Windows driver exposes no user-mode Level Zero adapter - engine 0.1.40-sycl decodes **22.3 tok/s** with `--spec 2`
where 0.1.39-sycl decodes **53.6 tok/s** on the same card, same flags, interleaved runs.

The same A/B with the MTP draft layer (`--spec 4 --spec-min-p 0.5 --mtp`) is the other way round: 77.3 against 69.7
tok/s. So this is not "0.1.40 is slower on this backend" - it is specifically the no-draft-layer path.

### Numbers

Coder IQ1_M, 12,288/12,288 experts resident, `--stream-experts`, INT8 KV, greedy 256 tokens, run interleaved
(v1, v2, v1, v2, v1, v2) from the command line, `STRATA_VERIFY_EAGER=1` on both:

| run | 0.1.39-sycl | 0.1.40-sycl |
|---|---|---|
| `--spec 4 --mtp` | 69.86, 69.64, 69.69 | 77.30, 77.29, 77.24 |
| `--spec 2` | 53.64, 53.90, 53.38 | 21.98, 22.85, 22.32 |

Three pairs each, tight spread, and the `--spec 2` gap is 2.4x. Same machine, same model, same flags, the only
difference the engine build.

### A measurement caveat, which may matter more than the regression

Short generations are not comparable on this backend. The same build and the same flags gave **48.9 and 67.5 tok/s**
on two consecutive 128-token runs, and a 155-token answer through the API was ~25% off a 256-token run of the same
build. I read a regression off those short runs first and it did not survive interleaved medians over 256 tokens.

If the numbers in docs/INTEL.md come from runs of that length, they carry that variance. Median of 3+, interleaved
against the other build, would be a safer way to state them.

### Not attributed

I have not worked out which part of 0.1.40 makes `--spec 2` slower: the drafter-free window, or the suffix drafter's
own cost per round. `STRATA_VERIFY_PROFILE=1` on 0.1.40-sycl shows a 4-token window at ~65 ms with ~24 ms of it in the
GDN hyper-connection read, nearly flat from T=2 (23.6 ms) to T=6 (25.6 ms) - a per-layer fixed ~0.5 ms over 48 layers,
which looks like dispatch rather than arithmetic. Whether that differs on the 0.1.39 build I did not measure.

### Setup

Windows 10, driver 32.0.101.8976, i9-9900 / 64 GB, conda-forge `dpcpp_win-64` 2026.1.1, OpenCL backend
(`sycl-ls` lists only `opencl:gpu`), `--spec` / `--mtp` / `--prefill 4096` / `--vram-reserve-mib 2048` /
`--kv-resident 32768` / `--max-context 131072`.

Mehr auf der Site

Links zu Install, Modellen, Releases.